Model Pruning
Structured model optimization
Structured model optimization works in two key ways:
Eliminating parameter redundancy.
Structuring computations for efficient hardware execution through techniques like gradient checkpointing and parallel processing patterns.
Key structural model optimization techniques include pruning, knowledge distillation and neural architecture search.
Pruning
Deep neural networks normally need a large number of pathways to effectively explore the loss landscape and learn complex representations. On the other hand, inference only requires a small percentage of those weights for effective performance. In most cases there are redundant parameters within the neural network that don't necessarily help with performance.
Pruning is a structured model optimization technique that removes redundant parameters from neural networks while preserving performance, reducing model size and computational cost for efficient deployment.
What can we prune?
There are three main architectural structures on which pruning can be applied.
Neuron pruning removes entire neurons along with their associated weights and biases reducing the width of a layer. it is often applied to fully connected layers.
Channel Pruning reduces the depth of feature maps which impacts the networks ability to extract certain features. Applied in image processing tasks: i.e. can be used in CNNs to eliminate entire channels or filters hence is also known as filter pruning.
Layer pruning removes whole layers from the network significantly reducing depth. It requires careful balance to ensure the model retains sufficient capacity to capture complex patterns.
Three key approaches to pruning
Unstructured pruning removes individual weights while preserving the overall network architecture. It is useful in reducing model size and memory footprint.
Structured pruning eliminates entire neurons, channels or layers leading to more hardware friendly model by reducing the number of floating point operations (FLOPs) required during inference.
Dynamic pruning introduces adaptability into the pruning process by adjusting which parameters are pruned at runtime based on input data and training dynamics. It allows for a better balance between accuracy and efficiency as the model retains the flexibility to reintroduce previously pruned parameters if need be.
Pruning strategies
Iterative pruning gradually removes structures through multiple cycles of pruning and finetuning.
One-shot pruning removes multiple architectural components in a single step followed by an extensive finetuning phase to recover model accuracy.
Lottery ticket hypothesis proposes that within large neural networks, there exists small, well-initialized sub networks referred to as wining tickets that can achieve accuracy comparative to the full model when trained in isolation. A large network is first trained to convergence, the lowest magnitude weights are then pruned and the remaining weights are reset to their original initialization rather than being re-randomized. this process is iteratively repeated , gradually reducing the model size while preserving performance.
You don't need an entire model as it is, remove what you don't need and keep what you need. That's what pruning is all about. See you in the next one.
