Knowledge distillation
Knowledge distillation involves using a large teacher model to train a smaller student model. The student model not only learns from the correct labels but also from the teacher's output distribution.
In this post we will focus on response-based knowledge distillation, which is the most prevalent technique used.
Response-Based Knowledge Distillation
The most common model distillation approach is Response-Based Knowledge Distillation. Response-Based KD distills knowledge based on responses from the teacher model.
It treats the internal workings of the neural network as a black box and operates entirely at the output layer. The core premise is that the final unnormalized log-probabilities (logits) generated by a heavily parameterized teacher model encapsulate its entire reasoning process, which the student model must replicate.
The primary limitation of response-based KD is that it only transfers what the teacher thinks, not how it arrived at that conclusion.
If the student network has a vastly different architecture or is significantly shallower, matching the final logits becomes exponentially more difficult. The student may lack the non-linear capacity to map the raw input directly to the teacher's complex output space without intermediate guidance. .
Strengths and weaknesses of RBKD
Advantages of response-based KD
Architecture agnostic: Teacher and student can have completely different internal topologies (e.g., distilling a ResNet teacher into a MobileNet student).
Low memory footprint: You only need to store and compute losses over final output vectors, not deep intermediate tensors.
Disadvantages of response-based KD
ignores internal reasoning: Treats the teacher as a black box; the student learns what the teacher predicted, but not how it derived the answer.
Capacity gap failure: If the capacity gap between teacher and student is too large, the student fails to fit the complex soft output boundary.
Response-based KD is best suited for classification tasks (vision, sentiment, topic tagging), lightweight inference edge models, and cross-architecture compression.
Loss Functions for Knowledge distillation
There are two main loss functions to consider during knowledge distillation.
Distillation loss - is normally the KL divergence between the teachers and the students soft label distributions. KL (Kullback-Leibler) Divergence measures how much one probability distribution differs from another.
Student loss - is the standard Cross-Entropy loss between the ground truth hard labels and the Student's standard predictions. It ensures the Student actually learns the primary task by correctly classifying the hard labels.
Final take
Distillation allows smaller models to retain much of the predictive power of larger models helping optimize for memory footprint required for significant performance.
A distilled model also has the advantage of requiring fewer FLOPs per inference since the student is trained to approximate the teachers knowledge in a more compact representation.
Knowledge distillation preserves accuracy better than pruning but requires higher training complexity through training a new model rather than modifying an existing one.
Combining pruning and distillation is a good option. Pruning removes unnecessary parameters before distillation optimizes the final student model.
