Model Quantization
Quantization is a model compression technique that reduces numerical precision of weights and activations from floating-point to lower-bit representations, decreasing model size and computational cost with minimal accuracy loss.
Most models use high-precision floating-point formats like FP32 (32-bit floating point), which offer numerical stability and high accuracy but increase storage requirements, memory bandwidth usage, and power consumption.
Quantization is the process of mapping these high-precision numbers to a smaller set of lower-precision representations like 8-bit integers (INT8), 4-bit integers (INT4), or even fewer.
By reducing the precision, you shrink the model's memory footprint, lower its energy consumption, and significantly speed up inference (since integer math is faster than floating-point math), all while attempting to preserve as much of the model's original accuracy as possible
1. Post-training quantization
Post-training quantization (PTQ) reduces numerical precision after training, converting weights and activations from high-precision formats (FP32) to lower precision representations (INT8 or FP16) without retraining.
Strength of PTQ
PTQ’s key advantage is low computational cost—it requires no retraining or access to training data. However, reducing precision introduces quantization error that can degrade accuracy, particularly for tasks requiring fine-grained numerical precision.
It is fast, cheap, and requires no access to the original training data (which is great for privacy and proprietary models). You don't need a massive GPU cluster to do it.
Weaknesses of PTQ
Because the model wasn't trained to account for the sudden loss of precision, there is an inevitable drop in accuracy. This drop becomes severe as you push below 4 bits.
Suitability of PTQ
Deploying existing open-source models (like Llama) on consumer hardware for everyday local inference where you need a good balance of speed, efficiency, and acceptable quality without the massive cost of retraining.
2. Quantization-aware training
QAT integrates quantization constraints directly into the training process, simulating low-precision arithmetic during forward passes to allow the model to adapt to quantization effects.
This approach proves particularly important for models requiring fine-grained numerical precision, such as transformers used in NLP and speech recognition systems.
Strength of QAT
Recovers almost all the accuracy lost in PTQ. It allows models to be pushed to much lower bit-widths (sub-4-bit) while remaining highly performant.
Weaknesses of QAT
Computationally expensive and data-intensive. You are essentially training the model from scratch (or undertaking a heavy fine-tuning process), requiring access to large datasets and massive compute resources.
Suitability of QAT
Pushing models to extreme efficiency limits (like 2-bit or lower) where PTQ would completely break the model, or when deploying on specialized edge devices where maximum performance-per-watt is critical and training budgets are large.
3. Extreme quantization
Extreme quantization techniques use 1-bit (binarization) or 2-bit (ternarization) representations to reduce in memory usage and computational requirements
Binary Quantization
Binarization constrains weights and activations to two values (typically -1 and +1, or 0 and 1), reducing model size and accelerating inference on specialized hardware like binary neural networks as bitwise logic is incredibly cheap and fast.
Strengths of Binary Quantization
Since weights are binary, expensive floating-point multiplications are entirely replaced by simple bitwise operations like XNOR (exclusive NOT OR) and bit counting.
Weaknesses of Binary Quantization
Restricting a network to just two states causes massive gradient mismatch during training and significant accuracy drops as it severely limits model expressiveness.
Difficult to train and often fail to match full-precision performance on tasks requiring high precision such as image recognition or natural language processing.
Ternary Quantization
Ternarization extends binarization by allowing three values (-1, 0, +1), providing additional flexibility that slightly improves accuracy over pure binarization. The zero value enables greater sparsity while maintaining more representational power. Both techniques require gradient approximation methods like Straight-Through Estimator (STE) to handle non-differentiable quantization operations during training, with QAT integration helping mitigate accuracy loss.
Ternary quantization eliminates multiplication. You only perform additions and subtractions. Crucially, the introduction of the 0 state allows the network to completely ignore certain connections, effectively acting as built-in sparsity (ignoring useless data).
Strength of ternary quantization
Achieves nearly the same hardware efficiency as binary networks (no multiplications) but recovers significant accuracy.
Weakness of ternary quantization
Generally requires QAT (training from scratch with ternary constraints) to work effectively, which is expensive. (Though some new methods, like CAT-Q, are exploring PTQ for ternary networks). The software ecosystem and hardware kernels optimized for native ternary operations are also still maturing
Conclusion on extreme quantization
Performance maintenance proves difficult with such drastic quantization, requiring specialized hardware capable of efficiently handling binary or ternary operations. Traditional processors lack optimization for these computations, necessitating custom hardware accelerators.
Accuracy loss is a major concern. These methods suit tasks where high precision is not critical or where QAT can compensate for precision constraints. Despite challenges, the ability to drastically reduce model size while maintaining acceptable accuracy makes them attractive for edge AI and resource constrained environments.
