Skip to main content

Command Palette

Search for a command to run...

Model Optimization

Updated
•3 min read•View as Markdown
G
Am an AI engineer focused on Natural Language Processing and building NLP solutions

Model optimization is the systematic transformation of ML models to maximize computational efficiency while preserving task performance, enabling deployment across divers hardware constraints.

In most cases a trade off might have to be made between computational efficiency (speed, memory footprint) and performance (accuracy and generation quality).

There are three main dimensions to consider when optimizing an ML model:

  1. Structural efficiency in model representation (reduces what computations are performed).

  2. Numerical efficiency through precision optimization (changes how computations are executed).

  3. Computational efficiency via hardware implementation (ensures operations run efficiently on the target hardware).

Model representation optimization focuses on eliminating redundancies in ML models. Techniques include pruning, knowledge distillation and automated neural architecture search.

Numerical precision optimization changes how computations are executed by reducing the numeric fidelity of weights, activations and arithmetic operations. Quantization techniques map high-precision weights and activations to lower-bit representations. Mixed-precision training dynamically adjusts precision levels during training to strike a balance between efficiency and accuracy.

Computational efficiency optimizes the models computational graph to run on the hardware on which it is to run as efficiently as possible.

Here are some model optimization techniques:

  1. Pruning is a model optimization technique that removes redundant parameters from neural networks while preserving performance reducing model size and computational cost for efficient deployment. It can be applied in edge AI inference since it can allow us to fit models into kilobyte-scale memory without a problem.

  2. Knowledge distillation involves the use of a large pretrained teacher model to train a smaller student model. The student model not only learns from the correct and incorrect labels but also the teachers output distribution. It can be applied in a highly specialized narrow domain task. Distillation allows us to bake the reasoning capabilities of a frontier model into a lightweight domain specific model.

  3. Neural Architecture Search (NAS) is the process of automating deep learning model design. Instead of human engineers manually guessing the ideal depth, layer types, kernel sizes, and skip connections, NAS treats model architecture design as an optimization problem solved by an automated algorithm. NASNet, EfficientNet and MobileNetV3 are some renown computer vision algorithms developed with NAS.

  4. Quantization involves mapping high precision numbers (FP32) to lower precision formats (FP16, INT8, INT4). Quantization can be applied if one wants to run a model on hardware that cannot support it due to the models size. A 7B parameter model might require 28GB of VRAM, quantizing it to FP16 or BF16 brings the memory footprint down to 14GB VRAM while 4-Bit quantization reduces that further to 4GB of memory required making it easy to run on consumer hardware.

  5. Runtime optimization involves the use of tools like ONNX Runtime, OpenVino or TensorRT which analyze the entire network structure then fuses some operations together (e.g. combining a convolution step with a ReLU activation step into a single mathematical pass) and optimizing memory allocation as per the hardware on which it is running. Deploying a model through a runtime compiler ensures you get the best possible performance from the hardware you are running on, maximizing CPU throughput and minimizing latency for real time applications.

We'll delve deeper into the specifics of each individual technique in future blogs.

17 views