Home / Chapter 7 · Production and serving
    Last edited · 6 min read

    Use with AI

    Quantisation

    Quantisation stores the weights, and sometimes the activations and the KV cache too, in fewer bits by rounding them to a coarser grid of values. The model takes 2–4 times less memory, and generation gets faster because decode reads fewer bytes from memory.

    In plain wordsA photo saved as a JPEG instead of at full quality: it is several times smaller and looks the same on a phone screen. Artefacts appear only under heavy compression, and first in the fine detail. For a model, the fine detail is long reasoning, code and less common languages.

    Lower the bit count, add an outlier, then turn on a separate scale per block. Watch the rounding error

    levels on the grid
    rounding error relative to weight size
    rule of thumb for quality

    40 weights from a normal distribution, a grid symmetric around zero, a block is 8 weights. The error leaves out the outlier itself, so you can see what happens to the rest.

    Pick a model size and see what hardware the weights alone fit on. Change the bit count in the widget above

    for the weights alone
    at 16 bits

    Counted as the weights plus 15% headroom, enough for one person’s short context. Serving many conversations needs much more memory for the KV cache. Block scales add a few per cent.

    How it works

    Formats and what they speed up

    What you lose and how to decide

    Check yourself

    What does quantisation give you, what do you lose, and which format would you choose?

    The weights are stored in fewer bits, from 16 down to 8 or 4, with a scale computed per small block, because single outliers stretch the grid. The model needs 2–4 times less memory, and decode gets faster because it reads fewer bytes. Weight-only quantisation such as W4A16 helps at small batch sizes. FP8 for weights and activations also speeds up the maths, so on Hopper it wins under heavy traffic; on Blackwell, NVFP4 for weights and activations is faster still. 8 bits is practically lossless, 4 bits costs a little, and below that quality drops fast, first in reasoning and code. Your own evals decide.

    Po polsku

    Wagi zapisuje się mniejszą liczbą bitów, z 16 do 8 albo 4, ze skalą liczoną dla małych bloków, bo pojedyncze wartości odstające rozciągają siatkę. Model zajmuje 2–4 razy mniej pamięci, a decode przyspiesza, bo czyta mniej bajtów. Kwantyzacja samych wag, jak W4A16, pomaga przy małym batchu. FP8 dla wag i aktywacji przyspiesza też liczenie, więc na Hopperze wygrywa przy dużym ruchu; na Blackwellu jeszcze szybsze jest NVFP4 dla wag i aktywacji. 8 bitów jest praktycznie bezstratne, 4 bity to niewielka strata, niżej jakość szybko spada, najpierw w rozumowaniu i kodzie. Rozstrzygają własne ewaluacje.

    Follow-up questions (4)
    PTQ or QAT?
    PTQ quantises a finished model in minutes or hours, using a small sample of calibration data (GPTQ, AWQ). QAT simulates quantisation during training, so the model learns to tolerate it and holds its quality better at 4 bits and below. The cost is training, so QAT is usually done by the model’s author.
    Why are activations harder to quantise than weights?
    Weights are known in advance and can be carefully rescaled offline. Activations depend on the input and have large outliers in a few channels. Methods that shift the difficulty onto the weights (SmoothQuant) help, as does FP8, which has a wider range than INT8.
    After quantising to 4 bits the benchmarks look fine, but users complain about code. What do you do?
    A benchmark average hides losses in long reasoning chains and code. Compare both versions on your own eval set from that area. If the loss is real, go back to 8 bits or keep the sensitive layers at higher precision.
    When does quantising the KV cache give more than quantising the weights?
    With long contexts and large batches, when the cache of all conversations takes more memory than the weights. An FP8 cache holds twice as many tokens, i.e. twice as many conversations or twice the context length.

    Sources

    Report an error · Suggest a fix