Models · reviewed · reviewed Oct 5, 2026 · 4 min
How does model quantization work?
Quantization represents numerical values with a smaller set of codes plus decoding information such as scales. It reduces some storage and computation costs by introducing approximation. Rounding, clipping, calibration, sensitive features, and runtime support determine the practical trade-off.
Keep the network, change the representation
A model contains numerical parameters. Quantization changes how selected values are represented and used; it does not inherently remove layers, reduce the number of parameters, or create a different model architecture.
Imagine marking a measured length on a fine ruler and then replacing it with a ruler that has fewer marks. Nearby values now share a representable value. A scale tells the system how those marks relate to the original units.
For a simple uniform quantizer, the sequence is choose a range → map to a code → round and bound the code → reconstruct a value. Real methods can use groups, different precision for different quantities, and choices optimized to preserve useful model behaviour.
Move eight weights onto a coarser ruler
The experiment uses eight authored weights and a fixed range of −1 to 1. Two bits supply four code values, 0 through 3; four bits supply sixteen. Each code maps back to one level across the chosen range.
Change the bit width and inspect original weights, stored codes, reconstructed values, and one dot product against a fixed input. Then introduce a weight of 1.8 without changing the range. Predict whether additional bits can restore it.
Explore the mechanism
Move weights onto a coarser ruler
Representable levels: 16. Fixed range: −1 to 1. Level spacing: 0.1333.
- Weight payload (bytes)
- 4
- Largest absolute error
- 0.063
Original dot product: 0.8500
Quantized dot product: 0.8133
Rounding changes the numerical calculation; this does not measure downstream LLM quality.
Inspect codes, inputs, and reconstructed values
| Original / input | Code | Decoded | Error |
|---|---|---|---|
| -1 / 0.2 | 0 | -1.000 | 0.000 |
| -0.82 / -0.6 | 1 | -0.867 | 0.047 |
| -0.36 / 0.1 | 5 | -0.333 | 0.027 |
| -0.07 / 0.3 | 7 | -0.067 | 0.003 |
| 0.09 / 0.9 | 8 | 0.067 | 0.023 |
| 0.28 / -0.5 | 10 | 0.333 | 0.053 |
| 0.61 / 0.8 | 12 | 0.600 | 0.010 |
| 0.93 / 0.2 | 14 | 0.867 | 0.063 |
Payload counts packed weight codes only. It excludes scales, offsets, alignment, activations, KV state, and runtime buffers.
Eight authored weights, fixed inputs, and a uniform [-1, 1] range. Codes, reconstruction, absolute errors, payload estimates, and dot products are computed locally. No checkpoint is loaded; this is not GPTQ/AWQ or an LLM-quality benchmark.
More levels can reduce rounding error within the range. They cannot recover the portion of a value clipped outside that range: 1.8 still reconstructs as 1. The range and scaling policy are as important as the label “4-bit.”
The displayed errors and dot products are computed locally. They are not LLM accuracy, perplexity, latency, or an implementation of GPTQ, AWQ, or LLM.int8.
Compression needs its decoding information
Eight four-bit codes occupy a four-byte weight payload if packed. Reconstructing the values also needs the range or scale and offset. Actual artifacts have metadata, alignment, grouping, and sometimes values retained at higher precision.
The arithmetic is straightforward at larger scale: seven billion values at four bits each occupy 7,000,000,000 × 4 / 8 = 3.5 billion bytes of raw codes. That is 3.5 decimal GB before overhead, not a claim that a complete seven-billion-parameter model runs in 3.5 GB.
Serving also holds activations, attention/KV state, temporary buffers, runtime allocations, and concurrency-dependent data. Compressing weights does not eliminate those costs. A long-context workload can still run out of memory.
Real methods protect different parts of the calculation
GPTQ uses approximate second-order information to select weight quantization while accounting for reconstruction error. It is more deliberate than independently rounding every weight to the nearest value.
AWQ uses activation information to guide weight scaling and protect important behaviour. The size of a weight alone is not enough to explain its influence on outputs.
LLM.int8 combines an eight-bit path with higher-precision handling for important outlier features. A deployment described as low-bit can therefore contain several numerical precisions.
These methods illustrate why bit width alone is an incomplete compatibility or quality description. Distinguish what is quantized—weights, activations, or cache—from the calibration procedure, grouping, reconstruction rule, accumulation precision, and supported execution kernels.
Smaller storage is not a universal speedup
A runtime must execute the representation efficiently. Packing, unpacking, conversion, memory transfer, and available hardware kernels influence latency and throughput. A compact artifact can save memory without improving every workload's elapsed time.
Likewise, a small numerical error is not a product-quality guarantee. Repeated transformations can amplify differences; a close pair of output scores can change their ordering. Compare the original and quantized artifacts on the actual tasks, including important languages, long inputs, schemas, tools, and severe failure cases.
Keep model identity and test conditions explicit. Quantization is an engineering trade-off to measure, not a permanent claim that a smaller file is equivalent to its source in every setting.
The uniform quantizer used in this example
For bit width b, let levels = 2ᵇ, range minimum −1, and range maximum 1. The level spacing is scale = 2 / (levels − 1).
The code is round((weight + 1) / scale), bounded to 0 … levels − 1. Reconstruction is −1 + code × scale. The toy uses ordinary JavaScript numbers for calculation and estimates packed payload size; it does not store a real low-bit checkpoint.
The dot product multiplies each weight by its corresponding fixed input and sums the products. Comparing the original and reconstructed results exposes one numerical consequence. Different inputs or a different calibration range can change that consequence.
Sources
Sources and further reading
- 01GPTQ: Accurate Post-Training Quantization for Generative Pre-trained TransformersFrantar et al. · research · published Oct 31, 2022 · source checked Oct 5, 2026
Primary approximate second-order weight quantization method, illustrating calibration and reconstruction choices beyond independent nearest-level rounding.
- 02AWQ: Activation-aware Weight Quantization for LLM Compression and AccelerationLin et al. · research · published Jun 1, 2023 · source checked Oct 5, 2026
Primary activation-aware weight scaling and low-bit serving study; supports distinguishing numerical precision from sensitive features and practical kernel support.
- 03LLM.int8(): 8-bit Matrix Multiplication for Transformers at ScaleDettmers et al. · research · published Aug 15, 2022 · source checked Oct 5, 2026
Primary account of vector-wise quantization and mixed-precision treatment of outlier features, showing why an eight-bit deployment need not use one precision everywhere.
