Quantisation: Smaller, Faster Models
Store each weight in fewer bits. See weights snap to int8 and int4 levels, and how a 7B model shrinks from 28 GB to 3.5 GB.
๐ Illustrated notes ยท every chapter as a picture ยท printable
Quick quiz
3 questions to check your understanding.
Go deeper
University-level written lectures in The AI Lecture Hall:
Transcript
Introduction. Large models are heavy. A seven billion parameter model in full precision needs about 28 gigabytes just for its weights. Quantisation makes it much smaller.
Snapping to levels. Normally each weight is a 32-bit floating point number. Int8 quantisation snaps every weight to one of 256 levels. Int4 uses only 16 levels. Memory for a seven billion parameter model falls from 28 gigabytes to 14 in 16-bit, 7 in 8-bit and 3.5 in 4-bit.
The formula. Each weight is divided by a scale factor and rounded to the nearest integer. To use it, multiply back by the scale. The rounding error is the price we pay. Good methods choose scales per group of weights to keep errors small.
Trade-offs. Quantisation cuts memory and often speeds up inference, making large models run on laptops and phones. The cost is a small drop in quality, larger at very low bit widths, and outlier weights need careful handling.
Recap. To recap. Quantisation stores weights in fewer bits, shrinking a seven billion parameter model from 28 gigabytes to 3.5, with a small cost in quality.