AI in Motion

Quantisation: Smaller, Faster Models

Generative AIIntermediate1:165 chapters

Store each weight in fewer bits. See weights snap to int8 and int4 levels, and how a 7B model shrinks from 28 GB to 3.5 GB.

๐Ÿ“„ Illustrated notes ยท every chapter as a picture ยท printable

Shortcuts: Space play/pause ยท โ†/โ†’ 5 s ยท N/P chapter ยท M voice ยท C subtitles ยท F fullscreen

Quick quiz

3 questions to check your understanding.

Q1 How much memory do the weights of a 7B model need in INT4?
Q2 How many distinct levels can a 4-bit integer represent?
Q3 What is the main cost of quantisation?

Go deeper

University-level written lectures in The AI Lecture Hall:

Transcript

Introduction. Large models are heavy. A seven billion parameter model in full precision needs about 28 gigabytes just for its weights. Quantisation makes it much smaller.

Snapping to levels. Normally each weight is a 32-bit floating point number. Int8 quantisation snaps every weight to one of 256 levels. Int4 uses only 16 levels. Memory for a seven billion parameter model falls from 28 gigabytes to 14 in 16-bit, 7 in 8-bit and 3.5 in 4-bit.

The formula. Each weight is divided by a scale factor and rounded to the nearest integer. To use it, multiply back by the scale. The rounding error is the price we pay. Good methods choose scales per group of weights to keep errors small.

Trade-offs. Quantisation cuts memory and often speeds up inference, making large models run on laptops and phones. The cost is a small drop in quality, larger at very low bit widths, and outlier weights need careful handling.

Recap. To recap. Quantisation stores weights in fewer bits, shrinking a seven billion parameter model from 28 gigabytes to 3.5, with a small cost in quality.