LLM Inference Engineering — lecture notes
Why serving LLMs is hard and how it is made fast and cheap: prefill vs decode, memory bandwidth, the KV cache, continuous batching, PagedAttention, quantisation, speculative decoding and caching.
0:001. Introduction

Training a model happens once, but serving it happens billions of times. Every token a chatbot writes costs time and money. In this deep dive we look at why language model inference is hard, and the engineering tricks, from the KV cache to speculative decoding, that make it fast and affordable.
0:212. Two phases

Inference has two phases. In prefill, the whole prompt is processed in one parallel pass, which keeps the GPU’s arithmetic units busy. In decode, the reply is generated one token at a time, each depending on the one before. These phases stress the hardware in very different ways.
0:423. The life of a request

Follow one request. The prompt is tokenised. Prefill runs the whole prompt through the model and builds a cache of keys and values. The decode loop then produces one token per step, reusing that cache, and each token is detokenised and streamed back to the user immediately.
1:024. Metrics

Four numbers matter. Time to first token, dominated by prefill, decides how responsive the chat feels. Time per output token, dominated by decode, decides reading speed. Throughput counts tokens per second across all users. And cost per million tokens is what the business ultimately cares about.
1:215. Pause and think

Pause and think. With a very long prompt, why does the first token take noticeably longer than each later one? The first token needs the whole prefill pass, which processes every prompt token and builds its keys and values. After that, each step only processes one new token.
1:426. The memory wall

Why is decoding slow? To produce a single token, the GPU must read every one of the model’s weights from memory. For a seven billion parameter model in sixteen bit precision, that is fourteen gigabytes per token. At a batch size of one, the arithmetic units mostly sit idle, waiting for memory.
2:047. Pause and think

Pause and think. A GPU reads memory at about one terabyte per second, and the model is fourteen gigabytes. Roughly how many tokens per second can it decode for a single user? About a thousand divided by fourteen, roughly seventy tokens per second, an upper bound set purely by reading the weights.
2:268. The KV cache

The first essential trick is the KV cache. Without it, generating each new token would recompute keys and values for the entire sequence, so work would grow quadratically. With it, the keys and values of earlier tokens are stored and reused, and only the newest token is computed at each step.
2:479. KV cache size

How big does the cache get? Two, for keys and values, times the number of layers, times the key value width, times bytes per number. For Llama two seven B that is two times thirty two times four thousand and ninety six times two bytes: half a megabyte for every single token of context.
3:1010. Pause and think

Pause and think. Serving sixteen users at once, each with a four thousand token context, how much cache memory is needed? Half a megabyte times four thousand tokens is two gigabytes per user, so thirty two gigabytes in total, more than the fourteen gigabytes of weights themselves.
3:3011. Batching

The main lever for throughput is batching. If many users’ sequences advance together, the weights are read from memory once per step and reused for every sequence in the batch. A step that was memory bound for one user becomes efficient for thirty two, multiplying the tokens produced per second.
3:5112. Continuous batching

But replies have different lengths. With static batching, a batch waits for its longest request, leaving slots idle. Continuous batching fills a slot the moment its request finishes. On the same twelve requests, continuous batching finishes in fourteen steps instead of nineteen, raising utilisation from sixty two to eighty four percent.
4:1313. PagedAttention

Continuous batching creates a memory problem: requests grow unpredictably, and reserving contiguous cache space wastes memory. PagedAttention, introduced by the vLLM project in 2023, stores the cache in small pages, like an operating system’s virtual memory. Almost nothing is wasted, and throughput improved two to four times.
4:3314. Prefix caching

Many requests start with the same text, such as a long system prompt or a shared document. Prefix caching computes the key value cache for that shared beginning once and reuses it across requests. This cuts both the time to first token and the cost for applications with long, repeated prompts.
4:5415. Quantisation

Quantisation stores weights with fewer bits. Rounding sixteen bit weights to eight or four bits shrinks memory, and because decoding is memory bound, fewer bytes per weight also means faster tokens. Careful methods calibrate the rounding so accuracy drops only slightly.
5:1216. Memory by precision

For a seven billion parameter model, thirty two bit weights need twenty eight gigabytes. Sixteen bit halves that to fourteen. Eight bit integers need seven, and four bit only three and a half gigabytes, small enough for a laptop GPU or even a high end phone, with a modest loss in quality.
5:3417. Lower precision hardware

Hardware keeps adding lower precision formats. Recent data centre GPUs support eight bit floating point arithmetic natively, which can roughly double throughput compared with sixteen bit for models that tolerate it. Choosing number formats and efficient kernels matters as much as choosing the model.
5:5318. Grouped-query attention

The architecture itself can help. Grouped query attention lets many query heads share a few key and value heads. Llama two seventy B uses eight key value heads for sixty four query heads, making its cache eight times smaller, so many more users fit on the same hardware.
6:1319. Speculative decoding

Speculative decoding attacks the one token per step bottleneck. A small, fast draft model proposes several tokens. The large model checks all of them in a single parallel pass, accepts the ones it agrees with, and supplies a correction. Here three rounds produce twelve tokens with only three expensive passes.
6:3520. Lossless speed-up

Remarkably, speculative decoding does not change the output. A careful acceptance rule, based on rejection sampling, guarantees the text follows exactly the large model’s distribution. The speed up, often two to three times, depends on how often the draft model guesses correctly.
6:5321. Mixture of experts

Mixture of experts models change the trade off too. Only a few experts run for each token, so compute per token is modest even when total parameters are huge. Mixtral eight by seven B has about forty seven billion parameters but uses only about thirteen billion per token. All the experts still have to fit in memory, though.
7:1722. Decoding settings

Decoding settings affect quality as well as speed. Top k keeps only the k most likely tokens, and top p, nucleus sampling, keeps the smallest set whose probabilities add up to p. Lower temperature and smaller sets give focused, predictable text, while higher values give variety. Stop sequences and maximum lengths bound the cost.
7:3923. Distillation

Often the biggest saving is simply using a smaller model. Knowledge distillation trains a small student to imitate a large teacher’s outputs. For focused tasks, such as classification, extraction or routine support questions, a well tuned model with a few billion parameters can match a far larger one at a fraction of the cost.
8:0224. Serving many fine-tunes

Many products need dozens of specialised variants of one model. Instead of loading a separate copy for each, the base model is loaded once and small LoRA adapters are applied per request. Specialised serving systems can batch requests that use different adapters together, making thousands of fine tunes affordable.
8:2325. On-device inference

Inference is also moving onto devices. Quantised models with a few billion parameters run on laptops and phones using optimised kernels for CPUs, GPUs and neural processing units. There is no server round trip, data stays private, and the application works offline, at the cost of smaller, less capable models.
8:4426. Autoscaling the fleet

At the fleet level, serving looks like any other web service. Traffic rises and falls through the day, and an autoscaler adds and removes model replicas. When load rises faster than replicas can start, queues build and latency spikes, as here around midday. GPU startup time makes this especially tricky for language models.
9:0727. Serving with vLLM

In practice you rarely build this from scratch. With an inference engine such as vLLM, you load the model once, choose sampling settings like temperature, top p and a maximum length, and submit many prompts. Continuous batching and the paged key value cache are handled automatically.
9:2628. Latency vs throughput

Every serving system balances latency against throughput. Small batches give each user a fast first token and quick streaming, but cost more per token. Large batches keep GPUs busy and minimise cost, but each user waits a little longer. Interactive chat and overnight batch jobs sit at opposite ends.
9:4729. Pause and think

Pause and think. Your chatbot prepends the same three thousand token policy document to every request, and the first token is slow to appear. What should you try first? Prefix caching: compute the cache for that shared document once, so each request only has to prefill its own new tokens.
10:0830. Cost levers

To cut costs, pull these levers in order. Use the smallest model that meets your quality bar. Quantise the weights. Batch continuously with a paged cache. Cache shared prefixes, and whole responses when questions repeat. And keep prompts and outputs as short as the task allows.
10:2831. Recap

To recap. Prefill is compute bound, while decoding is limited by memory bandwidth. The KV cache avoids recomputation but consumes memory. Continuous batching with a paged cache raises throughput, quantisation and grouped query attention shrink memory, and speculative decoding produces the same output with fewer expensive passes.
Key takeaways
- Inference has a parallel prefill phase and a sequential, memory-bound decode phase.
- Key metrics: time to first token, time per output token, throughput and cost per token.
- The KV cache stores past keys/values; for Llama 2 7B it costs ~0.5 MB per token in fp16.
- Continuous batching and PagedAttention keep GPUs busy with little wasted memory.
- Quantisation (INT8/INT4) and grouped-query attention reduce memory and speed up decoding.
- Speculative decoding uses a draft model to propose tokens that the large model verifies — exactly preserving its output distribution.
Check yourself
- Why is decoding at batch size 1 usually memory-bound?
Show answer
Every weight must be read from memory for each new token — Arithmetic waits on memory bandwidth.
- What does the KV cache store?
Show answer
Keys and values of previous tokens so they are not recomputed — Only the newest token needs fresh keys/values.
- What problem does continuous batching solve?
Show answer
Idle GPU slots while a batch waits for its longest request — Finished slots are refilled immediately.
- A 7B model quantised to INT4 needs roughly…
Show answer
3.5 GB — 0.5 bytes × 7B ≈ 3.5 GB.
- Speculative decoding changes the model’s output distribution…
Show answer
Never, with the correct acceptance rule — Rejection sampling keeps it exact.
Go deeper
- LLM Inference: KV Caching, Batching and Speculative Decoding · The AI Lecture Hall
- Quantising Large Language Models for Efficient Inference · The AI Lecture Hall
- Decoding Strategies: Greedy, Beam Search, Temperature, Top-k and Top-p · The AI Lecture Hall
- Model Serving: Batch, Real-Time APIs and Streaming · The AI Lecture Hall
© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/llm-inference-engineering.html