Large Language Models
Deep dives into transformer internals, attention maths, pre-training at scale, alignment (SFT, RLHF, DPO), inference engineering and building LLM applications.
6 animated lectures · 66 minutes
▶ Start with lecture 1Inside a Large Language Model
Follow a sentence through a GPT-style model: tokens, embeddings, the residual stream, attention and MLP blocks, layer norms, unembedding and sampling — with real parameter counts.
The Attention Mechanism in Depth
Queries, keys and values; scaled dot-product attention computed by hand; masking; multi-head attention; positional encodings and RoPE; the quadratic cost and how FlashAttention, GQA and sliding windows tame it.
Pre-training LLMs at Scale
The recipe behind base models: web-scale data pipelines, next-token loss and perplexity, scaling laws and compute budgets, Chinchilla, distributed training, learning-rate schedules and what can go wrong.
Aligning LLMs: Instruction Tuning, RLHF and DPO
How a text predictor becomes a helpful assistant: supervised fine-tuning, preference data, reward models, RLHF with PPO and a KL leash, DPO, AI feedback, reward hacking and evaluation.
LLM Inference Engineering
Why serving LLMs is hard and how it is made fast and cheap: prefill vs decode, memory bandwidth, the KV cache, continuous batching, PagedAttention, quantisation, speculative decoding and caching.
Building with LLMs: Prompting, RAG and Agents
The practical toolkit: prompt design and in-context learning, chain-of-thought, structured outputs, retrieval-augmented generation end to end, tool-using agents, evaluation, prompt injection and cost.