AI in Motion

How Large Language Models Work

Artificial IntelligenceBeginner1:316 chapters

ChatGPT, Claude and friends predict the next token over and over. Watch a model choose words from probabilities, one token at a time.

πŸ“„ Illustrated notes Β· every chapter as a picture Β· printable

Shortcuts: Space play/pause Β· ←/β†’ 5 s Β· N/P chapter Β· M voice Β· C subtitles Β· F fullscreen

Quick quiz

3 questions to check your understanding.

Q1 What does a language model output at each step?
Q2 What happens after a token is chosen?
Q3 Why can language models β€œhallucinate”?

Go deeper

University-level written lectures in The AI Lecture Hall:

Transcript

Introduction. Large language models can write essays, code and poems. Underneath, they do one thing over and over. They predict the next piece of text.

The pipeline. First the text is split into tokens. Each token becomes a vector of numbers called an embedding. Many transformer layers use attention to mix information between tokens. Finally the model outputs a probability for every possible next token.

Generating text. Watch it write. Given the prompt, the model scores every candidate for the next token. It picks one, adds it to the text, and repeats with the longer text. To, build, small, projects. These probabilities are illustrative, but the process is exactly how generation works.

Training. During pre-training, the model predicts the next token across an enormous amount of text, and every mistake slightly adjusts billions of parameters. Then fine-tuning and human feedback teach it to follow instructions and be helpful.

Strengths and limits. Language models are excellent at writing, explaining and coding. But they generate plausible text, not verified truth. They can confidently state false things, and their knowledge stops at a cut-off date. Always check important facts.

Recap. To recap. Text becomes tokens, embeddings, attention and probabilities. The model generates one token at a time. It learns from huge text, then from instructions. And plausible is not the same as true.