How Large Language Models Work — lecture notes
ChatGPT, Claude and friends predict the next token over and over. Watch a model choose words from probabilities, one token at a time.
0:001. Introduction

Large language models can write essays, code and poems. Underneath, they do one thing over and over. They predict the next piece of text.
0:112. The pipeline

First the text is split into tokens. Each token becomes a vector of numbers called an embedding. Many transformer layers use attention to mix information between tokens. Finally the model outputs a probability for every possible next token.
0:273. Generating text

Watch it write. Given the prompt, the model scores every candidate for the next token. It picks one, adds it to the text, and repeats with the longer text. To, build, small, projects. These probabilities are illustrative, but the process is exactly how generation works.
0:464. Training

During pre-training, the model predicts the next token across an enormous amount of text, and every mistake slightly adjusts billions of parameters. Then fine-tuning and human feedback teach it to follow instructions and be helpful.
1:025. Strengths and limits

Language models are excellent at writing, explaining and coding. But they generate plausible text, not verified truth. They can confidently state false things, and their knowledge stops at a cut-off date. Always check important facts.
1:176. Recap

To recap. Text becomes tokens, embeddings, attention and probabilities. The model generates one token at a time. It learns from huge text, then from instructions. And plausible is not the same as true.
Key takeaways
- An LLM predicts the next token repeatedly to generate text.
- Tokens become embeddings, attention mixes context, and the output is a probability for each possible token.
- Pre-training on text is followed by instruction fine-tuning and human feedback.
- LLM output is plausible, not guaranteed true — verify important facts.
Check yourself
- What does a language model output at each step?
Show answer
A probability for each possible next token — It produces a probability distribution over the next token.
- What happens after a token is chosen?
Show answer
It is added to the text and the model predicts again — Generation is repeated, one token at a time.
- Why can language models “hallucinate”?
Show answer
They generate plausible text rather than looking up verified facts — They are trained to produce likely text, which is not always true.
Go deeper
- Large Language Models: What They Are and How They Are Built · The AI Lecture Hall
- The GPT Family: Autoregressive Language Models from GPT-1 to Today · The AI Lecture Hall
- Decoding Strategies: Greedy, Beam Search, Temperature, Top-k and Top-p · The AI Lecture Hall
- Hallucination in LLMs: Causes, Detection and Mitigation · The AI Lecture Hall
© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/how-large-language-models-work.html