AI in Motion

Artificial IntelligenceBeginner1:31 video6 chapters

How Large Language Models Work — lecture notes

ChatGPT, Claude and friends predict the next token over and over. Watch a model choose words from probabilities, one token at a time.

▶ Watch the animated lecture

0:001. Introduction

Introduction — How Large Language Models Work

Large language models can write essays, code and poems. Underneath, they do one thing over and over. They predict the next piece of text.

0:112. The pipeline

The pipeline — How Large Language Models Work

First the text is split into tokens. Each token becomes a vector of numbers called an embedding. Many transformer layers use attention to mix information between tokens. Finally the model outputs a probability for every possible next token.

0:273. Generating text

Generating text — How Large Language Models Work

Watch it write. Given the prompt, the model scores every candidate for the next token. It picks one, adds it to the text, and repeats with the longer text. To, build, small, projects. These probabilities are illustrative, but the process is exactly how generation works.

0:464. Training

Training — How Large Language Models Work

During pre-training, the model predicts the next token across an enormous amount of text, and every mistake slightly adjusts billions of parameters. Then fine-tuning and human feedback teach it to follow instructions and be helpful.

1:025. Strengths and limits

Strengths and limits — How Large Language Models Work

Language models are excellent at writing, explaining and coding. But they generate plausible text, not verified truth. They can confidently state false things, and their knowledge stops at a cut-off date. Always check important facts.

1:176. Recap

Recap — How Large Language Models Work

To recap. Text becomes tokens, embeddings, attention and probabilities. The model generates one token at a time. It learns from huge text, then from instructions. And plausible is not the same as true.

Key takeaways

  • An LLM predicts the next token repeatedly to generate text.
  • Tokens become embeddings, attention mixes context, and the output is a probability for each possible token.
  • Pre-training on text is followed by instruction fine-tuning and human feedback.
  • LLM output is plausible, not guaranteed true — verify important facts.

Check yourself

  1. What does a language model output at each step?
    Show answer

    A probability for each possible next token — It produces a probability distribution over the next token.

  2. What happens after a token is chosen?
    Show answer

    It is added to the text and the model predicts again — Generation is repeated, one token at a time.

  3. Why can language models “hallucinate”?
    Show answer

    They generate plausible text rather than looking up verified facts — They are trained to produce likely text, which is not always true.

Go deeper

© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/how-large-language-models-work.html