AI in Motion

Generative AIBeginner1:31 video6 chapters

Decoding: Temperature, Top-k and Top-p — lecture notes

How a language model picks each word from its probabilities — and how temperature, top-k and top-p change its personality.

▶ Watch the animated lecture

0:001. Introduction

Introduction — Decoding: Temperature, Top-k and Top-p

A language model outputs probabilities for the next token. How we pick from them, called decoding, changes whether the text is safe and predictable or creative and surprising.

0:122. Temperature

Temperature — Decoding: Temperature, Top-k and Top-p

After the sky is, the model gives blue 46 percent at temperature one. Lower the temperature to 0.4, and the distribution sharpens: blue jumps to about 77 percent. Raise it to 2 and it flattens: blue falls to about 31 percent, and unusual words like purple get a real chance.

0:333. The formula

The formula — Decoding: Temperature, Top-k and Top-p

Temperature simply divides the raw scores before the softmax. Below one, differences grow and the top choice dominates. Above one, differences shrink and choices even out.

0:454. Top-k and top-p

Top-k and top-p — Decoding: Temperature, Top-k and Top-p

Top k keeps only the k most likely tokens, here three, and samples among them. Top p, or nucleus sampling, keeps the smallest set whose probabilities add up to 90 percent. Here that is four tokens. Both cut off the long tail of unlikely, often nonsensical choices.

1:055. Practical settings

Practical settings — Decoding: Temperature, Top-k and Top-p

For factual work like code or data extraction, use a low temperature for consistent answers. For brainstorming and stories, a higher temperature with top p around 0.9 gives variety without nonsense.

1:196. Recap

Recap — Decoding: Temperature, Top-k and Top-p

To recap. Decoding chooses from probabilities. Temperature sharpens or flattens them. Top k and top p trim the unlikely tail. Low for facts, higher for creativity.

Key takeaways

  • Temperature divides logits before softmax: T < 1 sharpens, T > 1 flattens.
  • In the demo “blue” is 76.7% at T = 0.4, 46.2% at T = 1 and 31.4% at T = 2.
  • Top-k keeps the k most likely tokens; top-p keeps the smallest set covering p of the probability.
  • Use low temperature for factual tasks and higher for creative ones.

Check yourself

  1. What does lowering the temperature do?
    Show answer

    Makes the top choice even more likely — Dividing by a small T exaggerates differences.

  2. Top-p = 0.9 keeps…
    Show answer

    The smallest set of tokens whose probabilities sum to at least 90% — Nucleus sampling uses cumulative probability.

  3. Which setting suits extracting data from invoices?
    Show answer

    Low temperature — Precision matters more than variety.

Go deeper

© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/decoding-temperature-and-top-p.html