AI in Motion

Decoding: Temperature, Top-k and Top-p

Generative AIBeginner1:316 chapters

How a language model picks each word from its probabilities — and how temperature, top-k and top-p change its personality.

📄 Illustrated notes · every chapter as a picture · printable

Shortcuts: Space play/pause · ←/→ 5 s · N/P chapter · M voice · C subtitles · F fullscreen

Quick quiz

3 questions to check your understanding.

Q1 What does lowering the temperature do?
Q2 Top-p = 0.9 keeps…
Q3 Which setting suits extracting data from invoices?

Go deeper

University-level written lectures in The AI Lecture Hall:

Transcript

Introduction. A language model outputs probabilities for the next token. How we pick from them, called decoding, changes whether the text is safe and predictable or creative and surprising.

Temperature. After the sky is, the model gives blue 46 percent at temperature one. Lower the temperature to 0.4, and the distribution sharpens: blue jumps to about 77 percent. Raise it to 2 and it flattens: blue falls to about 31 percent, and unusual words like purple get a real chance.

The formula. Temperature simply divides the raw scores before the softmax. Below one, differences grow and the top choice dominates. Above one, differences shrink and choices even out.

Top-k and top-p. Top k keeps only the k most likely tokens, here three, and samples among them. Top p, or nucleus sampling, keeps the smallest set whose probabilities add up to 90 percent. Here that is four tokens. Both cut off the long tail of unlikely, often nonsensical choices.

Practical settings. For factual work like code or data extraction, use a low temperature for consistent answers. For brainstorming and stories, a higher temperature with top p around 0.9 gives variety without nonsense.

Recap. To recap. Decoding chooses from probabilities. Temperature sharpens or flattens them. Top k and top p trim the unlikely tail. Low for facts, higher for creativity.