A language model never writes a sentence. At each step it hands over a probability for every token in its vocabulary, and something else, the decoder, has to choose one. Always take the most likely token and a strong model can get stuck repeating "I don't know." Sample from the whole distribution and it drifts into nonsense. Holtzman and colleagues proposed a middle way: sample only from the nucleus, the few tokens that hold most of the probability.
Spotted in the wild
- “u sub i”The logit of token : the model's raw, unnormalised score.
- “temperature”Divides every logit before the softmax; leaves the model's distribution unchanged.
- “p sub i at temperature T”Probability of token after temperature scaling.
- “size of the vocabulary”Number of tokens the model can emit at each step.
- “top k vocabulary”The most probable tokens at this step.
- “top p vocabulary, the nucleus”The smallest set of most probable tokens whose total probability is at least .
- “p base times p max”The min- threshold: tokens less probable than this fraction of the top token are removed.
- “entropy of p”Spread of the next-token distribution in bits, between 0 and .
- “beam width”Number of partial sequences beam search keeps at each step.
| Symbol | Say it | Meaning | LaTeX |
|---|---|---|---|
| “u sub i” | The logit of token : the model's raw, unnormalised score. | ||
| “temperature” | Divides every logit before the softmax; leaves the model's distribution unchanged. | ||
| “p sub i at temperature T” | Probability of token after temperature scaling. | ||
| “size of the vocabulary” | Number of tokens the model can emit at each step. | ||
| “top k vocabulary” | The most probable tokens at this step. | ||
| “top p vocabulary, the nucleus” | The smallest set of most probable tokens whose total probability is at least . | ||
| “p base times p max” | The min- threshold: tokens less probable than this fraction of the top token are removed. | ||
| “entropy of p” | Spread of the next-token distribution in bits, between 0 and . | ||
| “beam width” | Number of partial sequences beam search keeps at each step. |
Logits in, one token out
At each step the model produces a logit for each of the tokens (the next-token prediction chamber built them), the softmax turns logits into probabilities, and a decoding strategy turns probabilities into one token. Append it, run the model again, and repeat until something says stop.
Greedy decoding takes every time. It is cheap and, in principle, deterministic. But it maximises one step at a time, not the whole sequence: a token that wins now can lead to a context where every continuation is mediocre. In the third coding problem below, you will measure how far greedy falls behind the most probable sequence.
Beam search, and why it degenerates
Beam search keeps the highest-scoring partial sequences. At each step it extends every beam by every token, scores each candidate by its total log-probability, and keeps the best . Width is greedy. It costs about times the compute and is still not exact.
Total log-probability favours short outputs, since each extra token adds a negative term, so practical beam search normalises by length and retires beams that emit the end-of-sequence token. Even so, more search is not always better: with an exact search, Stahlberg and Byrne (2019) found that a translation model's single best output was the empty translation for more than half of the test sentences.
Beam search suits tasks where the input pins the output down, such as translation, summarisation and speech recognition: there, the most probable answer is usually close to the right one. On open-ended text it degenerates. Holtzman et al. show that the probability of a repeated phrase rises with each repetition, a positive feedback loop that a likelihood-seeking search walks straight into. The text it finds is also bland, because human text is not the most probable text: its per-token probabilities swing between likely and surprising.
Why does beam search tend to produce repetitive, generic text from open-ended prompts?
Temperature
Temperature divides every logit before the softmax. Let be the index of the largest logit and multiply above and below by :
The second form is also how stable code computes it. Now take limits.
- . If the maximum is unique, every term with has an exponent that tends to , so it vanishes, while the term is always . Hence : a one-hot vector on the argmax, which is greedy decoding. Implementations special-case as greedy rather than divide by zero.
- . Every exponent tends to 0, every weight to 1, and : uniform.
- Order is preserved. For any , exceeds 1 exactly when . Temperature reshapes the distribution but never re-ranks it, so on its own it cannot remove the tail.
Entropy rises steadily in between. Writing , one can show , so in nats: from 0 bits as up to bits as . At you sample the model's own distribution.
Logits are . What is the probability of the first token at temperature ? Give three decimal places.
Truncation: top-k, top-p and min-p
A vocabulary of 100,000 tokens has a long tail. Each tail token is unlikely, but together they hold real mass, and one bad draw can derail everything after it. Holtzman et al. call this the unreliable tail. Truncation samplers keep a set and renormalise:
That is exactly conditioning, , so the ratios between survivors are unchanged.
| Method | Keeps | Adapts to the distribution? | Common values |
|---|---|---|---|
| Top- (Fan et al., 2018) | the most likely tokens | No | 10–50 |
| Top- or nucleus (Holtzman et al., 2020) | the smallest set with mass at least | Yes, through its mass | 0.9–0.95 |
| Min- (Nguyen et al., 2025) | tokens with | Yes, through the top token | 0.05–0.1 |
Fan et al. generated stories by sampling from the most likely words, and found it substantially more effective than beam search. A fixed is the wrong size most of the time, though: too many candidates when the model is sure, too few when hundreds of words fit. The nucleus grows and shrinks with the distribution; the proof below shows it is always a non-empty prefix of the sorted tokens. Min- scales its cut-off by the model's confidence instead. With probabilities , top- at 0.9 keeps three tokens, while min- at 0.1 sets the threshold at 0.08 and keeps one. Its authors argue this lets you raise the temperature for creativity while staying coherent.
Sorted next-token probabilities are . How many tokens does nucleus sampling with keep?
The sampler, step by step
A typical sampler edits the logits, applies temperature, truncates and renormalises, then draws:
def next_token(u, history, rng, T=0.8, top_k=50, top_p=0.95, min_p=0.0,
bias=None, freq=0.0, pres=0.0):
u = u.copy()
for tok, b in (bias or {}).items(): # logit bias
u[tok] += b
c = np.bincount(history, minlength=len(u)) # counts so far
u -= freq * c + pres * (c > 0) # frequency and presence penalties
if T == 0:
return int(np.argmax(u)) # greedy
p = softmax(u / T) # temperature
p = renormalise(p * top_k_mask(p, top_k)) # truncate, renormalise, repeat
p = renormalise(p * top_p_mask(p, top_p))
p = renormalise(p * (p >= min_p * p.max()))
return int(rng.choice(len(p), p=p))
Around it, the loop stops at the end-of-sequence token, at a stop sequence (a string such as </answer> that ends generation when it appears and is usually left out of the output), or at max tokens, a hard cap that cuts the output off mid-sentence. Many APIs report which one fired as a finish reason: check it before you trust a truncated answer.
The order is common, not universal. At the time of writing, vLLM applies penalties, then temperature, then min-, then top- and top-, while llama.cpp's default chain truncates first, applies temperature last, and lets you reorder the chain with --samplers. The min- authors found that two widely used libraries applied temperature at different points, which changed their high-temperature results. Compare settings across tools only after checking the order.
Penalties fight repetition by editing the logits of tokens already generated. With the count of token so far, frequency and presence penalties give : the first grows with each use, the second is a one-off charge for appearing at all. The repetition penalty from CTRL (Keskar et al., 2019) divides the logits of earlier tokens by ; implementations multiply negative logits instead, so the penalty always pushes down. Penalties also hit legitimate repeats: a variable name, a JSON key, the subject of a story. Logit bias adds a fixed to chosen tokens: a large negative bias bans a token, a positive one encourages it. It acts on tokens, not words, so check how your words tokenise.
Log-probabilities as a confidence signal
Many APIs and open-weight servers can return each output token's log-probability, often with the top alternatives. Summed, they give the log-probability of the whole output. For a one-token label they give a ready-made confidence score, and a small gap between the top two labels flags a borderline case. Treat them as a signal, not as truth. They measure the model's belief about the next token, not whether a fact is correct, and depending on the tool they are computed before or after your sampling settings reshape the distribution. Post-training also bends them: the GPT-4 technical report shows a pre-trained model whose answer probabilities were well calibrated on a multiple-choice benchmark, and a post-trained model that was markedly less so. Before you threshold on log-probabilities, bin outputs by confidence, compare with accuracy on labelled data, and recalibrate if needed.
Speculative decoding
Decoding is slow because every token needs a forward pass of a large model, one after another. Speculative decoding (Leviathan, Kalman and Matias, ICML 2023) lets a small draft model propose tokens, then runs the target model once over all of them in parallel. Each draft token is accepted with probability ; at the first rejection, a replacement is drawn from the normalised . This rejection scheme makes every token distributed exactly as a sample from , so quality is untouched and only the speed depends on how often the draft agrees. The paper reports 2–3× faster decoding of T5-XXL with identical outputs.
Settings by task
| Task | Temperature | Truncation | Notes |
|---|---|---|---|
| Extraction, classification | 0 | none needed | Pair with structured output (two chambers on); use log-probabilities for confidence |
| Code | 0–0.2 for one answer, about 0.8 when sampling many to test | top- 0.95 | In the Codex paper, a 679M-parameter model did best at for pass@1 and for pass@100 |
| Chat, assistants | 0.5–0.8 | top- 0.9–0.95 or min- 0.05 | Provider defaults are tuned for this |
| Creative writing | 0.8–1.2 | min- 0.05–0.1 | A light frequency or presence penalty against loops |
| Brainstorming | 1.0–1.5 | min- 0.1 | Generate many, deduplicate, then filter |
These are illustrative starting points, not laws: tune them on your own evals. Two cautions. Some reasoning models fix their sampling settings or ignore them: depending on the provider, a custom temperature or top- is rejected, silently ignored, or restricted while extended reasoning is on, and open-weight reasoning models usually publish recommended settings in their model cards. Read the documentation for the exact model you call. And temperature 0 is not a reproducibility guarantee: why identical requests can still differ, and what to do about it, is the next chamber.
Read beyond
Article · free online · ~25 min
How to generate text: using different decoding methods for language generation with TransformersPatrick von Platen (Hugging Face) · Greedy search, beam search, top-k and top-p sampling
Run greedy and beam search on one prompt, then mark every repeated phrase in each output.
Paper · free online · ~15 min
Hierarchical Neural Story GenerationFan, Lewis & Dauphin (ACL 2018) · Section 5.4: generation
List the reasons the authors give for sampling from the top 10 words instead of using beam search.
Paper · free online · ~20 min
Turning Up the Heat: Min-p Sampling for Creative and Coherent LLM OutputsNguyen, Baker, Neo, Roush, Kirsch & Shwartz-Ziv (ICLR 2025) · Section 3: min-p sampling
Pick a distribution of your own and find a temperature at which top-p 0.9 keeps junk that min-p 0.1 removes.
Paper · free online · ~20 min
Fast Inference from Transformers via Speculative DecodingLeviathan, Kalman & Matias (ICML 2023) · Section 2.3 and the appendix: speculative sampling and its correctness
Check on a two-token example that accepting with probability min(1, p/q) and resampling from max(0, p − q) gives back p.
Read the equation in context
The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes & Yejin Choi · ICLR, 2020Section 3 sets out every decoder the paper compares. Eq. (2) defines the top- vocabulary as the smallest set reaching mass . Eq. (3) divides by inside it, the conditioning derived above, and Section 3.2 reuses Eq. (3) for top-. Section 3.3 writes temperature as a rescaled softmax, Eq. (4), noting that lowers the mass in the tail at a cost in diversity. Decode Eq. (4), then read Figure 2 (per-token probabilities of beam search and of human text) and Figure 4 (the repetition loop) in this chamber's terms.
Decode the paper · Eq. (4), Section 3.3: sampling with temperature
The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes & Yejin Choi · ICLR, 2020
Section 3 sets out the decoders the paper compares. Eq. (2) defines the nucleus, Eq. (3) renormalises inside it, and Section 3.2 reuses Eq. (3) for top-. Section 3.3 then writes temperature as a rescaled softmax and notes that skews the distribution towards likely tokens and lowers the mass in the tail, at a cost in diversity.
Options
Your turn
Ten candidate tokens, fixed logits, four knobs. Shape the distribution three ways, sample 20 continuations to hear what each setting sounds like, and see how far you can push the temperature.
Interactive lab
Nucleus navigator
The lighthouse keeper opened the door and saw a ▁
Solid bars: the final distribution. Dashed outlines: after temperature, before truncation. Rose tokens are junk.
Survivors
10 / 10
Top token
39.5%
Entropy
2.39 bits
Mass kept
100.0%
Entropy can reach at most log₂ 10 ≈ 3.32 bits over these tokens.
- Three-way race. Exactly three tokens survive, and the most likely one has between 40% and 45%.
- Wide but clean. All three junk tokens are removed, yet the entropy is at least 2.6 bits.
- Greedy at temperature 1. At T = 1, with top-k and min-p off, only one token survives: top-p alone makes decoding greedy.
Match · Method ↔ What it does
Decoding methods
Options
Match · Expression ↔ Meaning
Knobs on the logits
Options
Proof puzzle
Zero temperature is greedy
Claim
If logit is strictly larger than every other logit, prove that tends to the one-hot vector on as .
Tap lines in the order they should appear. Not every line belongs. Tap a line in your proof to send it back.
Your proof
- Pick the first line below.
Available lines
Prove it yourself
The nucleus is a prefix
Claim
Let be a next-token distribution sorted in decreasing order, and let . Prove that a smallest set of tokens with total probability at least exists, is non-empty, and can be taken to be a prefix of the sorted list.
Your typeset proof appears here.
Coding problems
Problem 10·Warm-up
Entropy at a temperature
A model's logits for eight candidate tokens are . Apply temperature and compute the entropy of the resulting distribution in bits, to 4 decimal places.
Problem 11·Standard
How big is the nucleus?
Generate logits with the Lehmer generator , starting from . Each draw with gives the logit . Build 200 decoding steps, each from the next 50 draws in order (the logits of tokens 0 to 49). At every step apply temperature , then find the nucleus for : the smallest set of tokens, taken in decreasing order of probability, whose probabilities sum to at least 0.9. Report the sum of the 200 nucleus sizes.
Problem 12·Challenge
Greedy is not the mode
A bigram language model over tokens 0 to 4 has next-token logits for current token (row) and next token (column), turned into probabilities by a softmax over each row: Starting from token 0, generate exactly 12 more tokens. Greedy decoding takes the most probable next token at each step. Compute in nats, to 6 decimal places. There are continuations, so search cleverly.
Key takeaways
- Greedy and beam search maximise probability, and on open-ended text the most probable continuations are bland and repetitive.
- Temperature divides the logits: is greedy, is uniform, and the ranking never changes.
- Top-, top- and min- cut the unreliable tail and renormalise, which is conditioning on the kept set.
- Samplers differ in their order of operations, penalties and log-probabilities: check before comparing settings across tools.
- Choose settings by task, near 0 for extraction and higher with truncation for creative work, then confirm them with evals.
Checkpoint
Prove it to the labyrinth
Answer every question to clear this chamber. First-try answers earn the most XP.
With min- and , how many tokens survive from the probabilities ?
As , what does the temperature-scaled softmax tend to?
A distribution is . After top- truncation with and renormalising, what is the first token's probability?
You extract invoice fields into JSON and want the most likely answer on every call. Which setting is the best starting point?
What distinguishes a presence penalty from a frequency penalty?
In speculative decoding, a small draft model proposes tokens and a large target model verifies them. What is the distribution of the output?
A vocabulary has 1024 tokens. With no truncation, what entropy in bits does the next-token distribution approach as ?
A post-trained chat model gives the token of its answer probability 0.97. What can you safely conclude?
End of the chamber
Clear this chamber
- Questions in this chamber (0/11 solved)Next unsolved
- Bonus: Nucleus navigator (+40 XP)
- Bonus: Problem 10: Entropy at a temperature (+20 XP)
- Bonus: Problem 11: How big is the nucleus? (+35 XP)
- Bonus: Problem 12: Greedy is not the mode (+50 XP)
- Bonus: Proof: Zero temperature is greedy (+25 XP)
- Bonus: Proof: The nucleus is a prefix (+35 XP)
- Bonus: Decode the paper (+25 XP)
- Bonus: Match: Decoding methods (+20 XP)
- Bonus: Match: Knobs on the logits (+20 XP)