Skip to content
AriadneTechnology

The Middle Ring · Chamber 4 of 9

Sampling: Temperature, Top-k and Top-p

From logits to words: greedy and beam search, temperature, top-k, nucleus and min-p sampling, and when to use each.

50 min 60 XP + 11 questions + 1 challengeMathVideoPapersProofsCodeLab

In this chamber you will

  • Turn logits into probabilities at any temperature, and find the limits as T → 0 and T → ∞
  • Apply top-k, top-p and min-p truncation, then renormalise
  • Explain why greedy and beam search degenerate on open-ended text
  • Choose decoding settings for extraction, code and creative writing
DiscoverLearnRead beyondPapers & lecturesYour turn

A language model never writes a sentence. At each step it hands over a probability for every token in its vocabulary, and something else, the decoder, has to choose one. Always take the most likely token and a strong model can get stuck repeating "I don't know." Sample from the whole distribution and it drifts into nonsense. Holtzman and colleagues proposed a middle way: sample only from the nucleus, the few tokens that hold most of the probability.

Spotted in the wild

∑x∈V(p)P(x∣x1:i−1)≥p\sum_{x\in V^{(p)}}P(x\mid x_{1:i-1})\ge p
The Curious Case of Neural Text Degeneration
DiscoverLearnRead beyondPapers & lecturesYour turn
Symbols for this chamber
  • uiu_i“u sub i”
    The logit of token ii: the model's raw, unnormalised score.
    u=(3.2, 2.6, 2.3)u=(3.2,\,2.6,\,2.3)
  • TT“temperature”
    Divides every logit before the softmax; T=1T=1 leaves the model's distribution unchanged.
  • pi(T)p_i(T)“p sub i at temperature T”
    Probability of token ii after temperature scaling.
    pi(T)=eui/T/∑jeuj/Tp_i(T)=e^{u_i/T}/\sum_je^{u_j/T}
  • ∣V∣|V|“size of the vocabulary”
    Number of tokens the model can emit at each step.
  • V(k)V^{(k)}“top k vocabulary”
    The kk most probable tokens at this step.
  • V(p)V^{(p)}“top p vocabulary, the nucleus”
    The smallest set of most probable tokens whose total probability is at least pp.
  • pbase pmax⁡p_{\text{base}}\,p_{\max}“p base times p max”
    The min-pp threshold: tokens less probable than this fraction of the top token are removed.
  • H(p)H(p)“entropy of p”
    Spread of the next-token distribution in bits, between 0 and log⁡2∣V∣\log_2|V|.
  • BB“beam width”
    Number of partial sequences beam search keeps at each step.

Logits in, one token out

At each step the model produces a logit ulu_l for each of the ∣V∣|V| tokens (the next-token prediction chamber built them), the softmax turns logits into probabilities, and a decoding strategy turns probabilities into one token. Append it, run the model again, and repeat until something says stop.

Greedy decoding takes arg⁡max⁡lul\arg\max_l u_l every time. It is cheap and, in principle, deterministic. But it maximises one step at a time, not the whole sequence: a token that wins now can lead to a context where every continuation is mediocre. In the third coding problem below, you will measure how far greedy falls behind the most probable sequence.

Beam search, and why it degenerates

Beam search keeps the BB highest-scoring partial sequences. At each step it extends every beam by every token, scores each candidate by its total log-probability, and keeps the best BB. Width B=1B=1 is greedy. It costs about BB times the compute and is still not exact.

Total log-probability favours short outputs, since each extra token adds a negative term, so practical beam search normalises by length and retires beams that emit the end-of-sequence token. Even so, more search is not always better: with an exact search, Stahlberg and Byrne (2019) found that a translation model's single best output was the empty translation for more than half of the test sentences.

Beam search suits tasks where the input pins the output down, such as translation, summarisation and speech recognition: there, the most probable answer is usually close to the right one. On open-ended text it degenerates. Holtzman et al. show that the probability of a repeated phrase rises with each repetition, a positive feedback loop that a likelihood-seeking search walks straight into. The text it finds is also bland, because human text is not the most probable text: its per-token probabilities swing between likely and surprising.

Quick check +20 XP

Why does beam search tend to produce repetitive, generic text from open-ended prompts?

Temperature

Temperature T>0T>0 divides every logit before the softmax. Let mm be the index of the largest logit and multiply above and below by e−um/Te^{-u_m/T}:

pi(T)=eui/T∑jeuj/T=e(ui−um)/T∑je(uj−um)/T.p_i(T)=\frac{e^{u_i/T}}{\sum_je^{u_j/T}}=\frac{e^{(u_i-u_m)/T}}{\sum_je^{(u_j-u_m)/T}}.

The second form is also how stable code computes it. Now take limits.

  • T→0+T\to0^+. If the maximum is unique, every term with j≠mj\ne m has an exponent that tends to −∞-\infty, so it vanishes, while the j=mj=m term is always e0=1e^0=1. Hence pm→1p_m\to1: a one-hot vector on the argmax, which is greedy decoding. Implementations special-case T=0T=0 as greedy rather than divide by zero.
  • T→∞T\to\infty. Every exponent tends to 0, every weight to 1, and pi→1/∣V∣p_i\to1/|V|: uniform.
  • Order is preserved. For any T>0T>0, pi/pj=e(ui−uj)/Tp_i/p_j=e^{(u_i-u_j)/T} exceeds 1 exactly when ui>uju_i>u_j. Temperature reshapes the distribution but never re-ranks it, so on its own it cannot remove the tail.

Entropy rises steadily in between. Writing β=1/T\beta=1/T, one can show dH/dβ=−βVar⁡p(u)dH/d\beta=-\beta\operatorname{Var}_{p}(u), so dH/dT=Var⁡p(u)/T3≥0dH/dT=\operatorname{Var}_{p}(u)/T^3\ge0 in nats: from 0 bits as T→0T\to0 up to log⁡2∣V∣\log_2|V| bits as T→∞T\to\infty. At T=1T=1 you sample the model's own distribution.

Quick check +20 XP

Logits are (2,1,0)(2,1,0). What is the probability of the first token at temperature T=0.5T=0.5? Give three decimal places.

Truncation: top-k, top-p and min-p

A vocabulary of 100,000 tokens has a long tail. Each tail token is unlikely, but together they hold real mass, and one bad draw can derail everything after it. Holtzman et al. call this the unreliable tail. Truncation samplers keep a set SS and renormalise:

P′(x)=P(x)∑y∈SP(y)  for x∈S,P′(x)=0  otherwise.P'(x)=\frac{P(x)}{\sum_{y\in S}P(y)}\ \text{ for }x\in S,\qquad P'(x)=0\ \text{ otherwise}.

That is exactly conditioning, P′(x)=P(x∣X∈S)P'(x)=P(x\mid X\in S), so the ratios between survivors are unchanged.

MethodKeepsAdapts to the distribution?Common values
Top-kk (Fan et al., 2018)the kk most likely tokensNo10–50
Top-pp or nucleus (Holtzman et al., 2020)the smallest set with mass at least ppYes, through its mass0.9–0.95
Min-pp (Nguyen et al., 2025)tokens with P(x)≥pbase pmax⁡P(x)\ge p_{\text{base}}\,p_{\max}Yes, through the top token0.05–0.1

Fan et al. generated stories by sampling from the k=10k=10 most likely words, and found it substantially more effective than beam search. A fixed kk is the wrong size most of the time, though: too many candidates when the model is sure, too few when hundreds of words fit. The nucleus grows and shrinks with the distribution; the proof below shows it is always a non-empty prefix of the sorted tokens. Min-pp scales its cut-off by the model's confidence instead. With probabilities 0.8,0.07,0.03,0.02,0.010.8, 0.07, 0.03, 0.02, 0.01, top-pp at 0.9 keeps three tokens, while min-pp at 0.1 sets the threshold at 0.08 and keeps one. Its authors argue this lets you raise the temperature for creativity while staying coherent.

Quick check +20 XP

Sorted next-token probabilities are 0.5,0.2,0.15,0.1,0.050.5, 0.2, 0.15, 0.1, 0.05. How many tokens does nucleus sampling with p=0.9p=0.9 keep?

The sampler, step by step

A typical sampler edits the logits, applies temperature, truncates and renormalises, then draws:

Python
def next_token(u, history, rng, T=0.8, top_k=50, top_p=0.95, min_p=0.0,
               bias=None, freq=0.0, pres=0.0):
    u = u.copy()
    for tok, b in (bias or {}).items():          # logit bias
        u[tok] += b
    c = np.bincount(history, minlength=len(u))   # counts so far
    u -= freq * c + pres * (c > 0)               # frequency and presence penalties
    if T == 0:
        return int(np.argmax(u))                 # greedy
    p = softmax(u / T)                           # temperature
    p = renormalise(p * top_k_mask(p, top_k))    # truncate, renormalise, repeat
    p = renormalise(p * top_p_mask(p, top_p))
    p = renormalise(p * (p >= min_p * p.max()))
    return int(rng.choice(len(p), p=p))

Around it, the loop stops at the end-of-sequence token, at a stop sequence (a string such as </answer> that ends generation when it appears and is usually left out of the output), or at max tokens, a hard cap that cuts the output off mid-sentence. Many APIs report which one fired as a finish reason: check it before you trust a truncated answer.

The order is common, not universal. At the time of writing, vLLM applies penalties, then temperature, then min-pp, then top-kk and top-pp, while llama.cpp's default chain truncates first, applies temperature last, and lets you reorder the chain with --samplers. The min-pp authors found that two widely used libraries applied temperature at different points, which changed their high-temperature results. Compare settings across tools only after checking the order.

Penalties fight repetition by editing the logits of tokens already generated. With cjc_j the count of token jj so far, frequency and presence penalties give uj−αfreqcj−αpres1[cj>0]u_j-\alpha_{\text{freq}}c_j-\alpha_{\text{pres}}\mathbb 1[c_j>0]: the first grows with each use, the second is a one-off charge for appearing at all. The repetition penalty from CTRL (Keskar et al., 2019) divides the logits of earlier tokens by θ≈1.2\theta\approx1.2; implementations multiply negative logits instead, so the penalty always pushes down. Penalties also hit legitimate repeats: a variable name, a JSON key, the subject of a story. Logit bias adds a fixed bjb_j to chosen tokens: a large negative bias bans a token, a positive one encourages it. It acts on tokens, not words, so check how your words tokenise.

Log-probabilities as a confidence signal

Many APIs and open-weight servers can return each output token's log-probability, often with the top alternatives. Summed, they give the log-probability of the whole output. For a one-token label they give a ready-made confidence score, and a small gap between the top two labels flags a borderline case. Treat them as a signal, not as truth. They measure the model's belief about the next token, not whether a fact is correct, and depending on the tool they are computed before or after your sampling settings reshape the distribution. Post-training also bends them: the GPT-4 technical report shows a pre-trained model whose answer probabilities were well calibrated on a multiple-choice benchmark, and a post-trained model that was markedly less so. Before you threshold on log-probabilities, bin outputs by confidence, compare with accuracy on labelled data, and recalibrate if needed.

Speculative decoding

Decoding is slow because every token needs a forward pass of a large model, one after another. Speculative decoding (Leviathan, Kalman and Matias, ICML 2023) lets a small draft model qq propose γ\gamma tokens, then runs the target model pp once over all of them in parallel. Each draft token xx is accepted with probability min⁡(1,p(x)/q(x))\min(1,p(x)/q(x)); at the first rejection, a replacement is drawn from the normalised max⁡(0,p−q)\max(0,p-q). This rejection scheme makes every token distributed exactly as a sample from pp, so quality is untouched and only the speed depends on how often the draft agrees. The paper reports 2–3× faster decoding of T5-XXL with identical outputs.

Settings by task

TaskTemperatureTruncationNotes
Extraction, classification0none neededPair with structured output (two chambers on); use log-probabilities for confidence
Code0–0.2 for one answer, about 0.8 when sampling many to testtop-pp 0.95In the Codex paper, a 679M-parameter model did best at T=0.2T=0.2 for pass@1 and T=0.8T=0.8 for pass@100
Chat, assistants0.5–0.8top-pp 0.9–0.95 or min-pp 0.05Provider defaults are tuned for this
Creative writing0.8–1.2min-pp 0.05–0.1A light frequency or presence penalty against loops
Brainstorming1.0–1.5min-pp 0.1Generate many, deduplicate, then filter

These are illustrative starting points, not laws: tune them on your own evals. Two cautions. Some reasoning models fix their sampling settings or ignore them: depending on the provider, a custom temperature or top-pp is rejected, silently ignored, or restricted while extended reasoning is on, and open-weight reasoning models usually publish recommended settings in their model cards. Read the documentation for the exact model you call. And temperature 0 is not a reproducibility guarantee: why identical requests can still differ, and what to do about it, is the next chamber.

DiscoverLearnRead beyondPapers & lecturesYour turn

Read beyond

Article · free online · ~25 min

How to generate text: using different decoding methods for language generation with Transformers

Patrick von Platen (Hugging Face) · Greedy search, beam search, top-k and top-p sampling

Run greedy and beam search on one prompt, then mark every repeated phrase in each output.

Paper · free online · ~15 min

Hierarchical Neural Story Generation

Fan, Lewis & Dauphin (ACL 2018) · Section 5.4: generation

List the reasons the authors give for sampling from the top 10 words instead of using beam search.

Paper · free online · ~20 min

Turning Up the Heat: Min-p Sampling for Creative and Coherent LLM Outputs

Nguyen, Baker, Neo, Roush, Kirsch & Shwartz-Ziv (ICLR 2025) · Section 3: min-p sampling

Pick a distribution of your own and find a temperature at which top-p 0.9 keeps junk that min-p 0.1 removes.

Paper · free online · ~20 min

Fast Inference from Transformers via Speculative Decoding

Leviathan, Kalman & Matias (ICML 2023) · Section 2.3 and the appendix: speculative sampling and its correctness

Check on a two-token example that accepting with probability min(1, p/q) and resampling from max(0, p − q) gives back p.

DiscoverLearnRead beyondPapers & lecturesYour turn

Read the equation in context

The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes & Yejin Choi · ICLR, 2020

Section 3 sets out every decoder the paper compares. Eq. (2) defines the top-pp vocabulary as the smallest set reaching mass pp. Eq. (3) divides by p′=∑x∈V(p)P(x∣x1:i−1)p'=\sum_{x\in V^{(p)}}P(x\mid x_{1:i-1}) inside it, the conditioning derived above, and Section 3.2 reuses Eq. (3) for top-kk. Section 3.3 writes temperature as a rescaled softmax, Eq. (4), noting that t∈[0,1)t\in[0,1) lowers the mass in the tail at a cost in diversity. Decode Eq. (4), then read Figure 2 (per-token probabilities of beam search and of human text) and Figure 4 (the repetition loop) in this chamber's terms.

Decode the paper · Eq. (4), Section 3.3: sampling with temperature

The Curious Case of Neural Text Degeneration

Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes & Yejin Choi · ICLR, 2020

+25 XP
p(x=Vl∣x1:i−1)=exp⁡(ul/t)∑l′exp⁡(ul′/t)p(x=V_l\mid x_{1:i-1})=\frac{\exp(u_l/t)}{\sum_{l'}\exp(u_{l'}/t)}

Section 3 sets out the decoders the paper compares. Eq. (2) defines the nucleus, Eq. (3) renormalises inside it, and Section 3.2 reuses Eq. (3) for top-kk. Section 3.3 then writes temperature as a rescaled softmax and notes that t∈[0,1)t\in[0,1) skews the distribution towards likely tokens and lowers the mass in the tail, at a cost in diversity.

ulu_l
tt
VlV_l
x1:i−1x_{1:i-1}
∑l′\sum_{l'}

Options

Deep Dive into LLMs like ChatGPTAndrej Karpathy · 211 min · starts at the inference chapter (26:01), where the model samples one token at a time
DiscoverLearnRead beyondPapers & lecturesYour turn

Your turn

Ten candidate tokens, fixed logits, four knobs. Shape the distribution three ways, sample 20 continuations to hear what each setting sounds like, and see how far you can push the temperature.

Interactive lab

Nucleus navigator

Ten candidate next tokens with fixed logits. The sampler divides the logits by T, then applies top-k, top-p and min-p in that order, renormalising after each cut. T = 0 means greedy. Hit the three targets below.

The lighthouse keeper opened the door and saw a ▁

ship39.5%
storm21.7%
stranger16.1%
light9.7%
whale5.3%
ghost4.0%
letter2.7%
the0.6%
of0.3%
and0.2%

Solid bars: the final distribution. Dashed outlines: after temperature, before truncation. Rose tokens are junk.

Survivors

10 / 10

Top token

39.5%

Entropy

2.39 bits

Mass kept

100.0%

Entropy can reach at most log₂ 10 ≈ 3.32 bits over these tokens.

  • Three-way race. Exactly three tokens survive, and the most likely one has between 40% and 45%.
  • Wide but clean. All three junk tokens are removed, yet the entropy is at least 2.6 bits.
  • Greedy at temperature 1. At T = 1, with top-k and min-p off, only one token survives: top-p alone makes decoding greedy.

Challenge: Nucleus navigatorTune temperature, top-k and top-p to hit three target distributions.+40 XP

Match · Method ↔ What it does

Decoding methods

+20 XP
Greedy decoding
Beam search
Top-kk sampling
Nucleus (top-pp) sampling
Min-pp sampling

Options

Match · Expression ↔ Meaning

Knobs on the logits

+20 XP
lim⁡T→0+p(T)\lim_{T\to0^+}p(T)
lim⁡T→∞p(T)\lim_{T\to\infty}p(T)
uj−αfreqcj−αpres1[cj>0]u_j-\alpha_{\text{freq}}c_j-\alpha_{\text{pres}}\mathbb 1[c_j>0]
uj+bju_j+b_j
min⁡(1,p(x)/q(x))\min\bigl(1,p(x)/q(x)\bigr)

Options

Proof puzzle

Zero temperature is greedy

+25 XP

Claim

If logit umu_m is strictly larger than every other logit, prove that pi(T)=eui/T/∑jeuj/Tp_i(T)=e^{u_i/T}/\sum_je^{u_j/T} tends to the one-hot vector on mm as T→0+T\to0^+.

Tap lines in the order they should appear. Not every line belongs. Tap a line in your proof to send it back.

Your proof

  1. Pick the first line below.

Available lines

Prove it yourself

The nucleus is a prefix

+35 XP

Claim

Let p(1)≥p(2)≥⋯≥p(V)p_{(1)}\ge p_{(2)}\ge\dots\ge p_{(V)} be a next-token distribution sorted in decreasing order, and let 0<p≤10<p\le1. Prove that a smallest set of tokens with total probability at least pp exists, is non-empty, and can be taken to be a prefix {(1),…,(k∗)}\{(1),\dots,(k^*)\} of the sorted list.

Preview

Your typeset proof appears here.

Coding problems

Problem 10·Warm-up

Entropy at a temperature

+20 XP

A model's logits for eight candidate tokens are u=(4,3,3,2,1,0,−1,−2)u=(4,3,3,2,1,0,-1,-2). Apply temperature T=0.7T=0.7 and compute the entropy of the resulting distribution in bits, to 4 decimal places.

A number, rounded to 4 decimal places

Problem 11·Standard

How big is the nucleus?

+35 XP

Generate logits with the Lehmer generator xn+1=48271 xn mod (231−1)x_{n+1}=48271\,x_n \bmod (2^{31}-1), starting from x0=2026x_0=2026. Each draw xnx_n with n≥1n\ge1 gives the logit (xn mod 1000)/100(x_n \bmod 1000)/100. Build 200 decoding steps, each from the next 50 draws in order (the logits of tokens 0 to 49). At every step apply temperature T=0.8T=0.8, then find the nucleus for p=0.9p=0.9: the smallest set of tokens, taken in decreasing order of probability, whose probabilities sum to at least 0.9. Report the sum of the 200 nucleus sizes.

An exact integer (or a fraction like 7/12)

Problem 12·Challenge

Greedy is not the mode

+50 XP

A bigram language model over tokens 0 to 4 has next-token logits UabU_{ab} for current token aa (row) and next token bb (column), turned into probabilities by a softmax over each row: U=(5344202523224210214105241).U=\begin{pmatrix}5&3&4&4&2\\0&2&5&2&3\\2&2&4&2&1\\0&2&1&4&1\\0&5&2&4&1\end{pmatrix}. Starting from token 0, generate exactly 12 more tokens. Greedy decoding takes the most probable next token at each step. Compute log⁡P(most probable 12-token continuation)−log⁡P(greedy continuation)\log P(\text{most probable 12-token continuation})-\log P(\text{greedy continuation}) in nats, to 6 decimal places. There are 512≈2.4×1085^{12}\approx2.4\times10^8 continuations, so search cleverly.

A number, rounded to 6 decimal places

Key takeaways

  • Greedy and beam search maximise probability, and on open-ended text the most probable continuations are bland and repetitive.
  • Temperature divides the logits: T→0T\to0 is greedy, T→∞T\to\infty is uniform, and the ranking never changes.
  • Top-kk, top-pp and min-pp cut the unreliable tail and renormalise, which is conditioning on the kept set.
  • Samplers differ in their order of operations, penalties and log-probabilities: check before comparing settings across tools.
  • Choose settings by task, near 0 for extraction and higher with truncation for creative work, then confirm them with evals.

Checkpoint

Prove it to the labyrinth

Answer every question to clear this chamber. First-try answers earn the most XP.

0/8
Question 1 of 8 +20 XP

With min-pp and pbase=0.1p_{\text{base}}=0.1, how many tokens survive from the probabilities 0.6,0.2,0.1,0.05,0.03,0.020.6, 0.2, 0.1, 0.05, 0.03, 0.02?

Question 2 of 8 +20 XP

As T→∞T\to\infty, what does the temperature-scaled softmax tend to?

Question 3 of 8 +20 XP

A distribution is (0.5,0.3,0.2)(0.5, 0.3, 0.2). After top-kk truncation with k=2k=2 and renormalising, what is the first token's probability?

Question 4 of 8 +20 XP

You extract invoice fields into JSON and want the most likely answer on every call. Which setting is the best starting point?

Question 5 of 8 +20 XP

What distinguishes a presence penalty from a frequency penalty?

Question 6 of 8 +20 XP

In speculative decoding, a small draft model proposes tokens and a large target model verifies them. What is the distribution of the output?

Question 7 of 8 +20 XP

A vocabulary has 1024 tokens. With no truncation, what entropy in bits does the next-token distribution approach as T→∞T\to\infty?

Question 8 of 8 +20 XP

A post-trained chat model gives the token of its answer probability 0.97. What can you safely conclude?

End of the chamber

Clear this chamber

+60 XPTemperatureTop-k SamplingNucleus SamplingBeam Search