Skip to content
AriadneTechnology

The Middle Ring · Chamber 5 of 9

Non-Determinism: Taming and Using Randomness

Why temperature zero still varies, how to make runs reproducible, how to get variety on purpose, and how to test a system that never answers the same way twice.

60 min 60 XP + 11 questions + 1 challengeMathVideoPapersProofsCodeLab

In this chamber you will

  • Name every source of variation, from the sampler's seed to batch-dependent GPU kernels
  • Show that floating-point addition is not associative, and why that can flip greedy decoding
  • Make LLM calls reproducible by pinning, seeding, logging, caching and replaying them
  • Generate varied outputs on purpose, and turn many samples into a better answer with self-consistency and pass@k
DiscoverLearnRead beyondPapers & lecturesYour turn

Send a served model the same prompt twice at temperature zero and you can get two different answers. In 2025, Horace He and colleagues at Thinking Machines Lab asked Qwen3-235B for 1,000 temperature-zero completions of "Tell me about Richard Feynman" and got 80 different texts. All of them matched for 102 tokens. At token 103, after "Feynman was born on May 11, 1918, in", 992 went on with "Queens, New York" and 8 with "New York City". Nobody rolled a die. This chamber finds where that variation comes from, how to remove it when you need a run you can repeat, and how to put randomness to work when you want it. The most useful trick of the second half fits in one line: sample many answers and keep the most common.

Spotted in the wild

arg⁡max⁡a∑i=1m1(ai=a)\arg\max_a\sum_{i=1}^{m}\mathbb{1}(a_i=a)
Self-Consistency Improves Chain of Thought Reasoning in Language Models
DiscoverLearnRead beyondPapers & lecturesYour turn
Symbols for this chamber
  • fl(x)\mathrm{fl}(x)“f l of x”
    The floating-point number nearest to the real number xx: round to nearest, ties to even.
    fl(108+1)=108 in float32\mathrm{fl}(10^8+1)=10^8\ \text{in float32}
  • a⊕ba\oplus b“a circle-plus b”
    Computed floating-point addition, fl(a+b)\mathrm{fl}(a+b). It is commutative but not associative.
    (1⊕108)⊕(−108)=0(1\oplus10^8)\oplus(-10^8)=0
  • uu“unit roundoff”
    The largest relative error of a single rounding: 2−242^{-24} for float32.
    fl(x)=x(1+δ), ∣δ∣≤u\mathrm{fl}(x)=x(1+\delta),\ |\delta|\le u
  • 1(ai=a)\mathbb{1}(a_i=a)“indicator that a i equals a”
    1 if the iith sample's final answer is aa, and 0 otherwise.
  • mm“m”
    The number of reasoning paths sampled for self-consistency (40 in Wang et al.).
    a1,…,ama_1,\dots,a_m
  • pass@k\text{pass@}k“pass at k”
    The probability that at least one of kk samples passes the tests.
    pass@k=1−(1−p)k\text{pass@}k=1-(1-p)^k
  • (nk)\binom{n}{k}“n choose k”
    The number of ways to choose kk of nn samples.
    (52)=10\binom{5}{2}=10
  • p^\hat p“p hat”
    An observed pass rate: passes divided by trials.
    p^=x/n\hat p=x/n

Seven sources of variation

SourceWhat changes between runsRemedy
Sampler randomnessThe random numbers drawn at temperature above zeroA fixed seed, where supported
Floating-point roundingThe order in which a long sum is added upA fixed reduction order
Batch-dependent kernelsThe reduction strategy, chosen from a batch size that tracks server loadBatch-invariant kernels
Mixture-of-experts routingWhich experts a token visitsStable numerics; routing that ignores other requests
Near-tiesWhich of two almost equal logits wins the argmaxNone: detect them and tolerate them
Model aliasesThe weights behind a name such as "latest"Pin a dated snapshot
Upstream inputsRetrieved documents, tool results, prompt edits, the date in the system promptLog the full request and replay it

The first row is the only one temperature controls. At temperature above zero a sampler can draw r∼U(0,1)r\sim\mathrm U(0,1) and return the first token whose cumulative probability exceeds rr, so a pseudo-random generator with a fixed seed reproduces the draws. Engines that batch many users must give each request its own reproducible stream rather than one shared generator: SGLang's deterministic mode, for example, draws its sampling noise from a seeded hash function, so that "the same (inputs, seed) pair always yields the same sample". Hosted APIs that accept a seed describe it as best effort, some report a fingerprint of the serving configuration so you can see when the backend changed, and some offer no seed at all. And a seed fixes only rr: if the arithmetic upstream moves a cumulative probability across it, the token changes anyway.

Floating-point addition is not associative

A float32 number carries a 24-bit significand, so each rounding has relative error at most u=2−24u=2^{-24}. Between 2262^{26} and 2272^{27} consecutive float32 numbers are 23=82^{3}=8 apart. Write a⊕b=fl(a+b)a\oplus b=\mathrm{fl}(a+b) for computed addition.

Claim. In float32, (1⊕108)⊕(−108)=0(1\oplus10^8)\oplus(-10^8)=0 but 1⊕(108⊕(−108))=11\oplus(10^8\oplus(-10^8))=1.

Proof. 108=8×12,500,00010^8=8\times12{,}500{,}000 lies between 226≈6.7×1072^{26}\approx6.7\times10^7 and 227≈1.3×1082^{27}\approx1.3\times10^8, so it is representable. The exact sum 108+110^8+1 is 1 from 10810^8 and 7 from the next float, 108+810^8+8, so it rounds to 10810^8, and subtracting 10810^8 leaves exactly 0. On the right, 108⊕(−108)=010^8\oplus(-10^8)=0 exactly, and 1⊕0=11\oplus0=1. The two groupings differ, so ⊕\oplus is not associative.

A dot product of length KK is a long sum. Added left to right, the computed result obeys ∣s^−s∣≲(K−1) u∑i∣xi∣|\hat s-s|\lesssim(K-1)\,u\sum_i|x_i|; a tree of pairwise sums improves the factor to about ulog⁡2Ku\log_2K. Every order gives a slightly different answer within these bounds, and usually nobody cares. Greedy decoding cares when the top two logits are closer than the rounding error: then the order of additions picks the token, and since every later token is conditioned on it, the texts part ways from there.

Quick check +20 XP

In float32, evaluate (1⊕108)⊕(−108)(1\oplus10^8)\oplus(-10^8), adding left to right.

Why temperature zero still varies

The folk explanation runs: GPUs add in parallel, threads finish in unpredictable orders, floating-point addition is not associative, so results wobble. He and Thinking Machines Lab call this incomplete. Run the same matrix multiplication on the same data and you get bitwise identical results, run after run. An order truly decided by racing threads needs atomic adds, and a typical LLM forward pass usually contains none. Each forward pass is run-to-run deterministic.

What varies is the batch. A server packs concurrent requests into one forward pass, and a fast kernel chooses its strategy from the shapes it receives. With few rows a matrix multiplication has too few output tiles to keep every core busy, so it splits the long reduction across cores and adds the partial sums afterwards (split-K); with many rows it doesn't. A different split is a different order of additions, so the numbers computed for your row depend on how many strangers arrived with it. In the post's words, "from the perspective of an individual user, the other concurrent users are not an 'input' to the system but rather a nondeterministic property". CPU and TPU servers have the same problem.

Batch invariance. A kernel ff is batch-invariant if the row it computes for input xx is bit-for-bit the same in every batch that contains xx. The post states the requirement plainly: "the reduction order for each element must be fixed regardless of the batch-size of the kernel."

Three operations in a transformer reduce over many numbers: RMSNorm, matrix multiplication and attention. The post makes each one batch-invariant. Normalisation and matrix multiplication use one kernel configuration for every shape, giving up split-K. Attention must also add a token's keys in the same order whether they arrive in one prefill or one at a time from the KV cache (as with chunked prefill and prefix caching), so the cache is laid out consistently before the kernel runs, and where the keys are split across cores, the size of each split is fixed rather than the number of splits. With these kernels all 1,000 Feynman completions were identical. The price was speed: on one GPU running Qwen3-8B, 1,000 sequences took 26 s with vLLM's default kernels, 55 s with an unoptimised deterministic version and 42 s with an improved attention kernel.

Mixture-of-experts models add a discrete choice: a router sends each token to its top-scoring experts, so a rounding-level change in router scores can redirect a token, and where each expert has a capacity limit, other requests' tokens compete for its slots. The Feynman model was itself a mixture of experts, and batch-invariant kernels were enough to make it repeat itself.

Quick check +20 XP

According to He and Thinking Machines Lab (2025), what mainly makes a served LLM give different answers to identical temperature-zero requests?

Making runs reproducible

Temperature zero is not a guarantee: it removes the sampler's dice, not the arithmetic or the inputs. To repeat a result, whether to debug, to audit or to compare two prompts fairly, control what you can and record the rest:

  1. Pin the model. Request a dated snapshot, never an alias the provider can move to new weights.
  2. Pin every parameter. Temperature, top-p, maximum tokens, stop sequences, system prompt, tool schemas and response format. Defaults change too.
  3. Seed where supported, knowing that hosted APIs promise only best effort.
  4. Log everything. The full request and response, the model version string, any system fingerprint, timestamps, and the retrieved context and tool results that went into the prompt.
  5. Cache by a hash of the request, so identical requests return the stored response.
  6. Record and replay in tests. Unit tests replay recorded responses; a separate, scheduled suite calls the live model.
  7. Self-host for bitwise reproducibility. Some open-source inference engines, vLLM and SGLang among them, now offer a batch-invariant or deterministic mode, at a cost in throughput and with limits on supported hardware and models. Check the current documentation before you rely on it.
Python
import hashlib, json

def request_key(request: dict) -> str:
    # Sorted keys and fixed separators make the JSON canonical: same request, same key.
    canonical = json.dumps(request, sort_keys=True, separators=(",", ":"), ensure_ascii=False)
    return hashlib.sha256(canonical.encode("utf-8")).hexdigest()

def cached_call(call, request: dict, store: dict) -> dict:
    key = request_key(request)        # model snapshot, every parameter, full messages, tools
    if key not in store:
        store[key] = call(**request)  # record on a miss
    return store[key]                 # replay on a hit

Variety on purpose

Sometimes the same answer twice is the failure. Brainstorming, creative writing, test-case generation and synthetic training data all want spread.

LeverHowWatch out for
Temperature, top-pRaise them (Chamber 4)Coherence falls in the tail
nn samplesSeveral completions of one promptThey cluster around the prompt's modes
SeedsA new, recorded seed per callVaries only the draws
Prompt variationPersonas, constraints, shuffled example orderRecord which variant produced what
Ask for a list"Ten options that differ in setting and tone" in one callLater items may weaken
DeduplicateEmbed, then drop pairs above a cosine thresholdThe threshold is task-specific

Measure what you get. Distinct-n is the number of unique nn-grams across all samples divided by the total number of nn-grams (some papers divide by the total number of tokens instead); mean pairwise cosine distance between embeddings catches paraphrases that nn-grams miss. For synthetic data, check for near-duplicates before you train on it.

Many samples, one better answer

When the task has a checkable final answer, variety buys accuracy. Self-consistency (Wang et al., 2023) samples mm chain-of-thought reasoning paths, parses each final answer aia_i, and returns the most frequent one: the opening equation. Its intuition is that correct reasoning paths, however different, tend to agree on the final answer more often than incorrect ones. With PaLM-540B and 40 sampled paths at temperature 0.7, it raised GSM8K accuracy from 56.5% with greedy chain of thought to 74.4%.

A clean model explains when voting helps. If each sample is right with probability pp, independently, and answers are binary, the vote over an odd number nn of samples is right with probability

Pvote(n)=∑k>n/2(nk)pk(1−p)n−k,P_{\text{vote}}(n)=\sum_{k>n/2}\binom nk p^k(1-p)^{n-k},

which rises towards 1 when p>12p>\tfrac12 and falls towards 0 when p<12p<\tfrac12: voting amplifies whatever the model usually says. All nn samples agree with probability pn+(1−p)np^n+(1-p)^n, so low agreement flags a question worth routing to a stronger model or a person. Real samples share blind spots, so gains flatten sooner than the formula suggests; when wrong answers scatter over many values, plurality voting can help even with p<12p<\tfrac12 (Problem 3).

Quick check +20 XP

Each sample's binary answer is correct with probability 0.70.7, independently. What is the accuracy of a majority vote over 3 samples?

Best-of-n with a verifier samples nn answers and keeps the one a verifier scores highest: unit tests, a schema check, a calculator or a reward model. It works where voting can't, as with code and prose. pass@k asks whether at least one of kk samples passes. Chen et al. (2021) estimate it from n≥kn\ge k samples of which cc pass, as 1−(n−ck)/(nk)1-\binom{n-c}{k}\big/\binom nk; the plug-in 1−(1−c/n)k1-(1-c/n)^k is biased low, and the proof puzzle below shows why the first is unbiased.

Testing a system that never answers the same way twice

A test that compares the reply with a stored string fails on the first harmless rephrasing. Assert properties instead: the reply parses and validates against its schema; required fields or facts are present and forbidden ones absent; numbers, lengths and labels fall in their allowed ranges; for free text, similarity to a reference or a model grader's verdict clears a threshold.

Python
import json

def acceptable(reply: str) -> bool:
    try:
        data = json.loads(reply)
    except json.JSONDecodeError:
        return False
    return data.get("label") in {"refund", "exchange", "other"} and 0 <= data.get("confidence", -1) <= 1

def trials(generate, prompt: str, n: int = 30) -> tuple[int, int]:
    return sum(acceptable(generate(prompt)) for _ in range(n)), n   # passes, trials

Then treat the pass rate as an estimate. With xx passes in nn trials and p^=x/n\hat p=x/n, the Wilson interval

p^+z22n1+z2n±z1+z2np^(1−p^)n+z24n2\frac{\hat p+\frac{z^2}{2n}}{1+\frac{z^2}{n}}\pm\frac{z}{1+\frac{z^2}{n}}\sqrt{\frac{\hat p(1-\hat p)}{n}+\frac{z^2}{4n^2}}

stays honest near 0 and 1, where the textbook p^±zp^(1−p^)/n\hat p\pm z\sqrt{\hat p(1-\hat p)/n} shrinks to a point. Eighteen passes in twenty give p^=0.9\hat p=0.9 but a 95% interval of about [0.70,0.97][0.70,0.97]. Set thresholds before you look: ship only if the lower end clears, say, 0.9; call a test flaky when it both passes and fails at fixed inputs; and quarantine flaky tests instead of rerunning until green, which measures only your patience. Chamber 7 builds this into full eval suites, with graders and paired comparisons between prompt versions.

DiscoverLearnRead beyondPapers & lecturesYour turn

Read beyond

Article · free online · ~35 min

Defeating Nondeterminism in LLM Inference

Horace He and Thinking Machines Lab (2025) · From “The original sin: floating-point non-associativity” to “Experiments”

Find the sentence that rejects "concurrency plus floating point" as the whole story, and list the three operations they make batch-invariant.

Paper · free online · ~45 min

What Every Computer Scientist Should Know About Floating-Point Arithmetic

David Goldberg (ACM Computing Surveys, 1991) · Rounding Error: relative error, ulps and cancellation

Redo this chamber's 10810^8 example in units in the last place.

Paper · free online · ~30 min

Non-Determinism of “Deterministic” LLM Settings

Berk Atil et al. (arXiv, 2024) · Section 5.2 (metrics) and Section 8.1 (implications for practical engineering)

Note how far accuracy moved across ten identical runs, and how TARr@N differs from TARa@N.

Article · free online · ~20 min

Towards Deterministic Inference in SGLang and Reproducible RL Training

The SGLang Team (LMSYS, 2025)

Find how they keep sampling reproducible at temperatures above zero, and what deterministic mode costs.

DiscoverLearnRead beyondPapers & lecturesYour turn

Read the equation in context

Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery & Denny Zhou · ICLR, 2023

Section 2 defines the vote. Each sampled output is a pair (ri,ai)(r_i,a_i): a reasoning path rir_i and the final answer aia_i parsed from it, and voting over the aia_i marginalises out the paths. The paper also weights each vote by the model's probability of the whole output: normalised by length, that scored about the same as the plain vote on PaLM-540B, while the unnormalised version did clearly worse (Table 1).

Evaluating Large Language Models Trained on CodeMark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan et al. · arXiv, 2021

The Codex paper introduced HumanEval and made "try kk times" a metric. It draws n=200n=200 samples per problem, counts the cc that pass the unit tests, and estimates pass@k for every k≤100k\le100 from the same samples.

Decode the paper · Section 2.1, Equation (1)

Evaluating Large Language Models Trained on Code

Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan et al. · arXiv, 2021

+25 XP
pass@k:=EProblems[1−(n−ck)(nk)]\text{pass@}k:=\mathop{\mathbb E}_{\text{Problems}}\left[1-\frac{\binom{n-c}{k}}{\binom{n}{k}}\right]

Chen et al. judge generated code by running it: a problem counts as solved if any of kk samples passes its unit tests. Drawing exactly kk samples per problem gives a noisy estimate, so they draw n≥kn\ge k, count the cc that pass, and compute the probability that a random kk-subset contains at least one pass. Appendix A shows this is unbiased, while plugging the empirical pass rate into 1−(1−p^)k1-(1-\hat p)^k underestimates.

nn
cc
kk
(n−ck)/(nk)\binom{n-c}{k}\big/\binom{n}{k}
EProblems\mathbb E_{\text{Problems}}

Options

Floating Point Numbers - ComputerphileComputerphile · 9 min
DiscoverLearnRead beyondPapers & lecturesYour turn

Your turn

The lab builds the Feynman fork in miniature: two candidate tokens whose float32 logits differ by a tenth of a float32 step in exact arithmetic. Find a batch size that flips greedy decoding, then switch on the batch-invariant kernel and show the flip is gone. The second panel turns the voting formulas into curves.

Interactive lab

Batch flip

Two candidate tokens, two logits, each a 2048-term float32 dot product; in exact arithmetic they differ by about 1e-7. To keep its 16 cores busy, the default kernel splits each sum into chunks when the batch is small (split-K), so the order of additions depends on how many requests share the batch.
Feynman was born on May 11, 1918, in Queens, New York.

16 chunks of 128 terms, each summed in order, then the 16 partial sums added in order.

logit “Queens”

11.1170597

logit “New”

11.1170578

Gap, float32 steps

+2

Cores busy

16/16

BatchChunks“Queens”“New”Greedy
1–81611.117059711.1170578Queens
9–16811.117061611.1170568Queens
17–32411.117056811.1170588New
33–64211.117065411.1170635Queens
65–128111.117068311.1170607Queens
float6411.117059653311.1170595527Queens

Many samples, one answer

Unscored. Each sample is right with probability p, independently, and answers are binary; an even split is settled by a coin. Real samples share blind spots, so real gains are usually smaller.

One sample

0.650

Vote of 5

0.765

All 5 agree

0.121

Majority-vote accuracy and probability of unanimous samples against the number of samples10110.25210.5310.75411Samples nProbability
● Majority vote is right● One sample is right● All n agree● Current n
Challenge: Batch flipFlip a temperature-zero answer by changing nothing but the batch size, then make the answer batch-invariant.+40 XP

Match · Symptom ↔ Fix

Symptoms and fixes

+20 XP
Different text at temperature 1 for the same request
Different text at temperature 0, depending on server load
Behaviour changed overnight under the same model name
A replayed request differs because the search results moved

Options

Match · Expression ↔ Meaning

Formulas for many samples

+20 XP
pn+(1−p)np^n+(1-p)^n
1−(n−ck)/(nk)1-\binom{n-c}{k}\big/\binom nk
1−(1−c/n)k1-(1-c/n)^k
arg⁡max⁡a∑i=1m1(ai=a)\arg\max_a\sum_{i=1}^m\mathbb{1}(a_i=a)
∑k>n/2(nk)pk(1−p)n−k\sum_{k>n/2}\binom nk p^k(1-p)^{n-k}

Options

Proof puzzle

The pass@k estimator is unbiased

+25 XP

Claim

Draw n≥kn\ge k independent samples, each passing with probability pp, and let cc count the passes. Show that E[1−(n−ck)/(nk)]=1−(1−p)k\mathbb E\big[1-\binom{n-c}{k}/\binom nk\big]=1-(1-p)^k.

Tap lines in the order they should appear. Not every line belongs. Tap a line in your proof to send it back.

Your proof

  1. Pick the first line below.

Available lines

Prove it yourself

Two out of three beats one

+35 XP

Claim

Three independent samples each give the correct answer with probability pp and otherwise the same wrong answer. Prove that their majority vote is correct with probability 3p2−2p33p^2-2p^3, and that this beats a single sample exactly when 12<p<1\tfrac12<p<1.

Preview

Your typeset proof appears here.

Coding problems

Problem 13·Warm-up

A perfect record, how sure?

+20 XP

A regression test of an LLM feature is run nn times at fixed inputs and passes every time. Using the 95% Wilson score interval with z=1.96z=1.96, find the smallest nn for which the lower end of the interval for the pass rate is at least 0.950.95.

An exact integer (or a fraction like 7/12)

Problem 14·Standard

Same numbers, different sum

+35 XP

Let xkx_k be 1/k1/k rounded to float32, for k=1,…,106k=1,\dots,10^6. Add them in float32, rounding after every addition. FF adds them in increasing order of kk (largest term first) and BB in decreasing order (smallest term first). Give B−FB-F to 6 decimal places. Careful: np.sum adds pairwise, which is a third order.

A number, rounded to 6 decimal places

Problem 15·Challenge

Votes when mistakes scatter

+50 XP

A model's final answer to a question is correct with probability 0.40.4. Otherwise it gives one of three wrong answers, with probabilities 0.30.3, 0.20.2 and 0.10.1. You draw 15 independent samples and return the most frequent answer, breaking any tie uniformly at random among the tied answers. What is the probability that you return the correct answer? Give 6 decimal places.

A number, rounded to 6 decimal places

Key takeaways

  • Temperature zero removes the sampler's randomness, not the arithmetic: floating-point addition is not associative, and at a near-tie the order of additions picks the token.
  • In served models the main culprit is kernels that are not batch-invariant, so a request's numerics depend on how many others share its batch. Batch-invariant kernels fix this at a cost in speed.
  • Reproducibility is a practice: pin the snapshot and every parameter, seed where you can, log full requests and responses, cache by hash and replay in tests.
  • For variety, vary temperature, seeds and prompts and deduplicate; for accuracy, turn many samples into one answer with self-consistency, a verifier or pass@k.
  • Test properties over repeated trials and report a pass rate with a Wilson interval, never a single run.

Checkpoint

Prove it to the labyrinth

Answer every question to clear this chamber. First-try answers earn the most XP.

0/8
Question 1 of 8 +20 XP

Which practice protects you against a provider pointing a model name at new weights?

Question 2 of 8 +20 XP

Binary answers; each sample is correct with probability 0.80.8, independently. What is the probability that all 5 samples agree?

Question 3 of 8 +20 XP

You draw n=20n=20 samples for a coding problem and c=5c=5 pass the tests. What is the unbiased estimate of pass@4?

Question 4 of 8 +20 XP

A test asserts that the model's reply equals a stored string. What should replace it?

Question 5 of 8 +20 XP

Which statement about batch-invariant kernels is true?

Question 6 of 8 +20 XP

A test passes 9 of 10 trials. What is the lower end of the 95% Wilson interval for its pass rate, with z=1.96z=1.96?

Question 7 of 8 +20 XP

Binary answers; each sample is correct with probability 0.40.4, independently. What happens to majority-vote accuracy as the number of samples grows?

Question 8 of 8 +20 XP

You need 200 varied product names as synthetic data. Which plan gives the most varied set?

End of the chamber

Clear this chamber

+60 XPSeeded SamplingFloating-Point Non-AssociativityBatch InvarianceSelf-Consistency