Send a served model the same prompt twice at temperature zero and you can get two different answers. In 2025, Horace He and colleagues at Thinking Machines Lab asked Qwen3-235B for 1,000 temperature-zero completions of "Tell me about Richard Feynman" and got 80 different texts. All of them matched for 102 tokens. At token 103, after "Feynman was born on May 11, 1918, in", 992 went on with "Queens, New York" and 8 with "New York City". Nobody rolled a die. This chamber finds where that variation comes from, how to remove it when you need a run you can repeat, and how to put randomness to work when you want it. The most useful trick of the second half fits in one line: sample many answers and keep the most common.
Spotted in the wild
- “f l of x”The floating-point number nearest to the real number : round to nearest, ties to even.
- “a circle-plus b”Computed floating-point addition, . It is commutative but not associative.
- “unit roundoff”The largest relative error of a single rounding: for float32.
- “indicator that a i equals a”1 if the th sample's final answer is , and 0 otherwise.
- “m”The number of reasoning paths sampled for self-consistency (40 in Wang et al.).
- “pass at k”The probability that at least one of samples passes the tests.
- “n choose k”The number of ways to choose of samples.
- “p hat”An observed pass rate: passes divided by trials.
| Symbol | Say it | Meaning | LaTeX |
|---|---|---|---|
| “f l of x” | The floating-point number nearest to the real number : round to nearest, ties to even. | ||
| “a circle-plus b” | Computed floating-point addition, . It is commutative but not associative. | ||
| “unit roundoff” | The largest relative error of a single rounding: for float32. | ||
| “indicator that a i equals a” | 1 if the th sample's final answer is , and 0 otherwise. | ||
| “m” | The number of reasoning paths sampled for self-consistency (40 in Wang et al.). | ||
| “pass at k” | The probability that at least one of samples passes the tests. | ||
| “n choose k” | The number of ways to choose of samples. | ||
| “p hat” | An observed pass rate: passes divided by trials. |
Seven sources of variation
| Source | What changes between runs | Remedy |
|---|---|---|
| Sampler randomness | The random numbers drawn at temperature above zero | A fixed seed, where supported |
| Floating-point rounding | The order in which a long sum is added up | A fixed reduction order |
| Batch-dependent kernels | The reduction strategy, chosen from a batch size that tracks server load | Batch-invariant kernels |
| Mixture-of-experts routing | Which experts a token visits | Stable numerics; routing that ignores other requests |
| Near-ties | Which of two almost equal logits wins the argmax | None: detect them and tolerate them |
| Model aliases | The weights behind a name such as "latest" | Pin a dated snapshot |
| Upstream inputs | Retrieved documents, tool results, prompt edits, the date in the system prompt | Log the full request and replay it |
The first row is the only one temperature controls. At temperature above zero a sampler can draw and return the first token whose cumulative probability exceeds , so a pseudo-random generator with a fixed seed reproduces the draws. Engines that batch many users must give each request its own reproducible stream rather than one shared generator: SGLang's deterministic mode, for example, draws its sampling noise from a seeded hash function, so that "the same (inputs, seed) pair always yields the same sample". Hosted APIs that accept a seed describe it as best effort, some report a fingerprint of the serving configuration so you can see when the backend changed, and some offer no seed at all. And a seed fixes only : if the arithmetic upstream moves a cumulative probability across it, the token changes anyway.
Floating-point addition is not associative
A float32 number carries a 24-bit significand, so each rounding has relative error at most . Between and consecutive float32 numbers are apart. Write for computed addition.
Claim. In float32, but .
Proof. lies between and , so it is representable. The exact sum is 1 from and 7 from the next float, , so it rounds to , and subtracting leaves exactly 0. On the right, exactly, and . The two groupings differ, so is not associative.
A dot product of length is a long sum. Added left to right, the computed result obeys ; a tree of pairwise sums improves the factor to about . Every order gives a slightly different answer within these bounds, and usually nobody cares. Greedy decoding cares when the top two logits are closer than the rounding error: then the order of additions picks the token, and since every later token is conditioned on it, the texts part ways from there.
In float32, evaluate , adding left to right.
Why temperature zero still varies
The folk explanation runs: GPUs add in parallel, threads finish in unpredictable orders, floating-point addition is not associative, so results wobble. He and Thinking Machines Lab call this incomplete. Run the same matrix multiplication on the same data and you get bitwise identical results, run after run. An order truly decided by racing threads needs atomic adds, and a typical LLM forward pass usually contains none. Each forward pass is run-to-run deterministic.
What varies is the batch. A server packs concurrent requests into one forward pass, and a fast kernel chooses its strategy from the shapes it receives. With few rows a matrix multiplication has too few output tiles to keep every core busy, so it splits the long reduction across cores and adds the partial sums afterwards (split-K); with many rows it doesn't. A different split is a different order of additions, so the numbers computed for your row depend on how many strangers arrived with it. In the post's words, "from the perspective of an individual user, the other concurrent users are not an 'input' to the system but rather a nondeterministic property". CPU and TPU servers have the same problem.
Batch invariance. A kernel is batch-invariant if the row it computes for input is bit-for-bit the same in every batch that contains . The post states the requirement plainly: "the reduction order for each element must be fixed regardless of the batch-size of the kernel."
Three operations in a transformer reduce over many numbers: RMSNorm, matrix multiplication and attention. The post makes each one batch-invariant. Normalisation and matrix multiplication use one kernel configuration for every shape, giving up split-K. Attention must also add a token's keys in the same order whether they arrive in one prefill or one at a time from the KV cache (as with chunked prefill and prefix caching), so the cache is laid out consistently before the kernel runs, and where the keys are split across cores, the size of each split is fixed rather than the number of splits. With these kernels all 1,000 Feynman completions were identical. The price was speed: on one GPU running Qwen3-8B, 1,000 sequences took 26 s with vLLM's default kernels, 55 s with an unoptimised deterministic version and 42 s with an improved attention kernel.
Mixture-of-experts models add a discrete choice: a router sends each token to its top-scoring experts, so a rounding-level change in router scores can redirect a token, and where each expert has a capacity limit, other requests' tokens compete for its slots. The Feynman model was itself a mixture of experts, and batch-invariant kernels were enough to make it repeat itself.
According to He and Thinking Machines Lab (2025), what mainly makes a served LLM give different answers to identical temperature-zero requests?
Making runs reproducible
Temperature zero is not a guarantee: it removes the sampler's dice, not the arithmetic or the inputs. To repeat a result, whether to debug, to audit or to compare two prompts fairly, control what you can and record the rest:
- Pin the model. Request a dated snapshot, never an alias the provider can move to new weights.
- Pin every parameter. Temperature, top-p, maximum tokens, stop sequences, system prompt, tool schemas and response format. Defaults change too.
- Seed where supported, knowing that hosted APIs promise only best effort.
- Log everything. The full request and response, the model version string, any system fingerprint, timestamps, and the retrieved context and tool results that went into the prompt.
- Cache by a hash of the request, so identical requests return the stored response.
- Record and replay in tests. Unit tests replay recorded responses; a separate, scheduled suite calls the live model.
- Self-host for bitwise reproducibility. Some open-source inference engines, vLLM and SGLang among them, now offer a batch-invariant or deterministic mode, at a cost in throughput and with limits on supported hardware and models. Check the current documentation before you rely on it.
import hashlib, json
def request_key(request: dict) -> str:
# Sorted keys and fixed separators make the JSON canonical: same request, same key.
canonical = json.dumps(request, sort_keys=True, separators=(",", ":"), ensure_ascii=False)
return hashlib.sha256(canonical.encode("utf-8")).hexdigest()
def cached_call(call, request: dict, store: dict) -> dict:
key = request_key(request) # model snapshot, every parameter, full messages, tools
if key not in store:
store[key] = call(**request) # record on a miss
return store[key] # replay on a hit
Variety on purpose
Sometimes the same answer twice is the failure. Brainstorming, creative writing, test-case generation and synthetic training data all want spread.
| Lever | How | Watch out for |
|---|---|---|
| Temperature, top-p | Raise them (Chamber 4) | Coherence falls in the tail |
| samples | Several completions of one prompt | They cluster around the prompt's modes |
| Seeds | A new, recorded seed per call | Varies only the draws |
| Prompt variation | Personas, constraints, shuffled example order | Record which variant produced what |
| Ask for a list | "Ten options that differ in setting and tone" in one call | Later items may weaken |
| Deduplicate | Embed, then drop pairs above a cosine threshold | The threshold is task-specific |
Measure what you get. Distinct-n is the number of unique -grams across all samples divided by the total number of -grams (some papers divide by the total number of tokens instead); mean pairwise cosine distance between embeddings catches paraphrases that -grams miss. For synthetic data, check for near-duplicates before you train on it.
Many samples, one better answer
When the task has a checkable final answer, variety buys accuracy. Self-consistency (Wang et al., 2023) samples chain-of-thought reasoning paths, parses each final answer , and returns the most frequent one: the opening equation. Its intuition is that correct reasoning paths, however different, tend to agree on the final answer more often than incorrect ones. With PaLM-540B and 40 sampled paths at temperature 0.7, it raised GSM8K accuracy from 56.5% with greedy chain of thought to 74.4%.
A clean model explains when voting helps. If each sample is right with probability , independently, and answers are binary, the vote over an odd number of samples is right with probability
which rises towards 1 when and falls towards 0 when : voting amplifies whatever the model usually says. All samples agree with probability , so low agreement flags a question worth routing to a stronger model or a person. Real samples share blind spots, so gains flatten sooner than the formula suggests; when wrong answers scatter over many values, plurality voting can help even with (Problem 3).
Each sample's binary answer is correct with probability , independently. What is the accuracy of a majority vote over 3 samples?
Best-of-n with a verifier samples answers and keeps the one a verifier scores highest: unit tests, a schema check, a calculator or a reward model. It works where voting can't, as with code and prose. pass@k asks whether at least one of samples passes. Chen et al. (2021) estimate it from samples of which pass, as ; the plug-in is biased low, and the proof puzzle below shows why the first is unbiased.
Testing a system that never answers the same way twice
A test that compares the reply with a stored string fails on the first harmless rephrasing. Assert properties instead: the reply parses and validates against its schema; required fields or facts are present and forbidden ones absent; numbers, lengths and labels fall in their allowed ranges; for free text, similarity to a reference or a model grader's verdict clears a threshold.
import json
def acceptable(reply: str) -> bool:
try:
data = json.loads(reply)
except json.JSONDecodeError:
return False
return data.get("label") in {"refund", "exchange", "other"} and 0 <= data.get("confidence", -1) <= 1
def trials(generate, prompt: str, n: int = 30) -> tuple[int, int]:
return sum(acceptable(generate(prompt)) for _ in range(n)), n # passes, trials
Then treat the pass rate as an estimate. With passes in trials and , the Wilson interval
stays honest near 0 and 1, where the textbook shrinks to a point. Eighteen passes in twenty give but a 95% interval of about . Set thresholds before you look: ship only if the lower end clears, say, 0.9; call a test flaky when it both passes and fails at fixed inputs; and quarantine flaky tests instead of rerunning until green, which measures only your patience. Chamber 7 builds this into full eval suites, with graders and paired comparisons between prompt versions.
Read beyond
Article · free online · ~35 min
Defeating Nondeterminism in LLM InferenceHorace He and Thinking Machines Lab (2025) · From “The original sin: floating-point non-associativity” to “Experiments”
Find the sentence that rejects "concurrency plus floating point" as the whole story, and list the three operations they make batch-invariant.
Paper · free online · ~45 min
What Every Computer Scientist Should Know About Floating-Point ArithmeticDavid Goldberg (ACM Computing Surveys, 1991) · Rounding Error: relative error, ulps and cancellation
Redo this chamber's example in units in the last place.
Paper · free online · ~30 min
Non-Determinism of “Deterministic” LLM SettingsBerk Atil et al. (arXiv, 2024) · Section 5.2 (metrics) and Section 8.1 (implications for practical engineering)
Note how far accuracy moved across ten identical runs, and how TARr@N differs from TARa@N.
Article · free online · ~20 min
Towards Deterministic Inference in SGLang and Reproducible RL TrainingThe SGLang Team (LMSYS, 2025)
Find how they keep sampling reproducible at temperatures above zero, and what deterministic mode costs.
Read the equation in context
Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery & Denny Zhou · ICLR, 2023Section 2 defines the vote. Each sampled output is a pair : a reasoning path and the final answer parsed from it, and voting over the marginalises out the paths. The paper also weights each vote by the model's probability of the whole output: normalised by length, that scored about the same as the plain vote on PaLM-540B, while the unnormalised version did clearly worse (Table 1).
Evaluating Large Language Models Trained on CodeMark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan et al. · arXiv, 2021The Codex paper introduced HumanEval and made "try times" a metric. It draws samples per problem, counts the that pass the unit tests, and estimates pass@k for every from the same samples.
Decode the paper · Section 2.1, Equation (1)
Evaluating Large Language Models Trained on CodeMark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan et al. · arXiv, 2021
Chen et al. judge generated code by running it: a problem counts as solved if any of samples passes its unit tests. Drawing exactly samples per problem gives a noisy estimate, so they draw , count the that pass, and compute the probability that a random -subset contains at least one pass. Appendix A shows this is unbiased, while plugging the empirical pass rate into underestimates.
Options
Your turn
The lab builds the Feynman fork in miniature: two candidate tokens whose float32 logits differ by a tenth of a float32 step in exact arithmetic. Find a batch size that flips greedy decoding, then switch on the batch-invariant kernel and show the flip is gone. The second panel turns the voting formulas into curves.
Interactive lab
Batch flip
16 chunks of 128 terms, each summed in order, then the 16 partial sums added in order.
logit “Queens”
11.1170597
logit “New”
11.1170578
Gap, float32 steps
+2
Cores busy
16/16
| Batch | Chunks | “Queens” | “New” | Greedy |
|---|---|---|---|---|
| 1–8 | 16 | 11.1170597 | 11.1170578 | Queens |
| 9–16 | 8 | 11.1170616 | 11.1170568 | Queens |
| 17–32 | 4 | 11.1170568 | 11.1170588 | New |
| 33–64 | 2 | 11.1170654 | 11.1170635 | Queens |
| 65–128 | 1 | 11.1170683 | 11.1170607 | Queens |
| float64 | 11.1170596533 | 11.1170595527 | Queens | |
Many samples, one answer
Unscored. Each sample is right with probability p, independently, and answers are binary; an even split is settled by a coin. Real samples share blind spots, so real gains are usually smaller.
One sample
0.650
Vote of 5
0.765
All 5 agree
0.121
Match · Symptom ↔ Fix
Symptoms and fixes
Options
Match · Expression ↔ Meaning
Formulas for many samples
Options
Proof puzzle
The pass@k estimator is unbiased
Claim
Draw independent samples, each passing with probability , and let count the passes. Show that .
Tap lines in the order they should appear. Not every line belongs. Tap a line in your proof to send it back.
Your proof
- Pick the first line below.
Available lines
Prove it yourself
Two out of three beats one
Claim
Three independent samples each give the correct answer with probability and otherwise the same wrong answer. Prove that their majority vote is correct with probability , and that this beats a single sample exactly when .
Your typeset proof appears here.
Coding problems
Problem 13·Warm-up
A perfect record, how sure?
A regression test of an LLM feature is run times at fixed inputs and passes every time. Using the 95% Wilson score interval with , find the smallest for which the lower end of the interval for the pass rate is at least .
Problem 14·Standard
Same numbers, different sum
Let be rounded to float32, for . Add them in float32, rounding after every addition. adds them in increasing order of (largest term first) and in decreasing order (smallest term first). Give to 6 decimal places. Careful: np.sum adds pairwise, which is a third order.
Problem 15·Challenge
Votes when mistakes scatter
A model's final answer to a question is correct with probability . Otherwise it gives one of three wrong answers, with probabilities , and . You draw 15 independent samples and return the most frequent answer, breaking any tie uniformly at random among the tied answers. What is the probability that you return the correct answer? Give 6 decimal places.
Key takeaways
- Temperature zero removes the sampler's randomness, not the arithmetic: floating-point addition is not associative, and at a near-tie the order of additions picks the token.
- In served models the main culprit is kernels that are not batch-invariant, so a request's numerics depend on how many others share its batch. Batch-invariant kernels fix this at a cost in speed.
- Reproducibility is a practice: pin the snapshot and every parameter, seed where you can, log full requests and responses, cache by hash and replay in tests.
- For variety, vary temperature, seeds and prompts and deduplicate; for accuracy, turn many samples into one answer with self-consistency, a verifier or pass@k.
- Test properties over repeated trials and report a pass rate with a Wilson interval, never a single run.
Checkpoint
Prove it to the labyrinth
Answer every question to clear this chamber. First-try answers earn the most XP.
Which practice protects you against a provider pointing a model name at new weights?
Binary answers; each sample is correct with probability , independently. What is the probability that all 5 samples agree?
You draw samples for a coding problem and pass the tests. What is the unbiased estimate of pass@4?
A test asserts that the model's reply equals a stored string. What should replace it?
Which statement about batch-invariant kernels is true?
A test passes 9 of 10 trials. What is the lower end of the 95% Wilson interval for its pass rate, with ?
Binary answers; each sample is correct with probability , independently. What happens to majority-vote accuracy as the number of samples grows?
You need 200 varied product names as synthetic data. Which plan gives the most varied set?
End of the chamber
Clear this chamber
- Questions in this chamber (0/11 solved)Next unsolved
- Bonus: Batch flip (+40 XP)
- Bonus: Problem 13: A perfect record, how sure? (+20 XP)
- Bonus: Problem 14: Same numbers, different sum (+35 XP)
- Bonus: Problem 15: Votes when mistakes scatter (+50 XP)
- Bonus: Proof: The pass@k estimator is unbiased (+25 XP)
- Bonus: Proof: Two out of three beats one (+35 XP)
- Bonus: Decode the paper (+25 XP)
- Bonus: Match: Symptoms and fixes (+20 XP)
- Bonus: Match: Formulas for many samples (+20 XP)