A teammate edits one sentence of the support bot's prompt, tries it on a single customer question, and the answer looks better. Two questions decide whether that edit should reach production: exactly what changed, and did it help? The first is version control. The second is an experiment, and Evan Miller's statistical guide to evals gives it an error bar: run both versions on the same items and take the standard error of the per-item differences.
Spotted in the wild
- “s sub i”Score of item : 1 for a pass and 0 for a fail, or the fraction of sampled answers that pass.
- “p hat”Observed pass rate, the mean item score over the items of the suite.
- “standard error”How much an estimate would vary if the eval items were drawn again.
- “d sub i”Paired difference on item : the new version's score minus the old version's score on the same item.
- “covariance of a and b”How two versions' item scores move together; positive when both find the same items hard.
- “K”Number of answers sampled per item and version, averaged into the item score.
- “half-width”Half the width of a confidence interval: at 95%.
- “b and c”Discordant counts: items only the old version passes, and items only the new version passes.
- “semantic version”A version number whose parts signal a breaking change, new compatible behaviour, or a fix.
| Symbol | Say it | Meaning | LaTeX |
|---|---|---|---|
| “s sub i” | Score of item : 1 for a pass and 0 for a fail, or the fraction of sampled answers that pass. | ||
| “p hat” | Observed pass rate, the mean item score over the items of the suite. | ||
| “standard error” | How much an estimate would vary if the eval items were drawn again. | ||
| “d sub i” | Paired difference on item : the new version's score minus the old version's score on the same item. | ||
| “covariance of a and b” | How two versions' item scores move together; positive when both find the same items hard. | ||
| “K” | Number of answers sampled per item and version, averaged into the item score. | ||
| “half-width” | Half the width of a confidence interval: at 95%. | ||
| “b and c” | Discordant counts: items only the old version passes, and items only the new version passes. | ||
| “semantic version” | A version number whose parts signal a breaking change, new compatible behaviour, or a fix. |
Prompts are code
A one-word change to a prompt can shift behaviour as much as a code change, so give prompts the same discipline as code.
- Files, not literals. Keep each prompt in its own file in the repository, not as strings scattered through the app.
- Templates with typed variables. Render with Jinja or Python format strings, and validate the inputs first. A missing
policyshould raise an error (Jinja'sStrictUndefineddoes this), not render as an empty string. - Reviewed diffs. Changes go through pull requests, where reviewers read a line-by-line diff of the prompt text, not a blob of escaped JSON.
- Tested in CI. A change under
prompts/runs the eval suite, just as a code change runs the unit tests.
# prompts/support_answer.prompt
---
id: support-answer
version: 1.4.0
model: provider-model-2026-05-14 # a dated snapshot, never "latest"
params: {temperature: 0.2, top_p: 0.95, max_tokens: 400, seed: 7}
inputs: {question: str, policy: str}
tools: [tools/lookup_order.json]
retrieval: {index: policies-2026-09, embedding: embed-2026-03, k: 5}
output_schema: schemas/support_answer.v1.json
changelog: "1.4.0 (MINOR): quote the policy section relied on"
---
Answer {{ question }} using only the policy below.
Quote the section you relied on, e.g. (§4.2).
Policy: {{ policy }}
A release is everything that shapes the output
The text is only one part. A release pins every input that changes what the model sees or how it samples:
| Part | Example | Why it must be pinned |
|---|---|---|
| Template | support_answer.prompt at 1.4.0 | Wording changes behaviour |
| Model snapshot | provider-model-2026-05-14 | An alias such as latest can be pointed at a new model by the provider |
| Sampling parameters | temperature, top-p, max tokens, seed | The Middle Ring showed how far these move outputs |
| Tool schemas | lookup_order.json | The model reads them as part of its prompt |
| Retrieval configuration | index version, embedding model, k | They decide what enters the context |
Fingerprint the release by hashing a canonical serialisation of all of it, and log the fingerprint with every call, next to the inputs and outputs. Any output in your logs can then be traced to exactly what produced it. Two calls with the same fingerprint can still differ, but only through sampling randomness (the Non-Determinism chamber), never through a silent configuration change. A fingerprint is only as good as its parts, though: the string latest hashes the same before and after the provider moves it, which is why the release names a dated snapshot.
import hashlib, json
def fingerprint(release: dict) -> str:
# Sorted keys and fixed separators: equal releases always hash equally.
canonical = json.dumps(release, sort_keys=True, separators=(",", ":"))
return hashlib.sha256(canonical.encode("utf-8")).hexdigest()
release = {
"template": open("prompts/support_answer.prompt").read(),
"model": "provider-model-2026-05-14",
"params": {"temperature": 0.2, "top_p": 0.95, "max_tokens": 400, "seed": 7},
"tools": [json.load(open("tools/lookup_order.json"))],
"retrieval": {"index": "policies-2026-09", "embedding": "embed-2026-03", "k": 5},
}
fp = fingerprint(release) # log fp[:12] with every call
The version number is a name for people; the fingerprint is the content. Semantic Versioning's build metadata can carry both, as in 1.4.0+fp.3c9a71d2 (an illustrative hash), because build metadata is ignored when versions are ordered.
Versions, labels and rollout
Semantic Versioning fits prompts once you treat the output contract as the public API:
| Bump | When | Prompt example |
|---|---|---|
| MAJOR | The output contract changes incompatibly | A schema field renamed, the label set changed |
| MINOR | New behaviour that keeps the contract | Cite the policy section; add an optional field |
| PATCH | A fix meant to change nothing | A typo, clearer wording |
- Versions are immutable. Once released, 1.4.0 never changes; any edit is a new version. Precedence compares MAJOR, MINOR and PATCH numerically from the left, so 1.10.0 follows 1.9.3, and a pre-release such as 2.0.0-rc.1 comes before 2.0.0.
- Labels move.
prod,stagingandcanarypoint at versions, like container image tags. If the app resolves its label at runtime, with a cached fallback, rolling back means pointingprodat the previous version: no redeploy. - Every version gets a changelog entry: what changed, why, and the eval result that justified it.
- Roll out gradually. Point
canaryat the new version for a small slice of traffic, or run an A/B test, and watch online metrics before promoting. - Snapshots retire. Providers deprecate model snapshots on published schedules. Moving to a successor is a release like any other: rerun the full suite before any label moves, and expect some prompts to need retuning.
A PATCH that "changes nothing" is a claim, and claims get tested. The lab's third release is one.
Version 1.4.0 of a prompt returns JSON with a field answer. The new version renames it to reply and changes nothing else. What is the new version number?
Git or a prompt registry?
Prompt-management tools such as Langfuse, MLflow Prompt Registry, PromptLayer and Braintrust (among others) store versioned prompts with movable labels or aliases, an editing interface and an API for fetching prompts at runtime.
| Question | Prompts in Git | Prompt registry |
|---|---|---|
| Who can edit? | Engineers, through pull requests | Also product and domain experts, in a web interface |
| How does prod change? | A deploy | Moving a label, fetched at runtime |
| Audit trail | Commits, blame and review comments | The registry's version history |
| Main risk | Slow iteration for non-engineers | Prompts drift away from the code and the evals that test them |
Common hybrids keep Git as the source of truth and let CI publish to the registry, or keep the registry as the source and have CI run the evals before a label may move. Whichever you choose, the rule is the same: no text reaches prod without passing the same eval gate.
Eval suites
An eval suite is a versioned dataset, a set of graders and a decision rule.
The golden dataset comes from real traces: sample production logs, label them, add every bug you fix as a regression item, and cover the slices that matter (languages, long inputs, adversarial users). Version the dataset too, and record which version produced each score.
Graders, cheapest first:
| Grader | Good for | Watch out for |
|---|---|---|
| Code assertions | Valid schema, a citation present, length limits, banned strings | Only checks what can be written down exactly |
| LLM-as-a-judge | Faithfulness to sources, tone, helpfulness, against a rubric | Biases, cost, and drift when the judge's model changes |
| Human review | Gold labels, ambiguous cases, calibrating the judge | Slow, costly, and raters disagree |
An LLM judge works pointwise (grade one output against a rubric, best as a binary pass or fail per criterion) or pairwise (which of two outputs is better). Pairwise suits a head-to-head between two versions; pointwise gives an absolute pass rate that scales to many versions. Zheng et al. found that strong judges agreed with human preferences over 80% of the time, about as often as humans agreed with each other, but also measured position bias (favouring the first answer), verbosity bias (favouring longer answers) and possible self-enhancement (favouring answers from the judge's own model). Later work (Panickssery, Bowman and Feng, 2024) found that models do recognise and favour their own generations. Swap the order and count only consistent wins, use a judge from a different model family, and calibrate: label a sample by hand and measure the judge's agreement on passes and on failures separately. Shankar et al. observed criteria drift: grading outputs changes what people think the criteria should be, so calibration is never finished. The judge prompt is a prompt too, so version it.
Gate merges in CI: run the suite on every pull request that touches prompts, post the comparison table, and block the merge if the paired interval shows a regression on a critical slice. After release, watch online metrics (user feedback, escalations, task completion, cost and latency), because offline evals only cover the inputs you thought of.
A pairwise LLM judge prefers whichever answer it sees first more often than chance. What is the conservative fix described by Zheng et al. (2023)?
Error bars on a pass rate
Treat the items as a sample from every input you care about. With pass/fail scores, the pass rate has standard error (Miller's Equation 2) and a 95% interval . Since , the standard error is at most , so the 95% half-width is at most : points for 100 items, for 400 and for 1,000. Most prompt suites are small, so most differences you see on them are noise.
Non-determinism adds a second layer: one item can pass on one sample and fail on the next. Score each item as the mean of samples. If is item 's true pass probability and the variance of one sampled answer's score,
Resampling shrinks only the second term; the spread of item difficulty is a floor that only more items lower. Miller also warns against lowering the temperature to reduce variance: evaluate at the settings you ship.
An eval suite of 400 items has a pass rate of 80%. What is the standard error of the pass rate, as a proportion?
Compare versions on the same items
Both versions answer the same items, and an item that is hard for one is usually hard for the other. Compare item by item, with :
The unpaired standard error ignores the covariance. When the versions agree on which items are hard, and pairing gives a narrower interval at no extra cost; you only need to keep the item ids aligned.
For pass/fail scores, McNemar's test looks only at the discordant items: items only the old version passes, and items only the new one passes. If neither version is better, each discordant item is equally likely to fall either way, so compare with 3.84, the 95% point of . When is small, use the exact binomial test instead.
Plan the size before you run: for an interval of half-width you need items, with from a pilot run. To detect a difference with given power, Miller's sample-size formula adds a second quantile (the decoder below).
Don't overfit the eval set. If you try 40 variants and keep the best, its dev score is biased upward by the selection alone (Problem 3 computes how much). Keep a held-out set that is used only for the release decision, and refresh the golden set from new traces. And don't peek: stopping the first time the interval clears zero raises the false-positive rate well above 5%. In CI, fix in advance, or use a sequential test designed for repeated looks.
import numpy as np
def paired_ci(new, old, z=1.96):
d = np.asarray(new, float) - np.asarray(old, float) # same items, same order
se = d.std(ddof=1) / np.sqrt(len(d))
return d.mean() - z * se, d.mean() + z * se
def decide(lo, hi, margin=0.03):
if lo > 0: return "ship"
if hi < 0: return "roll back"
if -margin < lo and hi < margin: return "hold"
return "keep evaluating"
Read beyond
Article · free online · ~30 min
Your AI Product Needs EvalsHamel Husain · The Types Of Evaluation: levels 1–3
Sort your own product's checks into unit tests, model and human evals, and A/B tests.
Article · free online · ~35 min
Using LLM-as-a-Judge For Evaluation: A Complete GuideHamel Husain · Steps 3–5, then the FAQ on validating a judge
Note how the judge is checked against a domain expert's labels before anyone trusts it.
Article · free online · ~15 min
Semantic Versioning 2.0.0Tom Preston-Werner · Specification items 3, 6–8 and 11
Rewrite items 6–8 with "the output contract" in place of "the public API".
Article · free online · ~30 min
What We've Learned From A Year of Building with LLMsYan, Bischof, Frye, Husain, Liu & Shankar · Sections 1.4 (Evaluation & Monitoring) and 2.2.3 (Version and pin your models)
List the advice you would turn into a CI check.
Read the equation in context
Adding Error Bars to Evals: A Statistical Approach to Language Model EvaluationsEvan Miller · arXiv preprint, 2024Miller treats eval questions as draws from an unseen super-population, so a score is an estimate of a skill and deserves a standard error. The paper derives the standard error of a mean score, clusters it when questions come in related groups (several questions about one passage, or one question translated into many languages; in a product suite, several items cut from one conversation), reduces variance by resampling answers, compares two models with paired differences (the opening equation, Equation 7), and closes with the sample-size formula decoded here.
Decode the paper · Section 5 (Power analysis), Equation (9)
Adding Error Bars to Evals: A Statistical Approach to Language Model EvaluationsEvan Miller · arXiv preprint, 2024
Miller treats eval questions as draws from a large unseen population and asks how many questions an eval needs. The paired design enters through , the variance of the difference between the two models' conditional mean scores, while sampling noise enters divided by the number of answers per question.
Options
The study that measured judge agreement with human preferences, catalogued position and verbosity bias, and proposed swapping positions as a fix.
Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human PreferencesShreya Shankar, J. D. Zamfirescu-Pereira, Björn Hartmann, Aditya G. Parameswaran & Ian Arawjo · UIST 2024, 2024A tool and user study for aligning LLM graders with human grades, and the origin of the term criteria drift.
The talk surveys prompt engineering, retrieval and fine-tuning. For each step, ask how you would version it and which eval would show that it helped.
Your turn
Three releases wait on the canary label. Run the golden set in batches, watch the paired interval tighten inside the unpaired one, and make each call only when the interval backs it. The typo fix is the hardest: a hold needs a narrow interval, and there is more than one way to narrow it.
Interactive lab
Release manager
prompts/support_answer.promptMINOR
Ask the assistant to quote the policy section it relied on, and to hand over when the policy is silent.
Answer {{ question }} using only the policy below.+ Quote the section you relied on, e.g. (§4.2).+ If the policy does not cover it, say so and offer a human.Policy: {{ policy }}
v1.3.2 pass rate
n/a
v1.4.0 pass rate
n/a
Paired difference, pts
n/a
Paired 95% interval
at 30 items
Unpaired 95% interval
at 30 items
Unpaired ÷ paired width
n/a
Items 0/500 · model calls 0 · correlation of item scores n/a · paired SE n/a
Match · Change ↔ Release
Starting from 1.4.0, which release?
answer to replyconfidence field to the outputOptions
Match · Expression ↔ Meaning
Eval statistics
Options
Proof puzzle
Why pairing shrinks the error bar
Claim
For one randomly drawn item, let and be the old and new versions' scores. Show that , so that scoring both versions on the same items shrinks the variance of the difference whenever the scores are positively correlated.
Tap lines in the order they should appear. Not every line belongs. Tap a line in your proof to send it back.
Your proof
- Pick the first line below.
Available lines
Prove it yourself
Worst-case error bars
Claim
Prove that the standard error of a pass rate is largest at . Deduce that items guarantee a 95% half-width of at most , whatever the true pass rate.
Your typeset proof appears here.
Coding problems
Problem 19·Warm-up
Fingerprint a release
A release is the JSON object below. Serialise it canonically, with keys sorted at every level and no spaces (in Python, json.dumps(release, sort_keys=True, separators=(",", ":"))). Compute the SHA-256 of the UTF-8 bytes, read the first 8 hexadecimal digits as a base-16 integer, and give that integer in decimal.
{"template": "Answer {{ question }} using only {{ policy }}. Quote the section you relied on.", "model": "provider-model-2026-05-14", "params": {"temperature": 0.2, "top_p": 0.95, "max_tokens": 400, "seed": 7}, "tools": [{"name": "lookup_order", "parameters": {"order_id": "string"}}], "retrieval": {"index": "policies-2026-09", "k": 5}}
Problem 20·Standard
Paired beats unpaired
Old and new versions of a prompt were run on the same 600 golden items. Both passed 410 items, only the old version passed 38, only the new version passed 64, and both failed 88. Let . Compute the paired 95% interval , where is the sample standard deviation of the (dividing by ). Give its lower end to 4 decimal places.
Problem 21·Challenge
The winner's curse
You try prompt variants that are, in truth, identical: each passes each of dev items independently with probability . You keep the variant with the highest dev pass rate. Compute exactly, without simulation, the expected dev pass rate of the winner minus . Give 4 decimal places. Hint: for a random variable taking values in , , and the maximum of independent scores is below only if every score is.
Key takeaways
- A release is template, model snapshot, parameters, tools and retrieval together: version it immutably, fingerprint it, and log the fingerprint with every call.
- Semantic versions name the change for people; movable labels such as
prodmake rollback a pointer change. - An eval suite combines real-trace data, code assertions, calibrated LLM judges and human review, and gates every merge.
- A pass rate on items carries a 95% error bar of up to about . Compare versions on the same items, where pairing removes the difficulty they share.
- Plan before you run, don't stop at the first lucky look, and keep a held-out set for the final decision.
Checkpoint
Prove it to the labyrinth
Answer every question to clear this chamber. First-try answers earn the most XP.
Two prompt versions each have a standard error of 0.02 on the same 400 items, and their item scores have correlation 0.6. What is the standard error of the paired difference?
A pilot run shows the per-item differences between two versions have standard deviation . How many items do you need for the 95% interval on the mean difference to have half-width at most 0.02? Round up.
Why should a release pin a dated model snapshot rather than an alias such as latest?
In McNemar's test for two versions scored pass or fail on the same items, which items carry the evidence about which version is better?
You tried 40 prompt variants on a 200-item dev set and kept the best, which scored 84%. Which number estimates its production pass rate without bias?
Item difficulty contributes variance to item scores, and sampling noise contributes per answer. With items and answers per item, what is the standard error of the pass rate?
Production fetches a prompt from a registry at runtime by its prod label. What is the main risk to guard against?
Which version has the highest precedence under Semantic Versioning 2.0.0?
How should you decide whether an LLM judge can stand in for human reviewers on your task?
End of the chamber
Clear this chamber
- Questions in this chamber (0/12 solved)Next unsolved
- Bonus: Release manager (+50 XP)
- Bonus: Problem 19: Fingerprint a release (+20 XP)
- Bonus: Problem 20: Paired beats unpaired (+35 XP)
- Bonus: Problem 21: The winner's curse (+50 XP)
- Bonus: Proof: Why pairing shrinks the error bar (+25 XP)
- Bonus: Proof: Worst-case error bars (+35 XP)
- Bonus: Decode the paper (+25 XP)
- Bonus: Match: Starting from 1.4.0, which release? (+20 XP)
- Bonus: Match: Eval statistics (+20 XP)