Skip to content
AriadneTechnology

The Inner Ring · Chamber 7 of 9

Prompt Version Control and Evals

Treat prompts as code: version, review, pin and roll them back, and decide every release with an eval suite and honest error bars.

60 min 60 XP + 12 questions + 1 challengeMathVideoPapersProofsCodeLab

In this chamber you will

  • Store prompts as versioned, reviewed artefacts with semantic versions and changelogs
  • Define a release as prompt, model snapshot, parameters and tools, fingerprinted together
  • Build an eval suite from assertions, model graders and human review
  • Compare two prompt versions with paired differences and a confidence interval
DiscoverLearnRead beyondPapers & lecturesYour turn

A teammate edits one sentence of the support bot's prompt, tries it on a single customer question, and the answer looks better. Two questions decide whether that edit should reach production: exactly what changed, and did it help? The first is version control. The second is an experiment, and Evan Miller's statistical guide to evals gives it an error bar: run both versions on the same items and take the standard error of the per-item differences.

Spotted in the wild

SEA−B, paired=Var(sA−B)/n=(1n−1∑i(sA−B,i−sˉA−B)2)/n\mathrm{SE}_{A-B,\,\mathrm{paired}}=\sqrt{\mathrm{Var}(s_{A-B})/n}=\sqrt{\left(\frac{1}{n-1}\sum_i(s_{A-B,i}-\bar{s}_{A-B})^2\right)/n}
Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations
DiscoverLearnRead beyondPapers & lecturesYour turn
Symbols for this chamber
  • sis_i“s sub i”
    Score of item ii: 1 for a pass and 0 for a fail, or the fraction of KK sampled answers that pass.
  • p^\hat p“p hat”
    Observed pass rate, the mean item score over the nn items of the suite.
    p^=1n∑isi\hat p=\tfrac1n\textstyle\sum_i s_i
  • SE\mathrm{SE}“standard error”
    How much an estimate would vary if the eval items were drawn again.
    SE=p^(1−p^)/n\mathrm{SE}=\sqrt{\hat p(1-\hat p)/n}
  • did_i“d sub i”
    Paired difference on item ii: the new version's score minus the old version's score on the same item.
    di=snew,i−sold,id_i=s_{\text{new},i}-s_{\text{old},i}
  • Cov⁡(a,b)\operatorname{Cov}(a,b)“covariance of a and b”
    How two versions' item scores move together; positive when both find the same items hard.
  • KK“K”
    Number of answers sampled per item and version, averaged into the item score.
  • hh“half-width”
    Half the width of a confidence interval: 1.96 SE1.96\,\mathrm{SE} at 95%.
  • b, cb,\ c“b and c”
    Discordant counts: items only the old version passes, and items only the new version passes.
  • MAJOR.MINOR.PATCH\text{MAJOR.MINOR.PATCH}“semantic version”
    A version number whose parts signal a breaking change, new compatible behaviour, or a fix.
    1.4.0→1.4.11.4.0\to1.4.1

Prompts are code

A one-word change to a prompt can shift behaviour as much as a code change, so give prompts the same discipline as code.

  • Files, not literals. Keep each prompt in its own file in the repository, not as strings scattered through the app.
  • Templates with typed variables. Render with Jinja or Python format strings, and validate the inputs first. A missing policy should raise an error (Jinja's StrictUndefined does this), not render as an empty string.
  • Reviewed diffs. Changes go through pull requests, where reviewers read a line-by-line diff of the prompt text, not a blob of escaped JSON.
  • Tested in CI. A change under prompts/ runs the eval suite, just as a code change runs the unit tests.
Text
# prompts/support_answer.prompt
---
id: support-answer
version: 1.4.0
model: provider-model-2026-05-14        # a dated snapshot, never "latest"
params: {temperature: 0.2, top_p: 0.95, max_tokens: 400, seed: 7}
inputs: {question: str, policy: str}
tools: [tools/lookup_order.json]
retrieval: {index: policies-2026-09, embedding: embed-2026-03, k: 5}
output_schema: schemas/support_answer.v1.json
changelog: "1.4.0 (MINOR): quote the policy section relied on"
---
Answer {{ question }} using only the policy below.
Quote the section you relied on, e.g. (§4.2).
Policy: {{ policy }}

A release is everything that shapes the output

The text is only one part. A release pins every input that changes what the model sees or how it samples:

PartExampleWhy it must be pinned
Templatesupport_answer.prompt at 1.4.0Wording changes behaviour
Model snapshotprovider-model-2026-05-14An alias such as latest can be pointed at a new model by the provider
Sampling parameterstemperature, top-p, max tokens, seedThe Middle Ring showed how far these move outputs
Tool schemaslookup_order.jsonThe model reads them as part of its prompt
Retrieval configurationindex version, embedding model, kThey decide what enters the context

Fingerprint the release by hashing a canonical serialisation of all of it, and log the fingerprint with every call, next to the inputs and outputs. Any output in your logs can then be traced to exactly what produced it. Two calls with the same fingerprint can still differ, but only through sampling randomness (the Non-Determinism chamber), never through a silent configuration change. A fingerprint is only as good as its parts, though: the string latest hashes the same before and after the provider moves it, which is why the release names a dated snapshot.

Python
import hashlib, json

def fingerprint(release: dict) -> str:
    # Sorted keys and fixed separators: equal releases always hash equally.
    canonical = json.dumps(release, sort_keys=True, separators=(",", ":"))
    return hashlib.sha256(canonical.encode("utf-8")).hexdigest()

release = {
    "template": open("prompts/support_answer.prompt").read(),
    "model": "provider-model-2026-05-14",
    "params": {"temperature": 0.2, "top_p": 0.95, "max_tokens": 400, "seed": 7},
    "tools": [json.load(open("tools/lookup_order.json"))],
    "retrieval": {"index": "policies-2026-09", "embedding": "embed-2026-03", "k": 5},
}
fp = fingerprint(release)   # log fp[:12] with every call

The version number is a name for people; the fingerprint is the content. Semantic Versioning's build metadata can carry both, as in 1.4.0+fp.3c9a71d2 (an illustrative hash), because build metadata is ignored when versions are ordered.

Versions, labels and rollout

Semantic Versioning fits prompts once you treat the output contract as the public API:

BumpWhenPrompt example
MAJORThe output contract changes incompatiblyA schema field renamed, the label set changed
MINORNew behaviour that keeps the contractCite the policy section; add an optional field
PATCHA fix meant to change nothingA typo, clearer wording
  • Versions are immutable. Once released, 1.4.0 never changes; any edit is a new version. Precedence compares MAJOR, MINOR and PATCH numerically from the left, so 1.10.0 follows 1.9.3, and a pre-release such as 2.0.0-rc.1 comes before 2.0.0.
  • Labels move. prod, staging and canary point at versions, like container image tags. If the app resolves its label at runtime, with a cached fallback, rolling back means pointing prod at the previous version: no redeploy.
  • Every version gets a changelog entry: what changed, why, and the eval result that justified it.
  • Roll out gradually. Point canary at the new version for a small slice of traffic, or run an A/B test, and watch online metrics before promoting.
  • Snapshots retire. Providers deprecate model snapshots on published schedules. Moving to a successor is a release like any other: rerun the full suite before any label moves, and expect some prompts to need retuning.

A PATCH that "changes nothing" is a claim, and claims get tested. The lab's third release is one.

Quick check +20 XP

Version 1.4.0 of a prompt returns JSON with a field answer. The new version renames it to reply and changes nothing else. What is the new version number?

Git or a prompt registry?

Prompt-management tools such as Langfuse, MLflow Prompt Registry, PromptLayer and Braintrust (among others) store versioned prompts with movable labels or aliases, an editing interface and an API for fetching prompts at runtime.

QuestionPrompts in GitPrompt registry
Who can edit?Engineers, through pull requestsAlso product and domain experts, in a web interface
How does prod change?A deployMoving a label, fetched at runtime
Audit trailCommits, blame and review commentsThe registry's version history
Main riskSlow iteration for non-engineersPrompts drift away from the code and the evals that test them

Common hybrids keep Git as the source of truth and let CI publish to the registry, or keep the registry as the source and have CI run the evals before a label may move. Whichever you choose, the rule is the same: no text reaches prod without passing the same eval gate.

Eval suites

An eval suite is a versioned dataset, a set of graders and a decision rule.

The golden dataset comes from real traces: sample production logs, label them, add every bug you fix as a regression item, and cover the slices that matter (languages, long inputs, adversarial users). Version the dataset too, and record which version produced each score.

Graders, cheapest first:

GraderGood forWatch out for
Code assertionsValid schema, a citation present, length limits, banned stringsOnly checks what can be written down exactly
LLM-as-a-judgeFaithfulness to sources, tone, helpfulness, against a rubricBiases, cost, and drift when the judge's model changes
Human reviewGold labels, ambiguous cases, calibrating the judgeSlow, costly, and raters disagree

An LLM judge works pointwise (grade one output against a rubric, best as a binary pass or fail per criterion) or pairwise (which of two outputs is better). Pairwise suits a head-to-head between two versions; pointwise gives an absolute pass rate that scales to many versions. Zheng et al. found that strong judges agreed with human preferences over 80% of the time, about as often as humans agreed with each other, but also measured position bias (favouring the first answer), verbosity bias (favouring longer answers) and possible self-enhancement (favouring answers from the judge's own model). Later work (Panickssery, Bowman and Feng, 2024) found that models do recognise and favour their own generations. Swap the order and count only consistent wins, use a judge from a different model family, and calibrate: label a sample by hand and measure the judge's agreement on passes and on failures separately. Shankar et al. observed criteria drift: grading outputs changes what people think the criteria should be, so calibration is never finished. The judge prompt is a prompt too, so version it.

Gate merges in CI: run the suite on every pull request that touches prompts, post the comparison table, and block the merge if the paired interval shows a regression on a critical slice. After release, watch online metrics (user feedback, escalations, task completion, cost and latency), because offline evals only cover the inputs you thought of.

Quick check +20 XP

A pairwise LLM judge prefers whichever answer it sees first more often than chance. What is the conservative fix described by Zheng et al. (2023)?

Error bars on a pass rate

Treat the nn items as a sample from every input you care about. With pass/fail scores, the pass rate p^\hat p has standard error p^(1−p^)/n\sqrt{\hat p(1-\hat p)/n} (Miller's Equation 2) and a 95% interval p^±1.96 SE\hat p\pm1.96\,\mathrm{SE}. Since p(1−p)≤1/4p(1-p)\le1/4, the standard error is at most 1/(2n)1/(2\sqrt n), so the 95% half-width is at most 0.98/n0.98/\sqrt n: ±9.8\pm9.8 points for 100 items, ±4.9\pm4.9 for 400 and ±3.1\pm3.1 for 1,000. Most prompt suites are small, so most differences you see on them are noise.

Non-determinism adds a second layer: one item can pass on one sample and fail on the next. Score each item as the mean of KK samples. If xix_i is item ii's true pass probability and σi2\sigma_i^2 the variance of one sampled answer's score,

Var⁡(p^)=Var⁡(x)+E[σi2]/Kn.\operatorname{Var}(\hat p)=\frac{\operatorname{Var}(x)+\mathbb E[\sigma_i^2]/K}{n}.

Resampling shrinks only the second term; the spread of item difficulty is a floor that only more items lower. Miller also warns against lowering the temperature to reduce variance: evaluate at the settings you ship.

Quick check +20 XP

An eval suite of 400 items has a pass rate of 80%. What is the standard error of the pass rate, as a proportion?

Compare versions on the same items

Both versions answer the same items, and an item that is hard for one is usually hard for the other. Compare item by item, with di=snew,i−sold,id_i=s_{\text{new},i}-s_{\text{old},i}:

Var⁡(d)=Var⁡(snew)+Var⁡(sold)−2Cov⁡(snew,sold),SEpaired=sdn.\operatorname{Var}(d)=\operatorname{Var}(s_{\text{new}})+\operatorname{Var}(s_{\text{old}})-2\operatorname{Cov}(s_{\text{new}},s_{\text{old}}),\qquad \mathrm{SE}_{\text{paired}}=\frac{s_d}{\sqrt n}.

The unpaired standard error SEnew2+SEold2\sqrt{\mathrm{SE}_{\text{new}}^2+\mathrm{SE}_{\text{old}}^2} ignores the covariance. When the versions agree on which items are hard, Cov⁡>0\operatorname{Cov}>0 and pairing gives a narrower interval at no extra cost; you only need to keep the item ids aligned.

For pass/fail scores, McNemar's test looks only at the discordant items: bb items only the old version passes, and cc items only the new one passes. If neither version is better, each discordant item is equally likely to fall either way, so compare (b−c)2/(b+c)(b-c)^2/(b+c) with 3.84, the 95% point of χ12\chi^2_1. When b+cb+c is small, use the exact binomial test instead.

Plan the size before you run: for an interval of half-width hh you need n≈(1.96 σd/h)2n\approx(1.96\,\sigma_d/h)^2 items, with σd\sigma_d from a pilot run. To detect a difference δ\delta with given power, Miller's sample-size formula adds a second quantile (the decoder below).

Don't overfit the eval set. If you try 40 variants and keep the best, its dev score is biased upward by the selection alone (Problem 3 computes how much). Keep a held-out set that is used only for the release decision, and refresh the golden set from new traces. And don't peek: stopping the first time the interval clears zero raises the false-positive rate well above 5%. In CI, fix nn in advance, or use a sequential test designed for repeated looks.

Python
import numpy as np

def paired_ci(new, old, z=1.96):
    d = np.asarray(new, float) - np.asarray(old, float)   # same items, same order
    se = d.std(ddof=1) / np.sqrt(len(d))
    return d.mean() - z * se, d.mean() + z * se

def decide(lo, hi, margin=0.03):
    if lo > 0: return "ship"
    if hi < 0: return "roll back"
    if -margin < lo and hi < margin: return "hold"
    return "keep evaluating"
DiscoverLearnRead beyondPapers & lecturesYour turn

Read beyond

Article · free online · ~30 min

Your AI Product Needs Evals

Hamel Husain · The Types Of Evaluation: levels 1–3

Sort your own product's checks into unit tests, model and human evals, and A/B tests.

Article · free online · ~35 min

Using LLM-as-a-Judge For Evaluation: A Complete Guide

Hamel Husain · Steps 3–5, then the FAQ on validating a judge

Note how the judge is checked against a domain expert's labels before anyone trusts it.

Article · free online · ~15 min

Semantic Versioning 2.0.0

Tom Preston-Werner · Specification items 3, 6–8 and 11

Rewrite items 6–8 with "the output contract" in place of "the public API".

Article · free online · ~30 min

What We've Learned From A Year of Building with LLMs

Yan, Bischof, Frye, Husain, Liu & Shankar · Sections 1.4 (Evaluation & Monitoring) and 2.2.3 (Version and pin your models)

List the advice you would turn into a CI check.

DiscoverLearnRead beyondPapers & lecturesYour turn

Read the equation in context

Adding Error Bars to Evals: A Statistical Approach to Language Model EvaluationsEvan Miller · arXiv preprint, 2024

Miller treats eval questions as draws from an unseen super-population, so a score is an estimate of a skill and deserves a standard error. The paper derives the standard error of a mean score, clusters it when questions come in related groups (several questions about one passage, or one question translated into many languages; in a product suite, several items cut from one conversation), reduces variance by resampling answers, compares two models with paired differences (the opening equation, Equation 7), and closes with the sample-size formula decoded here.

Decode the paper · Section 5 (Power analysis), Equation (9)

Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations

Evan Miller · arXiv preprint, 2024

+25 XP
n=(zα/2+zβ)2(ω2+σA2/KA+σB2/KB)/δ2n=(z_{\alpha/2}+z_\beta)^2(\omega^2+\sigma_A^2/K_A+\sigma_B^2/K_B)/\delta^2

Miller treats eval questions as draws from a large unseen population and asks how many questions an eval needs. The paired design enters through ω2=Var⁡(xA)+Var⁡(xB)−2Cov⁡(xA,xB)\omega^2=\operatorname{Var}(x_A)+\operatorname{Var}(x_B)-2\operatorname{Cov}(x_A,x_B), the variance of the difference between the two models' conditional mean scores, while sampling noise enters divided by the number of answers per question.

zα/2z_{\alpha/2}
zβz_\beta
ω2\omega^2
σA2\sigma_A^2
KAK_A
δ\delta

Options

Judging LLM-as-a-Judge with MT-Bench and Chatbot ArenaLianmin Zheng, Wei-Lin Chiang, Ying Sheng et al. · NeurIPS 2023 Datasets and Benchmarks Track, 2023

The study that measured judge agreement with human preferences, catalogued position and verbosity bias, and proposed swapping positions as a fix.

Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human PreferencesShreya Shankar, J. D. Zamfirescu-Pereira, Björn Hartmann, Aditya G. Parameswaran & Ian Arawjo · UIST 2024, 2024

A tool and user study for aligning LLM graders with human grades, and the origin of the term criteria drift.

A Survey of Techniques for Maximizing LLM PerformanceOpenAI · 46 min

The talk surveys prompt engineering, retrieval and fine-tuning. For each step, ask how you would version it and which eval would show that it helped.

DiscoverLearnRead beyondPapers & lecturesYour turn

Your turn

Three releases wait on the canary label. Run the golden set in batches, watch the paired interval tighten inside the unpaired one, and make each call only when the interval backs it. The typo fix is the hardest: a hold needs a narrow interval, and there is more than one way to narrow it.

Interactive lab

Release manager

Three prompt changes are running on the canary label. Both versions answer the same golden items, so compare them item by item. Ship when the paired 95% interval for new minus old lies above 0, roll back when it lies below 0, and hold when it lies entirely inside ±3 points. Intervals count from 30 items; the pool holds 500.

prompts/support_answer.promptMINOR

Ask the assistant to quote the policy section it relied on, and to hand over when the policy is silent.

Answer {{ question }} using only the policy below.
+ Quote the section you relied on, e.g. (§4.2).
+ If the policy does not cover it, say so and offer a human.
Policy: {{ policy }}
Answers per item:
Paired difference in pass rate with paired and unpaired 95% intervals, against the number of items0-2425-12500751210024Items evaluatedNew minus old, percentage points
● Zero and ±3 points● Unpaired 95% interval● Paired 95% interval● Paired difference

v1.3.2 pass rate

n/a

v1.4.0 pass rate

n/a

Paired difference, pts

n/a

Paired 95% interval

at 30 items

Unpaired 95% interval

at 30 items

Unpaired ÷ paired width

n/a

Items 0/500 · model calls 0 · correlation of item scores n/a · paired SE n/a

Challenge: Release managerCall three prompt releases correctly (ship, roll back or hold), with a 95% interval behind each call.+50 XP

Match · Change ↔ Release

Starting from 1.4.0, which release?

+20 XP
Rename the output field answer to reply
Add an optional confidence field to the output
Fix a typo in the instructions
Undo yesterday's release in production

Options

Match · Expression ↔ Meaning

Eval statistics

+20 XP
p^(1−p^)/n\sqrt{\hat p(1-\hat p)/n}
Var⁡(a)+Var⁡(b)−2Cov⁡(a,b)\operatorname{Var}(a)+\operatorname{Var}(b)-2\operatorname{Cov}(a,b)
(b−c)2/(b+c)(b-c)^2/(b+c)
(1.96 σd/h)2(1.96\,\sigma_d/h)^2

Options

Proof puzzle

Why pairing shrinks the error bar

+25 XP

Claim

For one randomly drawn item, let aa and bb be the old and new versions' scores. Show that Var⁡(b−a)=Var⁡(a)+Var⁡(b)−2Cov⁡(a,b)\operatorname{Var}(b-a)=\operatorname{Var}(a)+\operatorname{Var}(b)-2\operatorname{Cov}(a,b), so that scoring both versions on the same items shrinks the variance of the difference whenever the scores are positively correlated.

Tap lines in the order they should appear. Not every line belongs. Tap a line in your proof to send it back.

Your proof

  1. Pick the first line below.

Available lines

Prove it yourself

Worst-case error bars

+35 XP

Claim

Prove that the standard error p(1−p)/n\sqrt{p(1-p)/n} of a pass rate is largest at p=1/2p=1/2. Deduce that n≥(1.96/(2h))2n\ge\bigl(1.96/(2h)\bigr)^2 items guarantee a 95% half-width of at most hh, whatever the true pass rate.

Preview

Your typeset proof appears here.

Coding problems

Problem 19·Warm-up

Fingerprint a release

+20 XP

A release is the JSON object below. Serialise it canonically, with keys sorted at every level and no spaces (in Python, json.dumps(release, sort_keys=True, separators=(",", ":"))). Compute the SHA-256 of the UTF-8 bytes, read the first 8 hexadecimal digits as a base-16 integer, and give that integer in decimal.

{"template": "Answer {{ question }} using only {{ policy }}. Quote the section you relied on.", "model": "provider-model-2026-05-14", "params": {"temperature": 0.2, "top_p": 0.95, "max_tokens": 400, "seed": 7}, "tools": [{"name": "lookup_order", "parameters": {"order_id": "string"}}], "retrieval": {"index": "policies-2026-09", "k": 5}}

An exact integer (or a fraction like 7/12)

Problem 20·Standard

Paired beats unpaired

+35 XP

Old and new versions of a prompt were run on the same 600 golden items. Both passed 410 items, only the old version passed 38, only the new version passed 64, and both failed 88. Let di=snew,i−sold,i∈{−1,0,1}d_i=s_{\text{new},i}-s_{\text{old},i}\in\{-1,0,1\}. Compute the paired 95% interval dˉ±1.96 sd/n\bar d\pm1.96\,s_d/\sqrt n, where sds_d is the sample standard deviation of the did_i (dividing by n−1n-1). Give its lower end to 4 decimal places.

A number, rounded to 4 decimal places

Problem 21·Challenge

The winner's curse

+50 XP

You try m=20m=20 prompt variants that are, in truth, identical: each passes each of n=200n=200 dev items independently with probability 0.70.7. You keep the variant with the highest dev pass rate. Compute exactly, without simulation, the expected dev pass rate of the winner minus 0.70.7. Give 4 decimal places. Hint: for a random variable MM taking values in {0,1,…,n}\{0,1,\dots,n\}, E[M]=∑k=1nP(M≥k)\mathbb E[M]=\sum_{k=1}^{n}P(M\ge k), and the maximum of independent scores is below kk only if every score is.

A number, rounded to 4 decimal places

Key takeaways

  • A release is template, model snapshot, parameters, tools and retrieval together: version it immutably, fingerprint it, and log the fingerprint with every call.
  • Semantic versions name the change for people; movable labels such as prod make rollback a pointer change.
  • An eval suite combines real-trace data, code assertions, calibrated LLM judges and human review, and gates every merge.
  • A pass rate on nn items carries a 95% error bar of up to about ±1/n\pm1/\sqrt n. Compare versions on the same items, where pairing removes the difficulty they share.
  • Plan nn before you run, don't stop at the first lucky look, and keep a held-out set for the final decision.

Checkpoint

Prove it to the labyrinth

Answer every question to clear this chamber. First-try answers earn the most XP.

0/9
Question 1 of 9 +20 XP

Two prompt versions each have a standard error of 0.02 on the same 400 items, and their item scores have correlation 0.6. What is the standard error of the paired difference?

Question 2 of 9 +20 XP

A pilot run shows the per-item differences between two versions have standard deviation σd=0.4\sigma_d=0.4. How many items do you need for the 95% interval on the mean difference to have half-width at most 0.02? Round up.

Question 3 of 9 +20 XP

Why should a release pin a dated model snapshot rather than an alias such as latest?

Question 4 of 9 +20 XP

In McNemar's test for two versions scored pass or fail on the same items, which items carry the evidence about which version is better?

Question 5 of 9 +20 XP

You tried 40 prompt variants on a 200-item dev set and kept the best, which scored 84%. Which number estimates its production pass rate without bias?

Question 6 of 9 +20 XP

Item difficulty contributes variance Var⁡(x)=0.06\operatorname{Var}(x)=0.06 to item scores, and sampling noise contributes E[σi2]=0.12\mathbb E[\sigma_i^2]=0.12 per answer. With n=200n=200 items and K=4K=4 answers per item, what is the standard error of the pass rate?

Question 7 of 9 +20 XP

Production fetches a prompt from a registry at runtime by its prod label. What is the main risk to guard against?

Question 8 of 9 +20 XP

Which version has the highest precedence under Semantic Versioning 2.0.0?

Question 9 of 9 +20 XP

How should you decide whether an LLM judge can stand in for human reviewers on your task?

End of the chamber

Clear this chamber

+60 XPPrompt ReleaseEval SuiteLLM-as-a-JudgePaired Comparison