Ask a model to label a review and you may get Sure! Here is the JSON: {'sentiment': 'mixed'} Hope that helps!. Every word is reasonable, and your parser rejects all of it. This chamber has two complementary fixes: write the prompt precisely enough that the model usually does what you mean, then constrain decoding so that the output always parses. The second fits in one line of maths: hide the tokens the grammar forbids, then sample as usual.
Spotted in the wild
- “script V”The vocabulary: every token the model can emit, including the end token.
- “p of v given x before t”The model's probability that the next token is , given the prompt and the tokens generated so far.
- “delta of q and v”The automaton's transition: the state reached from state by reading token , if any.
- “A of q”The allowed set: tokens that lead from state to a state from which the schema can still be completed.
- “the mask at q”A 0/1 vector over the vocabulary with ones exactly on .
- “Z t, the kept mass”The probability the model put on allowed tokens at step .
- “one minus Z t, the removed mass”The probability the mask takes away at step .
- “p tilde of v given x before t”The masked, renormalised distribution the constrained sampler actually draws from.
| Symbol | Say it | Meaning | LaTeX |
|---|---|---|---|
| “script V” | The vocabulary: every token the model can emit, including the end token. | ||
| “p of v given x before t” | The model's probability that the next token is , given the prompt and the tokens generated so far. | ||
| “delta of q and v” | The automaton's transition: the state reached from state by reading token , if any. | ||
| “A of q” | The allowed set: tokens that lead from state to a state from which the schema can still be completed. | ||
| “the mask at q” | A 0/1 vector over the vocabulary with ones exactly on . | ||
| “Z t, the kept mass” | The probability the model put on allowed tokens at step . | ||
| “one minus Z t, the removed mass” | The probability the mask takes away at step . | ||
| “p tilde of v given x before t” | The masked, renormalised distribution the constrained sampler actually draws from. |
Roles, templates and the system prompt
A chat model never receives "messages". It receives one token sequence, rendered by a chat template that marks each turn with special tokens (Chamber 1) and leaves the assistant turn open:
<|system|>You label product reviews. Reply with JSON only.<|end|>
<|user|>Review: "Arrived late, but works perfectly."<|end|>
<|assistant|>
Markers differ by model family (ChatML's <|im_start|>, Llama 3's header tokens), and tokeniser libraries render them for you, for example Hugging Face's apply_chat_template. The system turn holds the operator's standing instructions, the user turn the request, and the assistant turn what the model writes. Instruction tuning (Chamber 3) teaches the model to give the system turn priority, but it is still just conditioning: the reply is sampled from . If you serve open weights, a missing or doubled special token in a hand-built template is a classic silent bug.
A chat API receives a system message and a user message. What does the model actually process?
Write the prompt like an interface
A function has a name, typed arguments, documented behaviour and a return type. A good prompt specifies the same things, for a reader who knows nothing you haven't written down:
| Part | Question it answers | Example |
|---|---|---|
| Task | What should happen? | Classify the sentiment of one product review |
| Audience | Who reads the output? | A dashboard parser, not a person |
| Context | What does the model need to know? | Reviews come from a hardware shop; sarcasm is common |
| Constraints | What must never happen? | No label outside the three allowed; no prose |
| Success criteria | How would you grade it? | Agrees with a careful human; mixed reviews get the dominant label |
| Output format | What exactly comes back? | One JSON object with the single key sentiment |
Separate instructions from data with delimiters such as XML-style tags, so the model can tell where a pasted document ends and you can refer to it by name. In code, a prompt is a template with variables, rendered per request; Chamber 7 versions and tests templates like any other code.
TEMPLATE = """Classify the sentiment of the review in <review> tags.
Allowed labels: positive, negative, neutral. A mixed review gets its dominant label.
Reply with exactly one JSON object: {{"sentiment": "<label>"}}
<review>
{review}
</review>"""
prompt = TEMPLATE.format(review=user_text)
Few-shot examples
Brown et al. (2020) showed that a large enough model can pick up a task from a few demonstrations in its context, with no gradient step: in-context learning. Choose examples as you would a test set: cover the hard cases, vary length and wording so the model doesn't copy surface features, balance the labels, and mind the order, since models can favour the majority label and labels near the end of the prompt (Zhao et al., 2021). Min et al. (2022) found something surprising: replacing the gold labels in demonstrations with random ones barely hurt classification and multiple-choice accuracy. The demonstrations mainly taught the label space, the distribution of inputs and the format. That is no licence for wrong labels: larger models can follow deliberately flipped ones (Wei et al., 2023).
Min et al. (2022) replaced the gold labels in few-shot demonstrations with random labels on classification and multiple-choice tasks. What happened?
Reasoning, decomposition and prefill
Chain of thought asks for intermediate steps before the answer: Wei et al. (2022) used a few worked exemplars, and Kojima et al. (2022) a zero-shot cue, "Let's think step by step". Each generated token gets one forward pass (Chamber 2), so writing reasoning first buys computation the answer can condition on: . Wei et al. saw gains mainly in models of about 100B parameters or more, and most on multi-step problems; on simple lookups, reasoning only adds latency and tokens. Reasoning models are trained with reinforcement learning to write such chains themselves. For them, state goals and constraints plainly instead of prescribing steps, and test whether examples still help: DeepSeek-R1's authors report that few-shot prompting degraded its performance (DeepSeek-AI, 2025).
Decomposition splits a long task into a prompt chain (extract, then classify, then summarise), with code checking each intermediate result; each stage is easier to test and can use its own model and temperature. Prefilling writes the first tokens of the assistant turn yourself, such as {"sentiment": ", so the model starts inside the format. Raw templates always allow it; some hosted APIs do and others don't.
Structured output that always parses
| Approach | Guarantees | Use when |
|---|---|---|
| Instructions and examples | Nothing, though usually right | Prototyping |
| JSON mode | JSON syntax, in any shape | Nothing stronger is available |
| Schema-constrained decoding | Output matches your grammar or JSON Schema | You need a fixed shape every time |
| Tool or function schema | Arguments for a named function | The output triggers an action (Chamber 8) |
| Validate and retry | Nothing per call; the loop ends valid or gives up | Always, as the last line of defence |
Validation also catches what no grammar expresses, such as "the quoted span must appear in the source". If each attempt passes independently with probability , the number of attempts is geometric, , so
At retries are cheap; at you pay ten calls on average, plus latency. Better to raise , or make it 1 with a mask.
A model's raw output passes your schema validator with probability 0.2 on each independent attempt. With unlimited retries, what is the expected number of calls?
Masking is conditioning
Compile the schema into a finite automaton with states and transitions . Before step , the allowed set holds the tokens leading to states from which the schema can still be completed. Zero the rest and renormalise:
For one step this is exact conditioning: if , then on and zero elsewhere. On logits, set masked entries to (not 0, since ) and let the softmax renormalise. If every reachable state allows some token and can still reach acceptance, the sampler never gets stuck and can only stop in an accepting state, so the output always parses; you will prove this below. One token can span several grammar symbols, which is why Willard and Louf precompute, for each automaton state, the vocabulary tokens it can read.
def masked_step(probs, allowed, rng):
kept = sum(probs[v] for v in allowed) # Z_t
weights = [probs[v] / kept if v in allowed else 0.0 for v in range(len(probs))]
return rng.choices(range(len(probs)), weights)[0], 1 - kept # token, removed mass
The kept masses measure how far the raw model was from compliance. Along a valid output, , so if the schema forces a single path, ; with branches, . But conditioning each step is not conditioning the whole string. The masked sampler produces each valid with probability , whereas , and the two agree only when is the same for every valid . Masking over-weights prefixes from which the model rarely finishes validly (Park et al., 2024). It guarantees syntax, not quality: describe the schema in the prompt too (some engines apply it only as a mask, so the model never sees it), keep schemas close to what the model would write anyway, and put any reasoning field before the answer field.
Only two tokens are allowed at this step, with model probabilities 0.12 and 0.28. After masking and renormalising, what is the probability of the second one?
Long contexts and caching
With long inputs, put the documents first and the question and instructions last; several providers' guides recommend this, some also repeating key instructions at both ends. Liu et al. (2024) showed why position matters: in multi-document question answering, accuracy was highest with the relevant document at the start or end of the context and dropped when it sat in the middle. Wrap each document in tags with its source, ask for the relevant quotes before the answer, and retrieve less but better (Chamber 8).
Prompt caching reuses the KV cache (Chamber 2) for a prefix the server has recently seen, cutting cost and time to first token. Hits need an exact prefix match, so order the prompt from most stable to least: system instructions, tool definitions, examples and long documents, then the per-request input. A timestamp at the top invalidates everything after it.
When a prompt fails
| Symptom | Likely cause | First fix |
|---|---|---|
| Different answers to near-identical inputs | Ambiguous task or success criteria | Define the edge cases; add an example of each |
| A rule is ignored | Conflicting or buried instructions | Remove the conflict; state the rule once, clearly |
| Outputs copy an example's wording | Examples too alike | Diversify them; say they are illustrations |
| Facts missed in a long input | Position effects | Move key content to the edges; quote first |
| Format breaks | Format described, not enforced | Prefill, constrain decoding, validate |
Debug a prompt like code: read the fully rendered prompt (many bugs live in the template, not the wording), keep a small set of hard inputs, change one thing at a time, and study failures rather than averages. Lower the temperature for extraction (Chamber 4), and remember Chamber 5: one good sample proves little. Chamber 7 turns this loop into versioned prompts and evals with error bars.
A prompt is not a security boundary. Any text the model reads, such as a retrieved web page or an uploaded file, can carry instructions the model may follow: prompt injection. Delimiters make accidents less likely, not attacks impossible. Chamber 8 covers the defences: least privilege for tools, keeping trusted and untrusted text apart, and checking actions outside the model.
Read beyond
Paper · free online · ~30 min
Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?Min, Lyu, Holtzman, Artetxe, Lewis, Hajishirzi & Zettlemoyer · Sections 4 and 5: random labels, then what demonstrations really teach
Find the random-label experiment and list the three aspects of demonstrations the authors say drive performance.
Paper · free online · ~25 min
Lost in the Middle: How Language Models Use Long ContextsLiu, Lin, Hewitt, Paranjape, Bevilacqua, Petroni & Liang · Section 2: multi-document question answering
Sketch the U-shaped curve and compare mid-context accuracy with the closed-book baseline.
Paper · free online · ~20 min
Large Language Models are Zero-Shot ReasonersKojima, Gu, Reid, Matsuo & Iwasawa · Section 3: Zero-shot-CoT and its two-stage prompting
Write out both stages of the prompt and explain why the second is needed.
Article · free online · ~30 min
Prompt EngineeringLilian Weng · Few-shot example selection and ordering, then chain of thought
Compare her tips on choosing and ordering examples with this chamber's checklist.
Tool · free online · ~15 min
GBNF grammars in llama.cppllama.cpp contributors · Grammar syntax and the JSON Schema conversion
Write a grammar for the lab's schema with exactly three allowed labels, and note what the guide says about whether the model ever sees a JSON schema.
Read the equation in context
Efficient Guided Generation for Large Language ModelsBrandon T. Willard & Rémi Louf · arXiv preprint, 2023Willard and Louf treat guided generation as ordinary sampling with a Boolean mask that restricts the support of the next-token distribution. The mask depends on the whole prefix through the state of a finite automaton built from a regular expression, and they extend the idea to context-free grammars with an LALR(1) parser. Their contribution is speed: an index from automaton states to vocabulary tokens turns each step's mask into a lookup instead of a scan over the vocabulary. Their library, Outlines, implements it, and other open-source engines and grammar libraries apply the same masked step.
Decode the paper · Section 2.2, Guiding generation (unnumbered display)
Efficient Guided Generation for Large Language ModelsBrandon T. Willard & Rémi Louf · arXiv preprint, 2023
Willard and Louf write the guided step as an ordinary sampling step with a Boolean mask applied first. The mask depends on the whole prefix, through the state of a finite automaton built from a regular expression or grammar. Read as nonnegative next-token weights: the paper calls them logits, but a 0/1 product only removes a token from probabilities (or exponentiated logits); on raw logits you would set masked entries to .
Options
Read Figure 1 for standard against chain-of-thought exemplars, then Section 3.2: the gains appear only at scale and grow with problem difficulty. Then ask what a mask does to a chain of thought: if the schema puts the answer field first, the model must answer before it reasons.
The talk uses one provider's models, but its habits (structure, context, examples, prefilled output, iterating on real failures) carry over to any model.
Your turn
The toy model below wants to chat, use single quotes, invent a "mixed" label and add pleasantries. With the mask on, produce all three valid labels and find the step where the mask cuts the most probability; then run 200 samples each way and compare with the exact rate.
Interactive lab
Make the model speak JSON
(the assistant turn starts here)
Step 1 · automaton expects {
Next-token distribution. Teal tokens are allowed; tap one or sample. Rose ones are masked.
Mask removes
55.0%
Kept mass Z
0.450
Valid labels under the mask: positivenegativeneutral· Largest cut found: not yet
Match · Approach ↔ What it guarantees
Ways to get structured output
Options
Match · Expression ↔ Meaning
Masking in symbols
Options
Proof puzzle
Masking then renormalising is conditioning
Claim
Let over and let have , with mask for and otherwise. Show that equals for every .
Tap lines in the order they should appear. Not every line belongs. Tap a line in your proof to send it back.
Your proof
- Pick the first line below.
Available lines
Prove it yourself
Masked sampling always finishes validly
Claim
Call an automaton state live if an accepting state can be reached from it. Suppose the start state is live, the automaton has no cycles, the mask at each live state allows exactly the tokens with live, plus the end token when is accepting, and the sampler stops only on the end token. If the model gives every token positive probability, prove that masked sampling never gets stuck and stops with a string the automaton accepts.
Your typeset proof appears here.
Coding problems
Problem 16·Warm-up
Validate and retry
A model returns schema-valid JSON on each call independently with probability . Your client validates every response and retries on failure, making at most 5 calls per request. What is the expected number of calls per request? Give 4 decimal places.
Problem 17·Standard
How often does the raw model comply?
Tokens are the digits 0–9 (ids 0–9) and an end token (id 10). A toy bigram model gives with , where is the previous token's id and generation starts with . The schema asks for a three-digit integer without a leading zero: exactly three digits, the first nonzero, then the end token. Compute the probability that unconstrained sampling produces a schema-valid output. Give 6 decimal places.
Problem 18·Challenge
The price of local renormalisation
Tokens are the digits 0–9 (ids 0–9) and an end token (id 10). A toy bigram model gives with , where is the previous token's id and generation starts with . The schema asks for a three-digit integer without a leading zero: exactly three digits, the first nonzero, then the end token. A masked sampler keeps only the allowed tokens at each step (1–9 first, then 0–9 twice, then only the end token) and renormalises. Compute the total variation distance between the masked sampler's distribution over the 900 valid strings and the model's own distribution conditioned on validity. Give 6 decimal places.
Key takeaways
- A prompt is one token sequence: write it like an interface, with task, context, constraints, success criteria and output format, and delimit data from instructions.
- Few-shot examples mostly teach format and label space; chain of thought helps multi-step problems on capable models. Test both rather than assuming.
- Masking then renormalising is exact conditioning at each step, and a live automaton guarantees output that parses. Validate anyway.
- Step-by-step masking is not sequence-level conditioning: it can shift which valid outputs the model prefers.
- Put long documents first and the question last, keep the stable prefix first for caching, and debug from the rendered prompt and real failures.
Checkpoint
Prove it to the labyrinth
Answer every question to clear this chamber. First-try answers earn the most XP.
A schema forces one fixed sequence of four tokens. Unconstrained, the model puts 0.9, 0.8, 0.95 and 0.5 on the required token at each step. What is the probability that unconstrained sampling produces the valid output?
What does a JSON mode guarantee on its own?
To mask a token at the logit level, before the softmax, you set its logit to:
Wei et al. (2022) found that chain-of-thought exemplars improved accuracy:
Liu et al. (2024) moved the one relevant document around a long multi-document prompt. Accuracy was usually:
Which layout gets the most prompt-cache hits across many requests?
Does token-by-token masking sample valid strings in proportion to the model's own probabilities ?
You move a prompt to a reasoning model that thinks before it answers. What is usually the better first change?
End of the chamber
Clear this chamber
- Questions in this chamber (0/12 solved)Next unsolved
- Bonus: Schema enforcer (+40 XP)
- Bonus: Problem 16: Validate and retry (+20 XP)
- Bonus: Problem 17: How often does the raw model comply? (+35 XP)
- Bonus: Problem 18: The price of local renormalisation (+50 XP)
- Bonus: Proof: Masking then renormalising is conditioning (+25 XP)
- Bonus: Proof: Masked sampling always finishes validly (+35 XP)
- Bonus: Decode the paper (+25 XP)
- Bonus: Match: Ways to get structured output (+20 XP)
- Bonus: Match: Masking in symbols (+20 XP)