Skip to content
AriadneTechnology

The Middle Ring · Chamber 6 of 9

Prompting as Programming

System prompts, examples, reasoning and structured output: write prompts like interfaces, and make the model speak valid JSON.

55 min 60 XP + 12 questions + 1 challengeMathVideoPapersProofsCodeLab

In this chamber you will

  • Structure a prompt: roles, context, instructions, examples and output format
  • Use few-shot examples and chain-of-thought, and know when each one helps
  • Get structured output reliably, and explain how grammar-constrained decoding works
  • Diagnose prompt failures: ambiguity, conflicting instructions and position effects
DiscoverLearnRead beyondPapers & lecturesYour turn

Ask a model to label a review and you may get Sure! Here is the JSON: {'sentiment': 'mixed'} Hope that helps!. Every word is reasonable, and your parser rejects all of it. This chamber has two complementary fixes: write the prompt precisely enough that the model usually does what you mean, then constrain decoding so that the output always parses. The second fits in one line of maths: hide the tokens the grammar forbids, then sample as usual.

Spotted in the wild

α=LM⁡(S~t,θ)α~=m⁡(S~t)⊙αs~t+1∼Categorical⁡(α~)\begin{aligned} \boldsymbol\alpha&=\operatorname{LM}(\tilde S_t,\boldsymbol\theta)\\ \tilde{\boldsymbol\alpha}&=\operatorname{m}(\tilde S_t)\odot\boldsymbol\alpha\\ \tilde s_{t+1}&\sim\operatorname{Categorical}(\tilde{\boldsymbol\alpha}) \end{aligned}
Efficient Guided Generation for Large Language Models
DiscoverLearnRead beyondPapers & lecturesYour turn
Symbols for this chamber
  • V\mathcal V“script V”
    The vocabulary: every token the model can emit, including the end token.
  • p(v∣x<t)p(v\mid x_{<t})“p of v given x before t”
    The model's probability that the next token is vv, given the prompt and the tokens generated so far.
  • δ(q,v)\delta(q,v)“delta of q and v”
    The automaton's transition: the state reached from state qq by reading token vv, if any.
    qt=δ(qt−1,xt)q_{t}=\delta(q_{t-1},x_t)
  • A(q)A(q)“A of q”
    The allowed set: tokens that lead from state qq to a state from which the schema can still be completed.
  • m(q)m(q)“the mask at q”
    A 0/1 vector over the vocabulary with ones exactly on A(q)A(q).
    m(q)∈{0,1}∣V∣m(q)\in\{0,1\}^{|\mathcal V|}
  • ZtZ_t“Z t, the kept mass”
    The probability the model put on allowed tokens at step tt.
    Zt=∑v∈A(qt−1)p(v∣x<t)Z_t=\sum_{v\in A(q_{t-1})}p(v\mid x_{<t})
  • 1−Zt1-Z_t“one minus Z t, the removed mass”
    The probability the mask takes away at step tt.
  • p~(v∣x<t)\tilde p(v\mid x_{<t})“p tilde of v given x before t”
    The masked, renormalised distribution the constrained sampler actually draws from.
    p~(v∣x<t)=mv p(v∣x<t)/Zt\tilde p(v\mid x_{<t})=m_v\,p(v\mid x_{<t})/Z_t

Roles, templates and the system prompt

A chat model never receives "messages". It receives one token sequence, rendered by a chat template that marks each turn with special tokens (Chamber 1) and leaves the assistant turn open:

Text
<|system|>You label product reviews. Reply with JSON only.<|end|>
<|user|>Review: "Arrived late, but works perfectly."<|end|>
<|assistant|>

Markers differ by model family (ChatML's <|im_start|>, Llama 3's header tokens), and tokeniser libraries render them for you, for example Hugging Face's apply_chat_template. The system turn holds the operator's standing instructions, the user turn the request, and the assistant turn what the model writes. Instruction tuning (Chamber 3) teaches the model to give the system turn priority, but it is still just conditioning: the reply is sampled from p(y∣s,u)p(y\mid s,u). If you serve open weights, a missing or doubled special token in a hand-built template is a classic silent bug.

Quick check +20 XP

A chat API receives a system message and a user message. What does the model actually process?

Write the prompt like an interface

A function has a name, typed arguments, documented behaviour and a return type. A good prompt specifies the same things, for a reader who knows nothing you haven't written down:

PartQuestion it answersExample
TaskWhat should happen?Classify the sentiment of one product review
AudienceWho reads the output?A dashboard parser, not a person
ContextWhat does the model need to know?Reviews come from a hardware shop; sarcasm is common
ConstraintsWhat must never happen?No label outside the three allowed; no prose
Success criteriaHow would you grade it?Agrees with a careful human; mixed reviews get the dominant label
Output formatWhat exactly comes back?One JSON object with the single key sentiment

Separate instructions from data with delimiters such as XML-style tags, so the model can tell where a pasted document ends and you can refer to it by name. In code, a prompt is a template with variables, rendered per request; Chamber 7 versions and tests templates like any other code.

Python
TEMPLATE = """Classify the sentiment of the review in <review> tags.
Allowed labels: positive, negative, neutral. A mixed review gets its dominant label.
Reply with exactly one JSON object: {{"sentiment": "<label>"}}

<review>
{review}
</review>"""

prompt = TEMPLATE.format(review=user_text)

Few-shot examples

Brown et al. (2020) showed that a large enough model can pick up a task from a few demonstrations in its context, with no gradient step: in-context learning. Choose examples as you would a test set: cover the hard cases, vary length and wording so the model doesn't copy surface features, balance the labels, and mind the order, since models can favour the majority label and labels near the end of the prompt (Zhao et al., 2021). Min et al. (2022) found something surprising: replacing the gold labels in demonstrations with random ones barely hurt classification and multiple-choice accuracy. The demonstrations mainly taught the label space, the distribution of inputs and the format. That is no licence for wrong labels: larger models can follow deliberately flipped ones (Wei et al., 2023).

Quick check +20 XP

Min et al. (2022) replaced the gold labels in few-shot demonstrations with random labels on classification and multiple-choice tasks. What happened?

Reasoning, decomposition and prefill

Chain of thought asks for intermediate steps before the answer: Wei et al. (2022) used a few worked exemplars, and Kojima et al. (2022) a zero-shot cue, "Let's think step by step". Each generated token gets one forward pass (Chamber 2), so writing reasoning zz first buys computation the answer can condition on: p(y∣x)=∑zp(z∣x) p(y∣x,z)p(y\mid x)=\sum_z p(z\mid x)\,p(y\mid x,z). Wei et al. saw gains mainly in models of about 100B parameters or more, and most on multi-step problems; on simple lookups, reasoning only adds latency and tokens. Reasoning models are trained with reinforcement learning to write such chains themselves. For them, state goals and constraints plainly instead of prescribing steps, and test whether examples still help: DeepSeek-R1's authors report that few-shot prompting degraded its performance (DeepSeek-AI, 2025).

Decomposition splits a long task into a prompt chain (extract, then classify, then summarise), with code checking each intermediate result; each stage is easier to test and can use its own model and temperature. Prefilling writes the first tokens of the assistant turn yourself, such as {"sentiment": ", so the model starts inside the format. Raw templates always allow it; some hosted APIs do and others don't.

Structured output that always parses

ApproachGuaranteesUse when
Instructions and examplesNothing, though usually rightPrototyping
JSON modeJSON syntax, in any shapeNothing stronger is available
Schema-constrained decodingOutput matches your grammar or JSON SchemaYou need a fixed shape every time
Tool or function schemaArguments for a named functionThe output triggers an action (Chamber 8)
Validate and retryNothing per call; the loop ends valid or gives upAlways, as the last line of defence

Validation also catches what no grammar expresses, such as "the quoted span must appear in the source". If each attempt passes independently with probability pp, the number of attempts NN is geometric, P(N=k)=(1−p)k−1pP(N=k)=(1-p)^{k-1}p, so

E[N]=∑k≥1P(N≥k)=∑k≥1(1−p)k−1=1p.E[N]=\sum_{k\ge1}P(N\ge k)=\sum_{k\ge1}(1-p)^{k-1}=\frac1p.

At p=0.95p=0.95 retries are cheap; at p=0.1p=0.1 you pay ten calls on average, plus latency. Better to raise pp, or make it 1 with a mask.

Quick check +20 XP

A model's raw output passes your schema validator with probability 0.2 on each independent attempt. With unlimited retries, what is the expected number of calls?

Masking is conditioning

Compile the schema into a finite automaton with states qq and transitions δ(q,v)\delta(q,v). Before step tt, the allowed set A(qt−1)A(q_{t-1}) holds the tokens leading to states from which the schema can still be completed. Zero the rest and renormalise:

p~(v∣x<t)=mv p(v∣x<t)Zt,Zt=∑u∈A(qt−1)p(u∣x<t).\tilde p(v\mid x_{<t})=\frac{m_v\,p(v\mid x_{<t})}{Z_t},\qquad Z_t=\sum_{u\in A(q_{t-1})}p(u\mid x_{<t}).

For one step this is exact conditioning: if X∼pX\sim p, then P(X=v∣X∈A)=p(v)/p(A)P(X=v\mid X\in A)=p(v)/p(A) on AA and zero elsewhere. On logits, set masked entries to −∞-\infty (not 0, since e0=1e^0=1) and let the softmax renormalise. If every reachable state allows some token and can still reach acceptance, the sampler never gets stuck and can only stop in an accepting state, so the output always parses; you will prove this below. One token can span several grammar symbols, which is why Willard and Louf precompute, for each automaton state, the vocabulary tokens it can read.

Python
def masked_step(probs, allowed, rng):
    kept = sum(probs[v] for v in allowed)                        # Z_t
    weights = [probs[v] / kept if v in allowed else 0.0 for v in range(len(probs))]
    return rng.choices(range(len(probs)), weights)[0], 1 - kept  # token, removed mass

The kept masses measure how far the raw model was from compliance. Along a valid output, p(x)=∏tZt p~(xt∣x<t)p(x)=\prod_t Z_t\,\tilde p(x_t\mid x_{<t}), so if the schema forces a single path, P(valid)=∏tZtP(\text{valid})=\prod_tZ_t; with branches, P(valid)=Ep~[∏tZt]P(\text{valid})=\mathbb E_{\tilde p}\bigl[\prod_tZ_t\bigr]. But conditioning each step is not conditioning the whole string. The masked sampler produces each valid xx with probability p(x)/∏tZtp(x)/\prod_tZ_t, whereas p(x∣valid)=p(x)/P(valid)p(x\mid\text{valid})=p(x)/P(\text{valid}), and the two agree only when ∏tZt\prod_tZ_t is the same for every valid xx. Masking over-weights prefixes from which the model rarely finishes validly (Park et al., 2024). It guarantees syntax, not quality: describe the schema in the prompt too (some engines apply it only as a mask, so the model never sees it), keep schemas close to what the model would write anyway, and put any reasoning field before the answer field.

Quick check +20 XP

Only two tokens are allowed at this step, with model probabilities 0.12 and 0.28. After masking and renormalising, what is the probability of the second one?

Long contexts and caching

With long inputs, put the documents first and the question and instructions last; several providers' guides recommend this, some also repeating key instructions at both ends. Liu et al. (2024) showed why position matters: in multi-document question answering, accuracy was highest with the relevant document at the start or end of the context and dropped when it sat in the middle. Wrap each document in tags with its source, ask for the relevant quotes before the answer, and retrieve less but better (Chamber 8).

Prompt caching reuses the KV cache (Chamber 2) for a prefix the server has recently seen, cutting cost and time to first token. Hits need an exact prefix match, so order the prompt from most stable to least: system instructions, tool definitions, examples and long documents, then the per-request input. A timestamp at the top invalidates everything after it.

When a prompt fails

SymptomLikely causeFirst fix
Different answers to near-identical inputsAmbiguous task or success criteriaDefine the edge cases; add an example of each
A rule is ignoredConflicting or buried instructionsRemove the conflict; state the rule once, clearly
Outputs copy an example's wordingExamples too alikeDiversify them; say they are illustrations
Facts missed in a long inputPosition effectsMove key content to the edges; quote first
Format breaksFormat described, not enforcedPrefill, constrain decoding, validate

Debug a prompt like code: read the fully rendered prompt (many bugs live in the template, not the wording), keep a small set of hard inputs, change one thing at a time, and study failures rather than averages. Lower the temperature for extraction (Chamber 4), and remember Chamber 5: one good sample proves little. Chamber 7 turns this loop into versioned prompts and evals with error bars.

A prompt is not a security boundary. Any text the model reads, such as a retrieved web page or an uploaded file, can carry instructions the model may follow: prompt injection. Delimiters make accidents less likely, not attacks impossible. Chamber 8 covers the defences: least privilege for tools, keeping trusted and untrusted text apart, and checking actions outside the model.

DiscoverLearnRead beyondPapers & lecturesYour turn

Read beyond

Paper · free online · ~30 min

Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?

Min, Lyu, Holtzman, Artetxe, Lewis, Hajishirzi & Zettlemoyer · Sections 4 and 5: random labels, then what demonstrations really teach

Find the random-label experiment and list the three aspects of demonstrations the authors say drive performance.

Paper · free online · ~25 min

Lost in the Middle: How Language Models Use Long Contexts

Liu, Lin, Hewitt, Paranjape, Bevilacqua, Petroni & Liang · Section 2: multi-document question answering

Sketch the U-shaped curve and compare mid-context accuracy with the closed-book baseline.

Paper · free online · ~20 min

Large Language Models are Zero-Shot Reasoners

Kojima, Gu, Reid, Matsuo & Iwasawa · Section 3: Zero-shot-CoT and its two-stage prompting

Write out both stages of the prompt and explain why the second is needed.

Article · free online · ~30 min

Prompt Engineering

Lilian Weng · Few-shot example selection and ordering, then chain of thought

Compare her tips on choosing and ordering examples with this chamber's checklist.

Tool · free online · ~15 min

GBNF grammars in llama.cpp

llama.cpp contributors · Grammar syntax and the JSON Schema conversion

Write a grammar for the lab's schema with exactly three allowed labels, and note what the guide says about whether the model ever sees a JSON schema.

DiscoverLearnRead beyondPapers & lecturesYour turn

Read the equation in context

Efficient Guided Generation for Large Language ModelsBrandon T. Willard & Rémi Louf · arXiv preprint, 2023

Willard and Louf treat guided generation as ordinary sampling with a Boolean mask that restricts the support of the next-token distribution. The mask depends on the whole prefix through the state of a finite automaton built from a regular expression, and they extend the idea to context-free grammars with an LALR(1) parser. Their contribution is speed: an index from automaton states to vocabulary tokens turns each step's mask into a lookup instead of a scan over the vocabulary. Their library, Outlines, implements it, and other open-source engines and grammar libraries apply the same masked step.

Decode the paper · Section 2.2, Guiding generation (unnumbered display)

Efficient Guided Generation for Large Language Models

Brandon T. Willard & Rémi Louf · arXiv preprint, 2023

+25 XP
α=LM⁡(S~t,θ)α~=m⁡(S~t)⊙αs~t+1∼Categorical⁡(α~)\begin{aligned}\boldsymbol\alpha&=\operatorname{LM}(\tilde S_t,\boldsymbol\theta)\\\tilde{\boldsymbol\alpha}&=\operatorname{m}(\tilde S_t)\odot\boldsymbol\alpha\\\tilde s_{t+1}&\sim\operatorname{Categorical}(\tilde{\boldsymbol\alpha})\end{aligned}

Willard and Louf write the guided step as an ordinary sampling step with a Boolean mask applied first. The mask depends on the whole prefix, through the state of a finite automaton built from a regular expression or grammar. Read α\boldsymbol\alpha as nonnegative next-token weights: the paper calls them logits, but a 0/1 product only removes a token from probabilities (or exponentiated logits); on raw logits you would set masked entries to −∞-\infty.

S~t\tilde S_t
θ\boldsymbol\theta
α\boldsymbol\alpha
m⁡(S~t)\operatorname{m}(\tilde S_t)
⊙\odot
s~t+1\tilde s_{t+1}

Options

Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le & Denny Zhou · NeurIPS, 2022

Read Figure 1 for standard against chain-of-thought exemplars, then Section 3.2: the gains appear only at scale and grow with problem difficulty. Then ask what a mask does to a chain of thought: if the schema puts the answer field first, the model must answer before it reasons.

Prompting 101 | Code w/ ClaudeAnthropic · 25 min

The talk uses one provider's models, but its habits (structure, context, examples, prefilled output, iterating on real failures) carry over to any model.

DiscoverLearnRead beyondPapers & lecturesYour turn

Your turn

The toy model below wants to chat, use single quotes, invent a "mixed" label and add pleasantries. With the mask on, produce all three valid labels and find the step where the mask cuts the most probability; then run 200 samples each way and compare with the exact rate.

Interactive lab

Make the model speak JSON

A toy model must answer {"sentiment": …} with positive, negative or neutral. It also likes chatty preambles, single quotes, a "mixed" label and trailing pleasantries. Pick tokens or sample them, with and without the grammar mask.

(the assistant turn starts here)

Step 1 · automaton expects {

Next-token distribution. Teal tokens are allowed; tap one or sample. Rose ones are masked.

0.45 → 1.00
0.40 → 0.00
0.15 → 0.00

Mask removes

55.0%

Kept mass Z

0.450

Valid labels under the mask: positivenegativeneutral· Largest cut found: not yet

Challenge: Schema enforcerGenerate three valid outputs with different labels under a grammar mask, and find the step where the mask removes the most probability.+40 XP

Match · Approach ↔ What it guarantees

Ways to get structured output

+20 XP
JSON mode
Schema-constrained decoding
Tool or function schema
Validate and retry

Options

Match · Expression ↔ Meaning

Masking in symbols

+20 XP
Zt=∑v∈A(qt−1)p(v∣x<t)Z_t=\sum_{v\in A(q_{t-1})}p(v\mid x_{<t})
1−Zt1-Z_t
p(v∣x<t)/Ztp(v\mid x_{<t})/Z_t
∏tZt\prod_t Z_t

Options

Proof puzzle

Masking then renormalising is conditioning

+25 XP

Claim

Let X∼pX\sim p over V\mathcal V and let A⊆VA\subseteq\mathcal V have p(A)>0p(A)>0, with mask mv=1m_v=1 for v∈Av\in A and 00 otherwise. Show that p~(v)=mv p(v)/∑umu p(u)\tilde p(v)=m_v\,p(v)/\sum_u m_u\,p(u) equals P(X=v∣X∈A)P(X=v\mid X\in A) for every vv.

Tap lines in the order they should appear. Not every line belongs. Tap a line in your proof to send it back.

Your proof

  1. Pick the first line below.

Available lines

Prove it yourself

Masked sampling always finishes validly

+35 XP

Claim

Call an automaton state live if an accepting state can be reached from it. Suppose the start state is live, the automaton has no cycles, the mask at each live state qq allows exactly the tokens vv with δ(q,v)\delta(q,v) live, plus the end token when qq is accepting, and the sampler stops only on the end token. If the model gives every token positive probability, prove that masked sampling never gets stuck and stops with a string the automaton accepts.

Preview

Your typeset proof appears here.

Coding problems

Problem 16·Warm-up

Validate and retry

+20 XP

A model returns schema-valid JSON on each call independently with probability p=0.2p=0.2. Your client validates every response and retries on failure, making at most 5 calls per request. What is the expected number of calls per request? Give 4 decimal places.

A number, rounded to 4 decimal places

Problem 17·Standard

How often does the raw model comply?

+35 XP

Tokens are the digits 0–9 (ids 0–9) and an end token (id 10). A toy bigram model gives p(j∣i)=w(i,j)/∑k=010w(i,k)p(j\mid i)=w(i,j)/\sum_{k=0}^{10}w(i,k) with w(i,j)=1+((3i+5j+2) mod 7)w(i,j)=1+((3i+5j+2)\bmod 7), where ii is the previous token's id and generation starts with i=10i=10. The schema asks for a three-digit integer without a leading zero: exactly three digits, the first nonzero, then the end token. Compute the probability that unconstrained sampling produces a schema-valid output. Give 6 decimal places.

A number, rounded to 6 decimal places

Problem 18·Challenge

The price of local renormalisation

+50 XP

Tokens are the digits 0–9 (ids 0–9) and an end token (id 10). A toy bigram model gives p(j∣i)=w(i,j)/∑k=010w(i,k)p(j\mid i)=w(i,j)/\sum_{k=0}^{10}w(i,k) with w(i,j)=1+((3i+5j+2) mod 7)w(i,j)=1+((3i+5j+2)\bmod 7), where ii is the previous token's id and generation starts with i=10i=10. The schema asks for a three-digit integer without a leading zero: exactly three digits, the first nonzero, then the end token. A masked sampler keeps only the allowed tokens at each step (1–9 first, then 0–9 twice, then only the end token) and renormalises. Compute the total variation distance 12∑x∣p~(x)−p(x∣x valid)∣\tfrac12\sum_x|\tilde p(x)-p(x\mid x\text{ valid})| between the masked sampler's distribution over the 900 valid strings and the model's own distribution conditioned on validity. Give 6 decimal places.

A number, rounded to 6 decimal places

Key takeaways

  • A prompt is one token sequence: write it like an interface, with task, context, constraints, success criteria and output format, and delimit data from instructions.
  • Few-shot examples mostly teach format and label space; chain of thought helps multi-step problems on capable models. Test both rather than assuming.
  • Masking then renormalising is exact conditioning at each step, and a live automaton guarantees output that parses. Validate anyway.
  • Step-by-step masking is not sequence-level conditioning: it can shift which valid outputs the model prefers.
  • Put long documents first and the question last, keep the stable prefix first for caching, and debug from the rendered prompt and real failures.

Checkpoint

Prove it to the labyrinth

Answer every question to clear this chamber. First-try answers earn the most XP.

0/8
Question 1 of 8 +20 XP

A schema forces one fixed sequence of four tokens. Unconstrained, the model puts 0.9, 0.8, 0.95 and 0.5 on the required token at each step. What is the probability that unconstrained sampling produces the valid output?

Question 2 of 8 +20 XP

What does a JSON mode guarantee on its own?

Question 3 of 8 +20 XP

To mask a token at the logit level, before the softmax, you set its logit to:

Question 4 of 8 +20 XP

Wei et al. (2022) found that chain-of-thought exemplars improved accuracy:

Question 5 of 8 +20 XP

Liu et al. (2024) moved the one relevant document around a long multi-document prompt. Accuracy was usually:

Question 6 of 8 +20 XP

Which layout gets the most prompt-cache hits across many requests?

Question 7 of 8 +20 XP

Does token-by-token masking sample valid strings in proportion to the model's own probabilities p(x∣x valid)p(x\mid x\text{ valid})?

Question 8 of 8 +20 XP

You move a prompt to a reasoning model that thinks before it answers. What is usually the better first change?

End of the chamber

Clear this chamber

+60 XPSystem PromptFew-Shot PromptingChain of ThoughtConstrained Decoding