In 2021 DeepMind trained Gopher: 280 billion parameters on 300 billion tokens. A few months later the same team spent the same compute on Chinchilla, with a quarter of the parameters and more than four times the data, and Chinchilla beat Gopher across a wide range of tasks. The difference was arithmetic, not architecture: how to split a fixed budget between a bigger model and more text. Their answer rests on a five-constant formula fitted to hundreds of training runs.
Spotted in the wild
This chamber follows a model from that budget to an assistant: pretraining and its compute, scaling laws, then instruction tuning, preference optimisation and cheap fine-tuning with LoRA.
- “N”Number of parameters in the model.
- “D”Number of training tokens.
- “C is about six N D”Training compute in FLOPs: about per token forward and backward.
- “L hat of N and D”Predicted final pretraining loss, in nats per token.
- “E, A, B”Constants of the Chinchilla fit: irreducible loss and term scales and .
- “alpha, beta (scaling law)”Exponents of the model and data terms, fitted as and .
- “N opt of C”The model size that minimises the predicted loss at budget .
- “pi theta of y given x”The policy: the model being tuned, as a distribution over responses to prompt .
- “pi ref”The frozen reference model, usually the supervised fine-tuned model.
- “r phi of x and y”A learned reward model's scalar score for response to prompt .
- “y w, y l”The preferred (winning) and dispreferred (losing) responses in a comparison.
- “beta (RLHF and DPO)”Strength of the KL penalty to the reference model. Not the scaling-law exponent.
- “sigma of z”The logistic sigmoid .
- “delta W equals B A”LoRA's low-rank update, with and (a matrix, not the Chinchilla constant).
| Symbol | Say it | Meaning | LaTeX |
|---|---|---|---|
| “N” | Number of parameters in the model. | ||
| “D” | Number of training tokens. | ||
| “C is about six N D” | Training compute in FLOPs: about per token forward and backward. | ||
| “L hat of N and D” | Predicted final pretraining loss, in nats per token. | ||
| “E, A, B” | Constants of the Chinchilla fit: irreducible loss and term scales and . | ||
| “alpha, beta (scaling law)” | Exponents of the model and data terms, fitted as and . | ||
| “N opt of C” | The model size that minimises the predicted loss at budget . | ||
| “pi theta of y given x” | The policy: the model being tuned, as a distribution over responses to prompt . | ||
| “pi ref” | The frozen reference model, usually the supervised fine-tuned model. | ||
| “r phi of x and y” | A learned reward model's scalar score for response to prompt . | ||
| “y w, y l” | The preferred (winning) and dispreferred (losing) responses in a comparison. | ||
| “beta (RLHF and DPO)” | Strength of the KL penalty to the reference model. Not the scaling-law exponent. | ||
| “sigma of z” | The logistic sigmoid . | ||
| “delta W equals B A” | LoRA's low-rank update, with and (a matrix, not the Chinchilla constant). |
Pretraining: one objective, trillions of tokens
Pretraining is the previous chamber's next-token prediction run at enormous scale: minimise over a corpus of trillions of tokens. Most of it starts as web crawl, joined by code, books and papers. Raw crawl is mostly unusable, so the data pipeline does as much work as the optimiser:
- Extraction and language identification: pull the text out of HTML and keep the languages you want.
- Quality filtering: heuristic rules (length, symbol ratios, repeated lines) plus classifiers that score pages for quality or educational value.
- Deduplication: drop exact and near-duplicate documents, commonly with MinHash, so repeated text is not memorised and over-weighted.
- Hygiene: remove personal data and toxic content, and decontaminate by deleting benchmark test items so later evals stay honest.
FineWeb, an open dataset built this way, keeps 15 trillion tokens from 96 Common Crawl snapshots. The result of pretraining is a base model: it continues documents. Ask it a question and it may answer, or it may write three more questions, because a page of questions is plausible text.
Counting compute: C ≈ 6ND
In the forward pass each parameter takes part in one multiply and one add per token: about FLOPs. The backward pass computes gradients for both activations and weights, roughly twice that, . So training on tokens costs
Kaplan et al. (2020) note that attention adds a context-length term, small while the model width exceeds about a twelfth of the context. Check the rule on a public number: Llama 3's 405B model saw 15.6T tokens, and , the figure its report gives. Inference runs only the forward pass, about FLOPs per generated token. Remember that; it returns below.
A model with parameters is trained on tokens. Using , how many training FLOPs is that, in units of ?
Scaling laws and the compute-optimal split
Kaplan et al. found that loss falls as a power law in , and , with some trends spanning more than seven orders of magnitude, and advised spending most extra compute on a larger model (). Hoffmann et al. trained over 400 models, matched each run's learning-rate schedule to its length, and found instead that parameters and tokens should grow roughly equally. Their third method fits the Glimpse formula with , , , and . Here plays the entropy of text, and the other two terms are the penalties for a finite model and for finite data.
Fix the budget . Each model size then buys tokens, and these pairs trace an iso-compute curve (the paper's IsoFLOP profile). Find its lowest point by substitution. Put into the loss, differentiate in , set the result to zero and multiply by :
At the optimum, times the model term equals times the data term. Solving for gives the paper's Eq. (4):
With the fitted exponents, and .
With the Chinchilla fit (, ), the compute-optimal size grows as . By what factor does grow when the budget grows tenfold?
Training past the optimum
Chinchilla minimises training compute. A deployed model also pays about FLOPs for every token it generates, for as long as it serves. If it will generate trillions of tokens, a smaller model trained on more data reaches the same loss with a lower lifetime bill, as Sardana et al. (2024) work out with the Chinchilla fit. That is why open-weight families such as Llama 3 and Qwen2.5 report pretraining sets of 15–18 trillion tokens even for models of a few billion parameters: hundreds to thousands of tokens per parameter. Loss keeps falling out there, only slowly. The last coding problem puts numbers on the trade.
From base model to assistant
Supervised fine-tuning (SFT), or instruction tuning, continues training on prompt–response pairs: human-written demonstrations, curated datasets or a stronger model's outputs. Each conversation is rendered in the model's chat template, special tokens that mark roles and turns. An illustrative one:
<|system|>You are a concise assistant.<|end|>
<|user|>Name a prime between 10 and 15.<|end|>
<|assistant|>11 (or 13).<|end|>
The loss is ordinary cross-entropy, counted only on the assistant's tokens. Templates differ between model families, so render with the model's own (for example apply_chat_template in Hugging Face Transformers, or the template your inference server loads with the weights). A wrong template is a silent bug: the model still answers, just worse. SFT mostly teaches format and behaviour: LIMA got a strong assistant from 1,000 well-chosen examples.
Learning from preferences: RLHF
Many qualities are easier to judge than to write. RLHF collects comparisons: for a prompt , a labeller marks response as better than . A reward model is fitted with the Bradley–Terry likelihood
minimising : logistic regression on reward differences. Only differences enter, so adding a constant to every reward for a prompt changes nothing (you will prove it). The policy is then optimised, classically with PPO, to
The KL penalty tethers the policy to the reference (the SFT model), where the reward model was trained and can be trusted. Without it the policy finds outputs the reward model overrates: reward hacking. In InstructGPT, labellers preferred a 1.3B model trained this way to the 175B GPT-3.
A reward model scores a preferred response and a rejected one . What probability does the Bradley–Terry model give that a labeller prefers the first?
DPO: preferences without a reward model
The KL-regularised objective has a closed-form optimum, . Take logs and solve for the reward: . The partition function is intractable, but it depends only on the prompt, so it cancels in the Bradley–Terry difference. Fitting the policy straight to the preferences gives the DPO loss, a classifier on log-probability ratios:
import torch.nn.functional as F
def dpo_loss(pol_w, pol_l, ref_w, ref_l, beta=0.1):
# Summed log-probabilities of the chosen (w) and rejected (l) responses.
margin = beta * ((pol_w - ref_w) - (pol_l - ref_l))
return -F.logsigmoid(margin).mean()
DPO still needs preference pairs and a frozen reference, but no reward model and no sampling during training. Online methods that sample from the current policy remain common too; which wins depends on the task and the data.
Verifiable rewards and reasoning models
When a program can check the answer, you need no learned reward model. RL with verifiable rewards scores sampled solutions with unit tests, exact final answers or format checks. DeepSeek-R1 showed that this kind of RL on a strong base model, with rule-based rewards and no supervised reasoning traces (its R1-Zero variant), produces long chains of reasoning with self-checking. Reasoning can then be distilled: fine-tune a smaller student on a teacher's outputs, or match its token distributions. That is how R1's reasoning reached smaller open models.
LoRA: fine-tuning a low-rank update
Full fine-tuning updates every weight and keeps optimiser state for each. LoRA (Hu et al., 2021) freezes a pretrained and learns
The adapter has parameters instead of : at on a matrix, instead of . Its rank is at most , because every column of is times a column of and so lies in the span of 's columns. starts at zero, so training begins exactly at the pretrained model. Afterwards can be merged into at no inference cost, or kept apart so that many small adapters share one base. QLoRA (Dettmers et al., 2023) keeps the frozen base in 4-bit NormalFloat and trains 16-bit adapters through it, enough to fine-tune a 65B model on one 48 GB GPU.
Fine-tuning on narrow data can erase skills the model had: catastrophic forgetting. Mix some general data into the set, use a small learning rate and few epochs, and evaluate broadly before and after. LoRA helps here too: Biderman et al. (2024) found that it learns less than full fine-tuning but forgets less.
LoRA with rank adapts one weight matrix. How many trainable parameters does the adapter have?
Prompt, retrieve or fine-tune?
| Situation | Reach first for | Why |
|---|---|---|
| A new task you can describe in a paragraph | Prompting with examples | Minutes to iterate, no training (Prompting as Programming) |
| Answers depend on private or changing documents | Retrieval | Current, citable facts (Retrieval, Tools and Agents) |
| A fixed format at high volume, or a narrow task on a smaller model | SFT, usually with LoRA | Behaviour moves into the weights; prompts shrink |
| Quality people can judge but not specify | Preference tuning such as DPO | Learns from comparisons |
| New factual knowledge | Retrieval, not fine-tuning | Fine-tuned facts are unreliable and go stale |
Go down the table in order of cost, and let an eval suite decide each step (Prompt Version Control and Evals).
Read beyond
Book · free online · ~30 min
How to Scale Your Model: All the Transformer Math You Need to KnowAustin, Douglas, Frostig et al. · Forward and reverse FLOPs; the general rule of thumb for Transformer FLOPs
Rederive the 6 in 6ND from the forward and backward matrix multiplications, and find when attention's extra FLOPs stop being negligible.
Article · free online · ~40 min
FineWeb: decanting the web for the finest text data at scalePenedo et al., Hugging Face · The deduplication and filtering sections
List each filtering and deduplication step, and note which ones improved downstream benchmarks.
Book · free online · ~45 min
Reinforcement Learning from Human FeedbackNathan Lambert · The chapters on reward modelling, regularisation and direct alignment
Find where the KL penalty enters the RLHF objective, and follow how DPO removes the reward model.
Article · free online · ~25 min
Practical Tips for Finetuning LLMs Using LoRA (Low-Rank Adaptation)Sebastian Raschka · QLoRA trade-offs; balancing r and alpha; enabling LoRA for more layers
Note how rank, scaling and the choice of layers changed results, and what QLoRA traded for its memory savings.
Read the equation in context
Training Compute-Optimal Large Language ModelsJordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, et al. · NeurIPS, 2022Section 3.3 proposes the parametric loss as a risk decomposition, fits its five constants with a Huber loss over all runs, and gives the efficient frontier in closed form as Eq. (4). Compare its exponents with Table 2, and the 20-tokens rule with Table 3.
Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning & Chelsea Finn · NeurIPS, 2023Section 4 derives DPO in four moves: the optimal KL-regularised policy (Eq. 4), the reward rewritten through it (Eq. 5), the partition function cancelling inside Bradley–Terry (Eq. 6), and the loss (Eq. 7). Decode that loss below.
Decode the paper · Eq. (7), Section 4
Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning & Chelsea Finn · NeurIPS, 2023
Section 4 derives DPO in four moves: the optimal KL-regularised policy (Eq. 4), the reward rewritten through that policy (Eq. 5), the partition function cancelling inside Bradley–Terry (Eq. 6), and this maximum-likelihood loss (Eq. 7).
Options
Your turn
Pick a budget, slide the model size along its iso-compute curve and find the bottom of the bowl. Watch the ratio of the two loss terms as you go: the balance condition says it settles at at the optimum.
Interactive lab
Spend a compute budget
Tokens D = C/(6N)
1.67 × 10¹⁰
Tokens per parameter
16.7
Predicted loss
2.60813
Model term
0.3540
Data term
0.5642
Model ÷ data term
0.627
Serving this model costs about 2N = 2.00 × 10⁹ FLOPs per generated token, whatever it was trained on.
Match · Stage ↔ Objective
What each stage optimises
Options
Match · Expression ↔ Meaning
Formulas of the pipeline
Options
Match · Situation ↔ First reach for
Prompt, retrieve or fine-tune?
Options
Proof puzzle
The compute-optimal model size
Claim
Under , the loss is minimised at , where .
Tap lines in the order they should appear. Not every line belongs. Tap a line in your proof to send it back.
Your proof
- Pick the first line below.
Available lines
Prove it yourself
Rewards are defined only up to a constant
Claim
In the Bradley–Terry model , show that replacing by , for any function of the prompt alone, leaves every preference probability unchanged. Deduce that the term cancels in the derivation of DPO.
Your typeset proof appears here.
Coding problems
Problem 7·Warm-up
Count the adapter
A 32-layer transformer has no biases, and each layer has seven weight matrices, written as output input: query and output projections , key and value projections (grouped-query attention), MLP gate and up projections , and an MLP down projection . You attach a LoRA adapter of rank to every one of these matrices. How many trainable parameters does LoRA add in total?
Problem 8·Standard
DPO on six pairs
Six preference pairs have these sequence log-probabilities (each is the sum of a response's token log-probabilities), listed as :
, , , , , .
With , compute the mean DPO loss in nats, to 4 decimal places.
Problem 9·Challenge
Train it smaller, serve it cheaper
Use the Chinchilla fit . You need a model with loss exactly and expect it to generate tokens over its lifetime. Training costs FLOPs and inference FLOPs per generated token, so lifetime compute is .
Model 1 is training-optimal: among all with , it has the least training compute . Model 2 minimises lifetime compute under the same loss constraint. What percentage of Model 1's lifetime compute does Model 2 save? Give 1 decimal place.
Key takeaways
- Training costs about FLOPs and serving about per token; both shape the choice of model size.
- At a fixed budget the Chinchilla fit balances its two loss terms, so ; models built for heavy use are trained far past that point.
- SFT on chat-templated demonstrations makes an assistant; RLHF and DPO then learn from comparisons, tethered to a reference model by a KL penalty.
- LoRA learns a rank- update with parameters per matrix, and QLoRA adds a 4-bit frozen base.
- Prompt first, retrieve for facts, fine-tune for behaviour, and let evals decide.
Checkpoint
Prove it to the labyrinth
Answer every question to clear this chamber. First-try answers earn the most XP.
Why are many open-weight models trained on hundreds or thousands of tokens per parameter, far past the Chinchilla optimum?
A budget of FLOPs trains a model of parameters. With , how many tokens per parameter is that?
What does the penalty do in RLHF?
What does DPO remove from the RLHF pipeline?
With and , where , what is the largest possible rank of ?
A support assistant must answer questions about a price list that changes every day. What should you reach for first?
In supervised fine-tuning on chat transcripts, which tokens usually contribute to the loss?
What makes a reward verifiable in reinforcement learning with verifiable rewards?
End of the chamber
Clear this chamber
- Questions in this chamber (0/12 solved)Next unsolved
- Bonus: Compute-optimal (+40 XP)
- Bonus: Problem 7: Count the adapter (+20 XP)
- Bonus: Problem 8: DPO on six pairs (+35 XP)
- Bonus: Problem 9: Train it smaller, serve it cheaper (+50 XP)
- Bonus: Proof: The compute-optimal model size (+25 XP)
- Bonus: Proof: Rewards are defined only up to a constant (+35 XP)
- Bonus: Decode the paper (+25 XP)
- Bonus: Match: What each stage optimises (+20 XP)
- Bonus: Match: Formulas of the pipeline (+20 XP)
- Bonus: Match: Prompt, retrieve or fine-tune? (+20 XP)