Skip to content
AriadneTechnology

The Outer Ring · Chamber 3 of 9

From Pretraining to Assistant

Compute budgets and scaling laws, then instruction tuning, preference optimisation and low-rank fine-tuning.

55 min 60 XP + 12 questions + 1 challengeMathVideoPapersProofsCodeLab

In this chamber you will

  • Estimate training compute with C ≈ 6ND and read a scaling law
  • Split a compute budget between parameters and data, Chinchilla style
  • Explain supervised fine-tuning, RLHF and DPO, and what each one changes
  • Count LoRA's parameters and choose between prompting, retrieval and fine-tuning
DiscoverLearnRead beyondPapers & lecturesYour turn

In 2021 DeepMind trained Gopher: 280 billion parameters on 300 billion tokens. A few months later the same team spent the same compute on Chinchilla, with a quarter of the parameters and more than four times the data, and Chinchilla beat Gopher across a wide range of tasks. The difference was arithmetic, not architecture: how to split a fixed budget between a bigger model and more text. Their answer rests on a five-constant formula fitted to hundreds of training runs.

Spotted in the wild

L^(N,D)≜E+ANα+BDβ\hat{L}(N, D) \triangleq E + \frac{A}{N^{\alpha}} + \frac{B}{D^{\beta}}
Hoffmann et al. (2022), “Training Compute-Optimal Large Language Models”, Eq. (2)

This chamber follows a model from that budget to an assistant: pretraining and its compute, scaling laws, then instruction tuning, preference optimisation and cheap fine-tuning with LoRA.

DiscoverLearnRead beyondPapers & lecturesYour turn
Symbols for this chamber
  • NN“N”
    Number of parameters in the model.
    N=7×109N=7\times10^{9}
  • DD“D”
    Number of training tokens.
    D=1.4×1012D=1.4\times10^{12}
  • C≈6NDC\approx6ND“C is about six N D”
    Training compute in FLOPs: about 2N2N per token forward and 4N4N backward.
  • L^(N,D)\hat L(N,D)“L hat of N and D”
    Predicted final pretraining loss, in nats per token.
    L^=E+A/Nα+B/Dβ\hat L=E+A/N^{\alpha}+B/D^{\beta}
  • E, A, BE,\ A,\ B“E, A, B”
    Constants of the Chinchilla fit: irreducible loss 1.691.69 and term scales 406.4406.4 and 410.7410.7.
  • α, β\alpha,\ \beta“alpha, beta (scaling law)”
    Exponents of the model and data terms, fitted as 0.340.34 and 0.280.28.
  • Nopt(C)N_{\text{opt}}(C)“N opt of C”
    The model size that minimises the predicted loss at budget CC.
    Nopt∝Cβ/(α+β)N_{\text{opt}}\propto C^{\beta/(\alpha+\beta)}
  • πθ(y∣x)\pi_\theta(y\mid x)“pi theta of y given x”
    The policy: the model being tuned, as a distribution over responses yy to prompt xx.
  • πref\pi_{\text{ref}}“pi ref”
    The frozen reference model, usually the supervised fine-tuned model.
  • rϕ(x,y)r_\phi(x,y)“r phi of x and y”
    A learned reward model's scalar score for response yy to prompt xx.
  • yw, yly_w,\ y_l“y w, y l”
    The preferred (winning) and dispreferred (losing) responses in a comparison.
  • β\beta“beta (RLHF and DPO)”
    Strength of the KL penalty to the reference model. Not the scaling-law exponent.
    β=0.1\beta=0.1
  • σ(z)\sigma(z)“sigma of z”
    The logistic sigmoid 1/(1+e−z)1/(1+e^{-z}).
  • ΔW=BA\Delta W=BA“delta W equals B A”
    LoRA's low-rank update, with B∈Rd×rB\in\mathbb R^{d\times r} and A∈Rr×kA\in\mathbb R^{r\times k} (a matrix, not the Chinchilla constant).
    W′=W+BAW'=W+BA

Pretraining: one objective, trillions of tokens

Pretraining is the previous chamber's next-token prediction run at enormous scale: minimise −1T∑tlog⁡pθ(xt∣x<t)-\frac1T\sum_t\log p_\theta(x_t\mid x_{<t}) over a corpus of trillions of tokens. Most of it starts as web crawl, joined by code, books and papers. Raw crawl is mostly unusable, so the data pipeline does as much work as the optimiser:

  • Extraction and language identification: pull the text out of HTML and keep the languages you want.
  • Quality filtering: heuristic rules (length, symbol ratios, repeated lines) plus classifiers that score pages for quality or educational value.
  • Deduplication: drop exact and near-duplicate documents, commonly with MinHash, so repeated text is not memorised and over-weighted.
  • Hygiene: remove personal data and toxic content, and decontaminate by deleting benchmark test items so later evals stay honest.

FineWeb, an open dataset built this way, keeps 15 trillion tokens from 96 Common Crawl snapshots. The result of pretraining is a base model: it continues documents. Ask it a question and it may answer, or it may write three more questions, because a page of questions is plausible text.

Counting compute: C ≈ 6ND

In the forward pass each parameter takes part in one multiply and one add per token: about 2N2N FLOPs. The backward pass computes gradients for both activations and weights, roughly twice that, 4N4N. So training on DD tokens costs

C≈6ND FLOPs.C\approx6ND\ \text{FLOPs}.

Kaplan et al. (2020) note that attention adds a context-length term, small while the model width exceeds about a twelfth of the context. Check the rule on a public number: Llama 3's 405B model saw 15.6T tokens, and 6×4.05×1011×1.56×1013≈3.8×10256\times4.05\times10^{11}\times1.56\times10^{13}\approx3.8\times10^{25}, the figure its report gives. Inference runs only the forward pass, about 2N2N FLOPs per generated token. Remember that; it returns below.

Quick check +20 XP

A model with N=109N=10^9 parameters is trained on D=2×1010D=2\times10^{10} tokens. Using C≈6NDC\approx6ND, how many training FLOPs is that, in units of 102010^{20}?

Scaling laws and the compute-optimal split

Kaplan et al. found that loss falls as a power law in NN, DD and CC, with some trends spanning more than seven orders of magnitude, and advised spending most extra compute on a larger model (Nopt∝C0.73N_{\text{opt}}\propto C^{0.73}). Hoffmann et al. trained over 400 models, matched each run's learning-rate schedule to its length, and found instead that parameters and tokens should grow roughly equally. Their third method fits the Glimpse formula with E=1.69E=1.69, A=406.4A=406.4, B=410.7B=410.7, α=0.34\alpha=0.34 and β=0.28\beta=0.28. Here EE plays the entropy of text, and the other two terms are the penalties for a finite model and for finite data.

Fix the budget CC. Each model size NN then buys D=C/(6N)D=C/(6N) tokens, and these pairs trace an iso-compute curve (the paper's IsoFLOP profile). Find its lowest point by substitution. Put D=C/(6N)D=C/(6N) into the loss, differentiate in NN, set the result to zero and multiply by NN:

αAN−α=βB(C6)−βNβ=βBD−β.\alpha AN^{-\alpha}=\beta B\left(\tfrac C6\right)^{-\beta}N^{\beta}=\beta BD^{-\beta}.

At the optimum, α\alpha times the model term equals β\beta times the data term. Solving for NN gives the paper's Eq. (4):

Nopt=G(C6)βα+β,G=(αAβB)1α+β.N_{\text{opt}}=G\left(\frac C6\right)^{\frac{\beta}{\alpha+\beta}},\qquad G=\left(\frac{\alpha A}{\beta B}\right)^{\frac1{\alpha+\beta}}.

With the fitted exponents, Nopt∝C0.45N_{\text{opt}}\propto C^{0.45} and Dopt∝C0.55D_{\text{opt}}\propto C^{0.55}.

Quick check +20 XP

With the Chinchilla fit (α=0.34\alpha=0.34, β=0.28\beta=0.28), the compute-optimal size grows as Nopt∝Cβ/(α+β)N_{\text{opt}}\propto C^{\beta/(\alpha+\beta)}. By what factor does NoptN_{\text{opt}} grow when the budget grows tenfold?

Training past the optimum

Chinchilla minimises training compute. A deployed model also pays about 2N2N FLOPs for every token it generates, for as long as it serves. If it will generate trillions of tokens, a smaller model trained on more data reaches the same loss with a lower lifetime bill, as Sardana et al. (2024) work out with the Chinchilla fit. That is why open-weight families such as Llama 3 and Qwen2.5 report pretraining sets of 15–18 trillion tokens even for models of a few billion parameters: hundreds to thousands of tokens per parameter. Loss keeps falling out there, only slowly. The last coding problem puts numbers on the trade.

From base model to assistant

Supervised fine-tuning (SFT), or instruction tuning, continues training on prompt–response pairs: human-written demonstrations, curated datasets or a stronger model's outputs. Each conversation is rendered in the model's chat template, special tokens that mark roles and turns. An illustrative one:

Text
<|system|>You are a concise assistant.<|end|>
<|user|>Name a prime between 10 and 15.<|end|>
<|assistant|>11 (or 13).<|end|>

The loss is ordinary cross-entropy, counted only on the assistant's tokens. Templates differ between model families, so render with the model's own (for example apply_chat_template in Hugging Face Transformers, or the template your inference server loads with the weights). A wrong template is a silent bug: the model still answers, just worse. SFT mostly teaches format and behaviour: LIMA got a strong assistant from 1,000 well-chosen examples.

Learning from preferences: RLHF

Many qualities are easier to judge than to write. RLHF collects comparisons: for a prompt xx, a labeller marks response ywy_w as better than yly_l. A reward model rϕr_\phi is fitted with the Bradley–Terry likelihood

P(yw≻yl∣x)=σ(rϕ(x,yw)−rϕ(x,yl)),P(y_w\succ y_l\mid x)=\sigma\big(r_\phi(x,y_w)-r_\phi(x,y_l)\big),

minimising −log⁡σ(rw−rl)-\log\sigma(r_w-r_l): logistic regression on reward differences. Only differences enter, so adding a constant to every reward for a prompt changes nothing (you will prove it). The policy is then optimised, classically with PPO, to

max⁡πθ Ex, y∼πθ[rϕ(x,y)]−β DKL(πθ(⋅∣x) ∥ πref(⋅∣x)).\max_{\pi_\theta}\ \mathbb E_{x,\,y\sim\pi_\theta}\big[r_\phi(x,y)\big]-\beta\,D_{\mathrm{KL}}\big(\pi_\theta(\cdot\mid x)\,\|\,\pi_{\text{ref}}(\cdot\mid x)\big).

The KL penalty tethers the policy to the reference (the SFT model), where the reward model was trained and can be trusted. Without it the policy finds outputs the reward model overrates: reward hacking. In InstructGPT, labellers preferred a 1.3B model trained this way to the 175B GPT-3.

Quick check +20 XP

A reward model scores a preferred response 2.02.0 and a rejected one 0.50.5. What probability does the Bradley–Terry model give that a labeller prefers the first?

DPO: preferences without a reward model

The KL-regularised objective has a closed-form optimum, π∗(y∣x)=πref(y∣x) er(x,y)/β/Z(x)\pi^*(y\mid x)=\pi_{\text{ref}}(y\mid x)\,e^{r(x,y)/\beta}/Z(x). Take logs and solve for the reward: r=βlog⁡π∗πref+βlog⁡Z(x)r=\beta\log\frac{\pi^*}{\pi_{\text{ref}}}+\beta\log Z(x). The partition function Z(x)Z(x) is intractable, but it depends only on the prompt, so it cancels in the Bradley–Terry difference. Fitting the policy straight to the preferences gives the DPO loss, a classifier on log-probability ratios:

Python
import torch.nn.functional as F

def dpo_loss(pol_w, pol_l, ref_w, ref_l, beta=0.1):
    # Summed log-probabilities of the chosen (w) and rejected (l) responses.
    margin = beta * ((pol_w - ref_w) - (pol_l - ref_l))
    return -F.logsigmoid(margin).mean()

DPO still needs preference pairs and a frozen reference, but no reward model and no sampling during training. Online methods that sample from the current policy remain common too; which wins depends on the task and the data.

Verifiable rewards and reasoning models

When a program can check the answer, you need no learned reward model. RL with verifiable rewards scores sampled solutions with unit tests, exact final answers or format checks. DeepSeek-R1 showed that this kind of RL on a strong base model, with rule-based rewards and no supervised reasoning traces (its R1-Zero variant), produces long chains of reasoning with self-checking. Reasoning can then be distilled: fine-tune a smaller student on a teacher's outputs, or match its token distributions. That is how R1's reasoning reached smaller open models.

LoRA: fine-tuning a low-rank update

Full fine-tuning updates every weight and keeps optimiser state for each. LoRA (Hu et al., 2021) freezes a pretrained W∈Rd×kW\in\mathbb R^{d\times k} and learns

W′=W+BA,B∈Rd×r, A∈Rr×k, r≪min⁡(d,k).W'=W+BA,\qquad B\in\mathbb R^{d\times r},\ A\in\mathbb R^{r\times k},\ r\ll\min(d,k).

The adapter has r(d+k)r(d+k) parameters instead of dkdk: at r=16r=16 on a 4096×40964096\times4096 matrix, 131,072131{,}072 instead of 16,777,21616{,}777{,}216. Its rank is at most rr, because every column of BABA is BB times a column of AA and so lies in the span of BB's rr columns. BB starts at zero, so training begins exactly at the pretrained model. Afterwards BABA can be merged into WW at no inference cost, or kept apart so that many small adapters share one base. QLoRA (Dettmers et al., 2023) keeps the frozen base in 4-bit NormalFloat and trains 16-bit adapters through it, enough to fine-tune a 65B model on one 48 GB GPU.

Fine-tuning on narrow data can erase skills the model had: catastrophic forgetting. Mix some general data into the set, use a small learning rate and few epochs, and evaluate broadly before and after. LoRA helps here too: Biderman et al. (2024) found that it learns less than full fine-tuning but forgets less.

Quick check +20 XP

LoRA with rank r=8r=8 adapts one 4096×40964096\times4096 weight matrix. How many trainable parameters does the adapter have?

Prompt, retrieve or fine-tune?

SituationReach first forWhy
A new task you can describe in a paragraphPrompting with examplesMinutes to iterate, no training (Prompting as Programming)
Answers depend on private or changing documentsRetrievalCurrent, citable facts (Retrieval, Tools and Agents)
A fixed format at high volume, or a narrow task on a smaller modelSFT, usually with LoRABehaviour moves into the weights; prompts shrink
Quality people can judge but not specifyPreference tuning such as DPOLearns from comparisons
New factual knowledgeRetrieval, not fine-tuningFine-tuned facts are unreliable and go stale

Go down the table in order of cost, and let an eval suite decide each step (Prompt Version Control and Evals).

DiscoverLearnRead beyondPapers & lecturesYour turn

Read beyond

Book · free online · ~30 min

How to Scale Your Model: All the Transformer Math You Need to Know

Austin, Douglas, Frostig et al. · Forward and reverse FLOPs; the general rule of thumb for Transformer FLOPs

Rederive the 6 in 6ND from the forward and backward matrix multiplications, and find when attention's extra FLOPs stop being negligible.

Article · free online · ~40 min

FineWeb: decanting the web for the finest text data at scale

Penedo et al., Hugging Face · The deduplication and filtering sections

List each filtering and deduplication step, and note which ones improved downstream benchmarks.

Book · free online · ~45 min

Reinforcement Learning from Human Feedback

Nathan Lambert · The chapters on reward modelling, regularisation and direct alignment

Find where the KL penalty enters the RLHF objective, and follow how DPO removes the reward model.

Article · free online · ~25 min

Practical Tips for Finetuning LLMs Using LoRA (Low-Rank Adaptation)

Sebastian Raschka · QLoRA trade-offs; balancing r and alpha; enabling LoRA for more layers

Note how rank, scaling and the choice of layers changed results, and what QLoRA traded for its memory savings.

DiscoverLearnRead beyondPapers & lecturesYour turn

Read the equation in context

Training Compute-Optimal Large Language ModelsJordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, et al. · NeurIPS, 2022

Section 3.3 proposes the parametric loss as a risk decomposition, fits its five constants with a Huber loss over all runs, and gives the efficient frontier in closed form as Eq. (4). Compare its exponents with Table 2, and the 20-tokens rule with Table 3.

Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning & Chelsea Finn · NeurIPS, 2023

Section 4 derives DPO in four moves: the optimal KL-regularised policy (Eq. 4), the reward rewritten through it (Eq. 5), the partition function cancelling inside Bradley–Terry (Eq. 6), and the loss (Eq. 7). Decode that loss below.

Decode the paper · Eq. (7), Section 4

Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning & Chelsea Finn · NeurIPS, 2023

+25 XP
LDPO(πθ;πref)=−E(x,yw,yl)∼D[log⁡σ(βlog⁡πθ(yw∣x)πref(yw∣x)−βlog⁡πθ(yl∣x)πref(yl∣x))]\mathcal L_{\text{DPO}}(\pi_\theta;\pi_{\text{ref}})=-\mathbb E_{(x,y_w,y_l)\sim\mathcal D}\left[\log\sigma\left(\beta\log\frac{\pi_\theta(y_w\mid x)}{\pi_{\text{ref}}(y_w\mid x)}-\beta\log\frac{\pi_\theta(y_l\mid x)}{\pi_{\text{ref}}(y_l\mid x)}\right)\right]

Section 4 derives DPO in four moves: the optimal KL-regularised policy (Eq. 4), the reward rewritten through that policy (Eq. 5), the partition function cancelling inside Bradley–Terry (Eq. 6), and this maximum-likelihood loss (Eq. 7).

πθ\pi_\theta
πref\pi_{\text{ref}}
ywy_w
yly_l
β\beta
σ\sigma

Options

State of GPT | BRK216HFSMicrosoft Developer · 43 min
Direct Preference Optimization (DPO) explained: Bradley-Terry model, log probabilities, mathUmar Jamil · 49 min
DiscoverLearnRead beyondPapers & lecturesYour turn

Your turn

Pick a budget, slide the model size along its iso-compute curve and find the bottom of the bowl. Watch the ratio of the two loss terms as you go: the balance condition says it settles at β/α≈0.82\beta/\alpha\approx0.82 at the optimum.

Interactive lab

Spend a compute budget

Fix a training budget CC. Every model size NN then buys D=C/(6N)D=C/(6N) tokens. Slide along this iso-compute curve and watch the Chinchilla loss L=1.69+406.4/N0.34+410.7/D0.28L=1.69+406.4/N^{0.34}+410.7/D^{0.28}: a floor, a model term and a data term. Submit the size that minimises it for each budget.
Predicted loss along the iso-compute curve72.528.52.81103.111.53.39133.68log₁₀ N (parameters)Predicted loss L
● Loss at this budget● Your model● 20 tokens per parameter

Tokens D = C/(6N)

1.67 × 10¹⁰

Tokens per parameter

16.7

Predicted loss

2.60813

Model term

0.3540

Data term

0.5642

Model ÷ data term

0.627

Serving this model costs about 2N = 2.00 × 10⁹ FLOPs per generated token, whatever it was trained on.

Challenge: Compute-optimalFor three compute budgets, find the model size that minimises the Chinchilla loss, each to within 10%.+40 XP

Match · Stage ↔ Objective

What each stage optimises

+20 XP
Pretraining
Supervised fine-tuning
Reward modelling
RLHF policy step
DPO
RL with verifiable rewards

Options

Match · Expression ↔ Meaning

Formulas of the pipeline

+20 XP
6ND6ND
2N2N
r(d+k)r(d+k)
σ(rw−rl)\sigma(r_w-r_l)
βlog⁡πθ(y∣x)πref(y∣x)\beta\log\frac{\pi_\theta(y\mid x)}{\pi_{\text{ref}}(y\mid x)}
EE

Options

Match · Situation ↔ First reach for

Prompt, retrieve or fine-tune?

+20 XP
A new task you can describe in a paragraph
Answers must cite documents that change every week
A narrow, high-volume task must run on a much smaller model in a fixed format
People can judge good tone and helpfulness but cannot write rules for it

Options

Proof puzzle

The compute-optimal model size

+25 XP

Claim

Under C=6NDC=6ND, the loss L=E+AN−α+BD−βL=E+AN^{-\alpha}+BD^{-\beta} is minimised at Nopt=G(C/6)β/(α+β)N_{\text{opt}}=G(C/6)^{\beta/(\alpha+\beta)}, where G=(αA/βB)1/(α+β)G=(\alpha A/\beta B)^{1/(\alpha+\beta)}.

Tap lines in the order they should appear. Not every line belongs. Tap a line in your proof to send it back.

Your proof

  1. Pick the first line below.

Available lines

Prove it yourself

Rewards are defined only up to a constant

+35 XP

Claim

In the Bradley–Terry model P(yw≻yl∣x)=σ(r(x,yw)−r(x,yl))P(y_w\succ y_l\mid x)=\sigma(r(x,y_w)-r(x,y_l)), show that replacing r(x,y)r(x,y) by r(x,y)+c(x)r(x,y)+c(x), for any function cc of the prompt alone, leaves every preference probability unchanged. Deduce that the term βlog⁡Z(x)\beta\log Z(x) cancels in the derivation of DPO.

Preview

Your typeset proof appears here.

Coding problems

Problem 7·Warm-up

Count the adapter

+20 XP

A 32-layer transformer has no biases, and each layer has seven weight matrices, written as output ×\times input: query and output projections 4096×40964096\times4096, key and value projections 1024×40961024\times4096 (grouped-query attention), MLP gate and up projections 14336×409614336\times4096, and an MLP down projection 4096×143364096\times14336. You attach a LoRA adapter of rank r=16r=16 to every one of these matrices. How many trainable parameters does LoRA add in total?

An exact integer (or a fraction like 7/12)

Problem 8·Standard

DPO on six pairs

+35 XP

Six preference pairs have these sequence log-probabilities (each is the sum of a response's token log-probabilities), listed as (log⁡πθ(yw∣x), log⁡πref(yw∣x), log⁡πθ(yl∣x), log⁡πref(yl∣x))(\log\pi_\theta(y_w\mid x),\ \log\pi_{\text{ref}}(y_w\mid x),\ \log\pi_\theta(y_l\mid x),\ \log\pi_{\text{ref}}(y_l\mid x)):

(−12.0,−13.5,−15.0,−14.0)(-12.0,-13.5,-15.0,-14.0), (−20.4,−20.0,−18.1,−19.2)(-20.4,-20.0,-18.1,-19.2), (−8.2,−9.0,−11.5,−10.4)(-8.2,-9.0,-11.5,-10.4), (−30.0,−31.2,−29.5,−28.0)(-30.0,-31.2,-29.5,-28.0), (−15.6,−15.6,−16.3,−16.3)(-15.6,-15.6,-16.3,-16.3), (−41.3,−44.0,−39.8,−40.1)(-41.3,-44.0,-39.8,-40.1).

With β=0.1\beta=0.1, compute the mean DPO loss in nats, to 4 decimal places.

A number, rounded to 4 decimal places

Problem 9·Challenge

Train it smaller, serve it cheaper

+50 XP

Use the Chinchilla fit L(N,D)=1.69+406.4/N0.34+410.7/D0.28L(N,D)=1.69+406.4/N^{0.34}+410.7/D^{0.28}. You need a model with loss exactly L=2.0L=2.0 and expect it to generate T=1013T=10^{13} tokens over its lifetime. Training costs 6ND6ND FLOPs and inference 2N2N FLOPs per generated token, so lifetime compute is 6ND+2NT6ND+2NT.

Model 1 is training-optimal: among all (N,D)(N,D) with L(N,D)=2.0L(N,D)=2.0, it has the least training compute 6ND6ND. Model 2 minimises lifetime compute 6ND+2NT6ND+2NT under the same loss constraint. What percentage of Model 1's lifetime compute does Model 2 save? Give 1 decimal place.

A number, rounded to 1 decimal place

Key takeaways

  • Training costs about 6ND6ND FLOPs and serving about 2N2N per token; both shape the choice of model size.
  • At a fixed budget the Chinchilla fit balances its two loss terms, so Nopt∝C0.45N_{\text{opt}}\propto C^{0.45}; models built for heavy use are trained far past that point.
  • SFT on chat-templated demonstrations makes an assistant; RLHF and DPO then learn from comparisons, tethered to a reference model by a KL penalty.
  • LoRA learns a rank-rr update with r(d+k)r(d+k) parameters per matrix, and QLoRA adds a 4-bit frozen base.
  • Prompt first, retrieve for facts, fine-tune for behaviour, and let evals decide.

Checkpoint

Prove it to the labyrinth

Answer every question to clear this chamber. First-try answers earn the most XP.

0/8
Question 1 of 8 +20 XP

Why are many open-weight models trained on hundreds or thousands of tokens per parameter, far past the Chinchilla optimum?

Question 2 of 8 +20 XP

A budget of C=1.2×1022C=1.2\times10^{22} FLOPs trains a model of N=1010N=10^{10} parameters. With C≈6NDC\approx6ND, how many tokens per parameter is that?

Question 3 of 8 +20 XP

What does the penalty βDKL(πθ∥πref)\beta D_{\mathrm{KL}}(\pi_\theta\|\pi_{\text{ref}}) do in RLHF?

Question 4 of 8 +20 XP

What does DPO remove from the RLHF pipeline?

Question 5 of 8 +20 XP

With B∈Rd×rB\in\mathbb R^{d\times r} and A∈Rr×kA\in\mathbb R^{r\times k}, where r≤min⁡(d,k)r\le\min(d,k), what is the largest possible rank of ΔW=BA\Delta W=BA?

Question 6 of 8 +20 XP

A support assistant must answer questions about a price list that changes every day. What should you reach for first?

Question 7 of 8 +20 XP

In supervised fine-tuning on chat transcripts, which tokens usually contribute to the loss?

Question 8 of 8 +20 XP

What makes a reward verifiable in reinforcement learning with verifiable rewards?

End of the chamber

Clear this chamber

+60 XPScaling LawInstruction TuningPreference OptimisationLoRA