Skip to content
AriadneTechnology

The Outer Ring · Chamber 2 of 9

Next-Token Prediction and the Transformer

One distribution per position: logits and softmax, cross-entropy and perplexity, causal self-attention and the KV cache.

55 min 60 XP + 14 questions + 1 challengeMathVideoPapersProofsCodeLab

In this chamber you will

  • Factorise the probability of a text into next-token predictions
  • Compute cross-entropy and perplexity, and read them as compression
  • Compute scaled dot-product attention with a causal mask
  • Explain prefill, decoding and why the KV cache makes generation fast
DiscoverLearnRead beyondPapers & lecturesYour turn

Ask a large language model anything and, underneath, it does one thing over and over: it reads the tokens so far and outputs a probability for every token that could come next. Choose one, append it, repeat. Training, perplexity, attention and the KV cache all follow from taking that loop seriously. The GPT-2 report states the starting point in one line.

Spotted in the wild

p(x)=∏i=1np(sn∣s1,…,sn−1)p(x)=\prod_{i=1}^{n}p(s_n\mid s_1,\ldots,s_{n-1})
Radford et al. (2019), “Language Models are Unsupervised Multitask Learners”, Equation (1)Read the indices carefully: the product is written over i, but each factor uses n. It means one factor per position.
DiscoverLearnRead beyondPapers & lecturesYour turn
Symbols for this chamber
  • x<tx_{<t}“x before t”
    The prefix: every token before position tt.
    x<3=(x1,x2)x_{<3}=(x_1,x_2)
  • pθ(xt∣x<t)p_\theta(x_t\mid x_{<t})“p theta of x t given the prefix”
    The model's next-token distribution at position tt.
  • ztz_t“z t, the logits”
    One real score per vocabulary token at position tt.
    zt∈R∣V∣z_t\in\mathbb R^{|\mathcal V|}
  • softmax⁡(z)i\operatorname{softmax}(z)_i“softmax of z, entry i”
    Exponentiate and normalise: ezi/∑jezje^{z_i}/\sum_je^{z_j}.
  • L\mathcal L“script L, the loss”
    Cross-entropy: the mean negative log-likelihood per token.
    L=−1T∑tlog⁡pθ(xt∣x<t)\mathcal L=-\frac1T\sum_t\log p_\theta(x_t\mid x_{<t})
  • PPL⁡\operatorname{PPL}“perplexity”
    eLe^{\mathcal L}: the effective number of tokens the model is choosing between.
  • qi, kj, vjq_i,\ k_j,\ v_j“query i, key j, value j”
    Projections used to score position jj for position ii, and to average.
    qi=WQhiq_i=W_Qh_i
  • dkd_k“d k”
    Dimension of queries and keys; scores are divided by dk\sqrt{d_k}.
  • MijM_{ij}“the causal mask”
    00 when j≤ij\le i and −∞-\infty when j>ij>i, added to the scores before the softmax.
  • ∣V∣|\mathcal V|“size of the vocabulary”
    The number of classes in every next-token prediction.
  • NN“N, the parameter count”
    A forward pass costs about 2N2N FLOPs per token.

One classification problem per position

The chain rule of probability is exact for any sequence of tokens:

p(x1,…,xT)=∏t=1Tp(xt∣x<t).p(x_1,\ldots,x_T)=\prod_{t=1}^{T}p(x_t\mid x_{<t}).

No Markov assumption hides here. An nn-gram model truncates the history; a transformer conditions on the whole prefix that fits in its context window. Language modelling is therefore TT classification problems over the vocabulary V\mathcal V, one per position. From the prefix x<tx_{<t} the network computes logits zt∈R∣V∣z_t\in\mathbb R^{|\mathcal V|}, and the softmax turns them into a distribution:

pθ(xt=i∣x<t)=softmax⁡(zt)i=ezt,i∑jezt,j.p_\theta(x_t=i\mid x_{<t})=\operatorname{softmax}(z_t)_i=\frac{e^{z_{t,i}}}{\sum_{j}e^{z_{t,j}}}.

In code, the output at position t−1t-1 is the prediction for token tt, which is why the targets are the inputs shifted left by one. Adding the same constant cc to every logit changes nothing, because ezi+c/∑jezj+c=ecezi/(ec∑jezj)e^{z_i+c}/\sum_je^{z_j+c}=e^{c}e^{z_i}/(e^{c}\sum_je^{z_j}). Only differences between logits matter, so implementations subtract the largest logit before exponentiating and nothing can overflow:

Python
def log_softmax(z):
    m = z.max(axis=-1, keepdims=True)  # shift invariance: subtract the max
    return z - m - np.log(np.exp(z - m).sum(axis=-1, keepdims=True))
Quick check +20 XP

A three-token vocabulary has logits (3,1,1)(3,1,1). What probability does softmax give the first token? (Subtract 1 from every logit first.)

Cross-entropy and teacher forcing

Training maximises the likelihood of real text, which is the same as minimising the mean negative log-likelihood, the cross-entropy loss:

L(θ)=−1T∑t=1Tlog⁡pθ(xt∣x<t).\mathcal L(\theta)=-\frac1T\sum_{t=1}^{T}\log p_\theta(x_t\mid x_{<t}).

For one position with target yy, the loss is −zy+log⁡∑jezj-z_y+\log\sum_je^{z_j}, and its gradient is ∂ℓ/∂z=softmax⁡(z)−ey\partial\ell/\partial z=\operatorname{softmax}(z)-e_y: predicted probabilities minus the one-hot target. Each wrong token's logit is pushed down in proportion to the probability it took.

During training the model conditions on the true prefix, never on its own guesses. This is teacher forcing, and together with the causal mask below it lets one forward pass score all TT positions in parallel. At generation time the model must condition on its own samples instead, so early mistakes can compound (sometimes called exposure bias).

A useful check: a freshly initialised model predicts nearly uniformly, so its loss should start near ln⁡∣V∣\ln|\mathcal V|. A first logged loss far above that points at the initialisation or the loss code.

Quick check +20 XP

A freshly initialised model with a vocabulary of 50,257 tokens predicts almost uniformly. Near what cross-entropy, in nats per token, should training start?

Perplexity, bits and compression

Perplexity is the exponential of the cross-entropy, the reciprocal of the geometric-mean probability of the observed tokens:

PPL⁡=eL=(∏t=1Tpθ(xt∣x<t))−1/T.\operatorname{PPL}=e^{\mathcal L}=\Big(\prod_{t=1}^{T}p_\theta(x_t\mid x_{<t})\Big)^{-1/T}.

A model that gives every token probability 1/∣V∣1/|\mathcal V| has L=ln⁡∣V∣\mathcal L=\ln|\mathcal V|, so its perplexity is exactly ∣V∣|\mathcal V|: perplexity is the effective number of tokens the model is choosing between, and a perfect model scores 1. Dividing L\mathcal L by ln⁡2\ln2 gives bits per token, and the bits are literal. An arithmetic coder driven by the model's probabilities encodes a text in about ∑t−log⁡2pθ(xt∣x<t)\sum_t-\log_2p_\theta(x_t\mid x_{<t}) bits, within a couple of bits for the whole message, so a better predictor is a better compressor (Delétang et al., 2024).

QuantityFormulaUniform model over ∣V∣\lvert\mathcal V\rvert tokens
Cross-entropy, nats per tokenL\mathcal Lln⁡∣V∣\ln\lvert\mathcal V\rvert
Bits per tokenL/ln⁡2\mathcal L/\ln2log⁡2∣V∣\log_2\lvert\mathcal V\rvert
PerplexityeLe^{\mathcal L}∣V∣\lvert\mathcal V\rvert
Bits per bytetotal bits divided by bytes of textdepends on the tokeniser

Per-token perplexity depends on the tokeniser: splitting the same text into fewer, longer tokens raises it without changing how well the text is predicted. Compare models on the same text in bits per byte, with the same context length, and remember that perplexity measures prediction, not helpfulness or truth.

Quick check +20 XP

A model assigns probabilities 1/21/2, 1/41/4 and 1/81/8 to the three tokens of a text. What is its perplexity on that text?

The decoder-only transformer

In most current LLMs, the network that computes ztz_t is a stack of identical blocks:

Text
tokens x_1 … x_T
  → embedding lookup: one row of a |V| × d table per token (chamber 1)
  → positional information (learned vectors, or RoPE inside attention)
  → L blocks, each:
        h ← h + Attention(Norm(h))   mixes information across positions
        h ← h + MLP(Norm(h))         processes each position on its own
  → final Norm → unembedding (d × |V|) → logits z_t

Attention is the only place where positions exchange information. Residual connections let each block add a correction to the stream rather than replace it, and normalising before each sublayer (pre-norm, as in GPT-2; many newer models use RMSNorm) keeps deep stacks trainable. Apart from the mask, attention scores ignore token order, so position is injected explicitly: GPT-2 adds a learned vector per position, while rotary embeddings, RoPE (Su et al., 2021), rotate each query and key by an angle proportional to its position so that q⋅kq\cdot k depends on the relative offset. Open-weight families such as Llama, Mistral and Qwen use RoPE.

Each block holds about 4d24d^2 parameters in attention and 8d28d^2 in an MLP with a 4d4d hidden layer, so N≈12Ld2N\approx12Ld^2 without embeddings (Kaplan et al., 2020). For GPT-2 small, L=12L=12, d=768d=768 and ∣V∣=50,257|\mathcal V|=50{,}257 give 12⋅12⋅7682+50,257⋅768≈123.512\cdot12\cdot768^2+50{,}257\cdot768\approx123.5 million; position embeddings, biases and norms bring it to the 124 million of the released checkpoint.

Scaled dot-product attention, step by step

Let hi∈Rdh_i\in\mathbb R^d be the input at position ii. One attention head projects it three ways: a query qi=WQhiq_i=W_Qh_i (what this position looks for), a key kj=WKhjk_j=W_Kh_j (what position jj can be found by) and a value vj=WVhjv_j=W_Vh_j (what position jj hands over), with q,k∈Rdkq,k\in\mathbb R^{d_k}.

  1. Score each position: sij=qi⋅kj/dks_{ij}=q_i\cdot k_j/\sqrt{d_k}.
  2. Mask the future: add Mij=−∞M_{ij}=-\infty for j>ij>i (and 0 otherwise), so position ii cannot see the tokens it must predict.
  3. Normalise: aij=esij+Mij/∑lesil+Mila_{ij}=e^{s_{ij}+M_{ij}}/\sum_le^{s_{il}+M_{il}}. Masked entries get e−∞=0e^{-\infty}=0, exactly.
  4. Average the values: oi=∑j≤iaijvjo_i=\sum_{j\le i}a_{ij}v_j.

Stacking the rows into matrices gives the whole layer in one line:

Attention⁡(Q,K,V)=softmax⁡(QK⊤dk+M)V.\operatorname{Attention}(Q,K,V)=\operatorname{softmax}\Big(\frac{QK^\top}{\sqrt{d_k}}+M\Big)V.

Python
def causal_attention(Q, K, V):
    T, d_k = Q.shape
    S = Q @ K.T / np.sqrt(d_k)              # (T, T) scores
    S[np.triu_indices(T, k=1)] = -np.inf    # mask j > i
    A = np.exp(S - S.max(axis=1, keepdims=True))
    A /= A.sum(axis=1, keepdims=True)       # each row sums to 1
    return A @ V

Three facts are worth knowing cold. The weights in each row are nonnegative and sum to 1, so every output lies in the convex hull of the visible values: attention blends, it never extrapolates. If the components of qq and kk are independent with mean 0 and variance 1, then q⋅kq\cdot k has variance dkd_k; without the division, wide heads would push the softmax into nearly one-hot regions with tiny gradients. And multiplying a query by cc multiplies every score by cc, which is a softmax temperature of 1/c1/c in disguise. Temperature returns, applied to the output logits, in the chamber on sampling.

Quick check +20 XP

In a causal decoder, what attention weight does the query at position 3 place on position 5?

Multi-head attention runs hh heads in parallel, each with its own projections to dk=d/hd_k=d/h dimensions, concatenates their outputs and mixes them with WOW_O, at about the cost of one full-width head. Heads specialise: one may attend to the previous token, another may find an earlier copy of the current token and look at what followed it (an induction head). Grouped-query attention (Ainslie et al., 2023) lets several query heads share one key–value head, which shrinks the cache below.

Generation: prefill, decode and the KV cache

Generation runs the opening loop: compute logits at the last position, choose a token (greedily, or by sampling with temperature, top-k or top-p, which is the subject of the sampling chamber), append it, repeat. It has two phases. Prefill processes the whole prompt in one parallel pass. Decode then produces one token per forward pass, and each new query must attend to the keys and values of every earlier position.

Because of the causal mask, earlier positions never see later tokens, so their keys and values never change. The model therefore stores them. The KV cache holds kjk_j and vjv_j for every layer, head and position so far; a decode step computes q,k,vq,k,v for the new token only, appends its k,vk,v and attends over the cache. Without it, every step would reprocess the whole prefix: generating GG tokens after a prompt of PP would take about GP+G2/2GP+G^2/2 token passes instead of P+GP+G.

PrefillDecode step
Tokens processedthe whole prompt, in parallelone
Weight-matrix workabout 2N2N FLOPs per prompt tokenabout 2N2N FLOPs
Attention workquadratic in prompt lengthlinear in current length
Usually limited byarithmeticreading weights and cache from memory

The 2N2N comes from one multiply and one add per parameter; Kaplan et al. write the whole forward pass as Cforward≈2N+2nlayernctxdmodelC_{\text{forward}}\approx2N+2n_{\text{layer}}n_{\text{ctx}}d_{\text{model}} FLOPs per token. Over a context of length TT the score matrices have T2T^2 entries per head and layer, so attention's total cost is quadratic in TT. Kernels such as FlashAttention (Dao et al., 2022) avoid storing those matrices but still do the quadratic arithmetic. The cache costs 2⋅L⋅nkv⋅dhead2\cdot L\cdot n_{\text{kv}}\cdot d_{\text{head}} numbers per token; the chamber on running models turns that into gigabytes.

Quick check +20 XP

Roughly how many GFLOPs does an 8-billion-parameter model spend on the matrix work for one generated token?

GFLOPs
DiscoverLearnRead beyondPapers & lecturesYour turn

Read beyond

Book · free online · ~25 min

Speech and Language Processing (3rd ed. draft)

Dan Jurafsky & James H. Martin · Chapter “N-gram Language Models”: evaluating language models and perplexity

Work through the perplexity examples and check that a uniform model over a vocabulary of size V scores exactly V.

Book · free online · ~25 min

Dive into Deep Learning

Zhang, Lipton, Li & Smola · Section 11.3: attention scoring functions and the masked softmax

Compare their masked softmax, which uses a large negative number, with setting masked scores to minus infinity.

Article · free online · ~30 min

The Illustrated GPT-2 (Visualizing Transformer Language Models)

Jay Alammar · Parts 1 and 2: the decoder stack and masked self-attention

Trace one token through the stack and name the matrix that produces each query, key and value.

Article · free online · ~30 min

Transformer Inference Arithmetic

kipply · Sections “kv cache” and “flops counting”

Reproduce the cache size per token for a model you use, and explain why decoding is limited by memory rather than arithmetic.

DiscoverLearnRead beyondPapers & lecturesYour turn

Read the equation in context

Attention Is All You NeedAshish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser & Illia Polosukhin · NeurIPS, 2017

Section 3.2.1 defines Equation (1) with queries, keys and values packed into matrices, so all outputs come from two matrix products. Read three passages around it: the footnote that justifies dk\sqrt{d_k} with the variance argument above; Section 3.2.2, where h=8h=8 heads of dk=dmodel/h=64d_k=d_{\text{model}}/h=64 dimensions replace one wide head; and Section 3.2.3, where the decoder prevents leftward information flow by setting the scores of illegal connections to −∞-\infty inside the softmax. The paper's model is an encoder–decoder for translation. GPT-style models keep only the masked decoder stack, drop the cross-attention to an encoder, and train it with the factorisation from the opening glimpse.

Decode the paper · Section 3.2.1, Equation (1): scaled dot-product attention

Attention Is All You Need

Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser & Illia Polosukhin · NeurIPS, 2017

+25 XP
Attention⁡(Q,K,V)=softmax⁡(QKTdk)V\operatorname{Attention}(Q,K,V)=\operatorname{softmax}\left(\frac{QK^T}{\sqrt{d_k}}\right)V

The paper packs many queries into one matrix and computes all attention outputs with two matrix products. Its footnote justifies the scaling: with independent unit-variance components, q⋅kq\cdot k has variance dkd_k. Section 3.2.3 adds the decoder's mask by setting illegal connections to −∞-\infty inside the softmax.

QQ
KK
VV
QKTQK^T
dk\sqrt{d_k}
softmax⁡\operatorname{softmax}

Options

Transformers, the tech behind LLMs | Deep Learning Chapter 53Blue1Brown · 27 min

For the whole pipeline written and trained in code, from a bigram baseline to a small GPT with causal self-attention:

Let's build GPT: from scratch, in code, spelled out.Andrej Karpathy · 116 min
DiscoverLearnRead beyondPapers & lecturesYour turn

Your turn

Six toy tokens have fixed 2-D keys and values. Steer one query until a single token takes almost all the attention, then shorten or turn it until the attention spreads out, and watch the output stay inside the hull of the visible values. The panel underneath shows what goes wrong without the d\sqrt{d}.

Interactive lab

Attention steerer

Six toy tokens, each with a fixed 2-D key (left) and value (right). Choose the position whose query you steer: later tokens are masked. Drag in the key plane, or use the sliders, to aim the query q. The weights are softmax(q·k/√d) with d = 2, and the output is the weighted average of the visible values.

Query position

AriadnegaveTheseusaredthread
Keys and the query (drag to aim)
AriadnegaveTheseusaredthreado
Values, their hull and the output o
Ariadne0.71 · 8.6%
gave1.98 · 30.9%
Theseus−1.70 · 0.8%
a−0.28 · 3.2%
red−0.71 · 2.1%
thread2.55 · 54.4%

Each row shows the score q·k/√2 and its weight.

Top weight

54.4%

Temperature √d/‖q‖

1.41

Output o

(0.60, −0.07)

Visible tokens

6 of 6

Focus ≥ 90% ·   Spread ≤ 40% ·

Why divide by √d?

Draw a query and six keys with independent standard normal components in d dimensions, forty times. Raw dot products spread out like √d and the softmax saturates; dividing by √d keeps the spread near 1.

Spread of q·k (√d = 8)

8.66

Spread of q·k/√d

1.08

Mean top weight, raw

91%

Mean top weight, scaled

45%

Keys and values are illustrative; in a real head they are learned projections with d around 64 to 128. The lab uses d = 2.

Challenge: Attention steererAim a query so one token takes at least 90% of the attention, then spread it so no token takes more than 40%.+30 XP

Match · Expression ↔ Meaning

From logits to loss

+20 XP
−1T∑tlog⁡pθ(xt∣x<t)-\frac1T\sum_t\log p_\theta(x_t\mid x_{<t})
eLe^{\mathcal L}
L/ln⁡2\mathcal L/\ln2
ln⁡∣V∣\ln|\mathcal V|
softmax⁡(z)−ey\operatorname{softmax}(z)-e_y

Options

Match · Term ↔ Meaning

Attention and generation

+20 XP
Q @ K.T / np.sqrt(d_k)
Mij=−∞M_{ij}=-\infty for j>ij>i
Prefill
Decode step
2N2N
O(T2)O(T^2)

Options

Proof puzzle

Causal attention only blends the past

+25 XP

Claim

With causal mask MM, show that each output oi=∑jaijvjo_i=\sum_ja_{ij}v_j of softmax⁡(QK⊤/dk+M)V\operatorname{softmax}(QK^\top/\sqrt{d_k}+M)V is a convex combination of v1,…,viv_1,\ldots,v_i.

Tap lines in the order they should appear. Not every line belongs. Tap a line in your proof to send it back.

Your proof

  1. Pick the first line below.

Available lines

Prove it yourself

Shift invariance and the stable log-softmax

+35 XP

Claim

Prove that softmax⁡(z+c1)=softmax⁡(z)\operatorname{softmax}(z+c\mathbf 1)=\operatorname{softmax}(z) for every real cc. Deduce that, with m=max⁡jzjm=\max_jz_j, log⁡softmax⁡(z)i=zi−m−log⁡∑jezj−m\log\operatorname{softmax}(z)_i=z_i-m-\log\sum_je^{z_j-m}, and explain why this formula cannot overflow.

Preview

Your typeset proof appears here.

Coding problems

Problem 4·Warm-up

Perplexity from raw logits

+20 XP

A model with a four-token vocabulary (tokens 0 to 3) outputs logits at four positions. At position 1 the logits are (2,1,0,−1)(2,1,0,-1) and the true next token is 0; at position 2 they are (0,0,0,0)(0,0,0,0) and the token is 3; at position 3 they are (1,3,0,2)(1,3,0,2) and the token is 1; at position 4 they are (999,1002,1002,1000)(999,1002,1002,1000) and the token is 2.

Compute the model's perplexity on these four tokens, using natural logarithms, to 4 decimal places.

A number, rounded to 4 decimal places

Problem 5·Standard

Causal attention at full size

+35 XP

Generate numbers with the Park–Miller generator x0=2026x_0=2026, xn+1=16807 xn mod 2147483647x_{n+1}=16807\,x_n \bmod 2147483647, and use rn=2xn/2147483647−1r_n=2x_n/2147483647-1 for n=1,2,3,…n=1,2,3,\ldots Fill an 8×48\times4 matrix QQ row by row with r1,…,r32r_1,\ldots,r_{32}, then KK with r33,…,r64r_{33},\ldots,r_{64}, then VV with r65,…,r96r_{65},\ldots,r_{96}.

Compute causal scaled dot-product attention O=softmax⁡(QK⊤/dk+M)VO=\operatorname{softmax}(QK^\top/\sqrt{d_k}+M)V with dk=4d_k=4, where Mij=−∞M_{ij}=-\infty for j>ij>i and the softmax runs along each row. Give the sum of all 32 entries of OO to 6 decimal places.

A number, rounded to 6 decimal places

Problem 6·Challenge

When the cache pays a hundredfold

+50 XP

Model a decoder with L=24L=24 layers and width d=1024d=1024 by its multiply–adds. Processing the token at position tt (counting from 1) costs, in each layer, 12d212d^2 for the weight matrices plus 2td2td for attention over the tt visible positions. Ignore embeddings, norms and the softmax.

A prompt has P=1000P=1000 tokens. Generating GG new tokens needs logits at positions P,P+1,…,P+G−1P,P+1,\ldots,P+G-1. With a KV cache, every position from 1 to P+G−1P+G-1 is processed exactly once. Without a cache, producing the kk-th new token (k=1,…,Gk=1,\ldots,G) reprocesses all positions 1,…,P+k−11,\ldots,P+k-1 from scratch.

Find the smallest GG for which generation without the cache needs at least 100 times as many multiply–adds as generation with it.

An exact integer (or a fraction like 7/12)

Key takeaways

  • A language model is a classifier over the vocabulary at every position; the chain rule turns its predictions into the probability of a whole text.
  • Cross-entropy with teacher forcing trains every position in parallel. Perplexity, its exponential, is the effective number of choices, exactly ∣V∣|\mathcal V| for a uniform guesser.
  • Causal scaled dot-product attention averages earlier values with softmax weights: the mask makes future weights exactly zero, dk\sqrt{d_k} keeps scores from saturating, and query length acts as an inverse temperature.
  • A decoder-only transformer is embeddings, positional information, and residual blocks of attention and MLP, with about 12Ld212Ld^2 parameters in the blocks.
  • Generation is prefill then decode. The KV cache reuses earlier keys and values, so each new token costs about 2N2N FLOPs plus attention that grows with the context.

Checkpoint

Prove it to the labyrinth

Answer every question to clear this chamber. First-try answers earn the most XP.

0/9
Question 1 of 9 +20 XP

Under teacher forcing, what does the model condition on when it predicts the token at position tt during training?

Question 2 of 9 +20 XP

If the components of q,k∈R128q,k\in\mathbb R^{128} are independent with mean 0 and variance 1, what is the standard deviation of the unscaled score q⋅kq\cdot k?

Question 3 of 9 +20 XP

Why is it valid to cache the keys and values of earlier positions while decoding?

Question 4 of 9 +20 XP

You multiply a query vector by 3. What happens to that query's attention weights?

Question 5 of 9 +20 XP

A model's validation loss is 1.5 nats per token. How many bits per token is that?

Question 6 of 9 +20 XP

Two models report per-token perplexities of 9 and 12 on the same text, but they use different tokenisers. What can you conclude?

Question 7 of 9 +20 XP

Estimate GPT-2 small's parameters in millions with 12Ld2+∣V∣d12Ld^2+|\mathcal V|d, where L=12L=12, d=768d=768 and ∣V∣=50257|\mathcal V|=50257.

million
Question 8 of 9 +20 XP

Which statement about a causal attention output oio_i is always true?

Question 9 of 9 +20 XP

For one position with logits zz and target yy, what is the gradient of −log⁡softmax⁡(z)y-\log\operatorname{softmax}(z)_y with respect to zz?

End of the chamber

Clear this chamber

+60 XPAutoregressive FactorisationPerplexityScaled Dot-Product AttentionKV Cache