Course · Intermediate
Large Language Models: From Tokens to Production
How LLMs work, how to steer them, and how to ship them without fooling yourself.
A large language model predicts the next token, and almost everything else follows from that. This course starts inside the model, with tokens, embeddings, attention and the training pipeline from pretraining to preference optimisation. Then it takes the controls: temperature and nucleus sampling, why the same prompt can give different answers even at temperature zero, and how to make that randomness reproducible or put it to work. The inner ring is engineering: prompts written and versioned like code, eval suites with honest error bars, retrieval and tool-using agents, prompt injection, and the choice between a closed API and open weights on your own hardware. Each chamber runs the research loop, from discovery to derivation to real papers such as Attention Is All You Need, Chinchilla, nucleus sampling, self-consistency and retrieval-augmented generation, with proofs to write and problems to code. The Sirens wait at the centre.
- Chambers
- 9 + boss
- Time
- ~11 hours
- Questions
- 109
- Coding problems
- 27
- Concept cards
- 37
- XP available
- 5,590+
The path through the labyrinth
Three rings, 9 chambers, one guardian. You can wander ahead, but the thread works best in order: each chamber builds on the last.
The Outer Ring
How a language model works- 1Tokens and Embeddings: How Text Becomes NumbersYou are hereByte-pair encoding, vocabularies and the token tax, then the lookup table that turns every token into a vector. 55 min 50 XP12 questions
- 2Next-Token Prediction and the TransformerOne distribution per position: logits and softmax, cross-entropy and perplexity, causal self-attention and the KV cache. 55 min 60 XP14 questions
- 3From Pretraining to AssistantCompute budgets and scaling laws, then instruction tuning, preference optimisation and low-rank fine-tuning. 55 min 60 XP12 questions
The Middle Ring
Steering the output- 4Sampling: Temperature, Top-k and Top-pFrom logits to words: greedy and beam search, temperature, top-k, nucleus and min-p sampling, and when to use each. 50 min 60 XP11 questions
- 5Non-Determinism: Taming and Using RandomnessWhy temperature zero still varies, how to make runs reproducible, how to get variety on purpose, and how to test a system that never answers the same way twice. 60 min 60 XP11 questions
- 6Prompting as ProgrammingSystem prompts, examples, reasoning and structured output: write prompts like interfaces, and make the model speak valid JSON. 55 min 60 XP12 questions
The Inner Ring
Language models in production- 7Prompt Version Control and EvalsTreat prompts as code: version, review, pin and roll them back, and decide every release with an eval suite and honest error bars. 60 min 60 XP12 questions
- 8Retrieval, Tools and AgentsGround answers in retrieved documents, let models call tools, close the agent loop, and defend it against prompt injection. 55 min 60 XP13 questions
- 9Open Weights or API? Choosing and Running ModelsWhen to call a closed model, when to run open weights yourself, and the memory, speed and cost arithmetic behind the choice. 60 min 60 XP12 questions