Every language model runs on somebody's hardware. Call a closed model through an API and the weights, the GPUs and the per-token price belong to the provider. Download open weights and the model becomes a file of numbers that you must fit into memory and stream through a chip once for every token you generate. The choice is partly about trust, control and capability, and partly plain arithmetic: bytes, bytes per second and money per month. The arithmetic starts with a rounding trick that stores each weight as a small integer plus a shared scale.
Spotted in the wild
- “N”Total parameters. Every one must be held in memory, including experts a token does not use.
- “N active”Parameters read per generated token: for a dense model, far fewer for a mixture of experts.
- “b”Bytes per stored parameter: 2 at 16-bit, 1 at 8-bit, about 0.56 at 4.5 bits per weight.
- “KV-cache memory”Bytes holding a key and a value for every layer, KV head, token and sequence.
- “delta, the step”Quantisation step: the gap between neighbouring representable values.
- “y rounded to the nearest integer”Round to nearest; it moves a number by at most one half.
- “beta”Memory bandwidth, in bytes per second.
- “F”Fixed monthly cost of running locally: amortised hardware, hosting, upkeep and people's time.
- “p and c”API price per token, and the local marginal cost per token (mostly electricity).
- “V star”Break-even monthly token volume.
| Symbol | Say it | Meaning | LaTeX |
|---|---|---|---|
| “N” | Total parameters. Every one must be held in memory, including experts a token does not use. | ||
| “N active” | Parameters read per generated token: for a dense model, far fewer for a mixture of experts. | ||
| “b” | Bytes per stored parameter: 2 at 16-bit, 1 at 8-bit, about 0.56 at 4.5 bits per weight. | ||
| “KV-cache memory” | Bytes holding a key and a value for every layer, KV head, token and sequence. | ||
| “delta, the step” | Quantisation step: the gap between neighbouring representable values. | ||
| “y rounded to the nearest integer” | Round to nearest; it moves a number by at most one half. | ||
| “beta” | Memory bandwidth, in bytes per second. | ||
| “F” | Fixed monthly cost of running locally: amortised hardware, hosting, upkeep and people's time. | ||
| “p and c” | API price per token, and the local marginal cost per token (mostly electricity). | ||
| “V star” | Break-even monthly token volume. |
Closed, open weights and open source
A closed (proprietary) model is reached only through its provider's API and apps. The weights never leave the provider, and its terms decide what you may send and do. Examples include OpenAI's GPT models, Anthropic's Claude, Google's Gemini and Amazon's Nova. An open-weights model publishes its trained parameters for download, so you can run it wherever you like: examples include Meta's Llama, Alibaba's Qwen, DeepSeek, Mistral's open models, Google's Gemma, Microsoft's Phi and OpenAI's gpt-oss. Licences range from permissive (Apache-2.0, MIT) to custom licences with acceptable-use rules or conditions for very large companies. Read the licence before you build on a model.
Open weights are not the same as open source AI. The Open Source Initiative's Open Source AI Definition 1.0 (October 2024) asks for the freedoms to use, study, modify and share, and for the preferred form for making modifications: the parameters, the complete code used to train and run the system, and enough information about the training data for a skilled person to build a substantially equivalent system. Most popular open-weights releases stop at the weights and inference code; fully open projects such as Ai2's OLMo go much further. For deployment the sharper question is simpler: can you hold the weights yourself?
Which release meets the OSI's Open Source AI Definition 1.0?
Which one, when
Neither side wins everywhere. Score your product on each row:
| Factor | Leans towards a closed API | Leans towards open weights you run |
|---|---|---|
| Capability | Your eval shows you need the strongest model available, and it is closed | Your eval shows an open model is good enough, as it often is for extraction, classification and routine code |
| Data sensitivity and residency | The provider's terms, regions and retention options pass your legal review | Data may not leave your network or jurisdiction |
| Latency and offline use | A network round trip is fine | You need on-device, air-gapped or offline use, or latency you control |
| Cost at your volume | Low or spiky traffic: you pay only for tokens used | High, steady traffic above the break-even volume (below) |
| Customisation | Prompts, retrieval and the provider's fine-tuning and structured outputs suffice | You need full fine-tuning, any sampler, raw logits or grammar-constrained decoding |
| Reproducibility | Dated snapshots and careful logging are enough | You hold the exact weights and runtime: no silent updates, so a pinned release stays pinned |
| Lock-in and deprecations | You accept migrating when a model is retired | You must keep a model for years, or switch vendors freely |
| Operational burden | No one on your team should run GPUs, drivers, scaling and on-call | You have, or will hire, the skills to keep a serving stack healthy |
There are middle grounds. Hosted inference providers serve open-weights models per token, so you can start without hardware and self-host later with the same model. Several clouds offer closed models inside your own cloud tenancy, under your region and enterprise terms. Dedicated deployments (reserved capacity at a fixed price) exist on both sides.
Many systems do not choose at all. A router sends each request to the cheapest model likely to handle it; a cascade tries a cheap model first and escalates when a scorer judges the answer unreliable. FrugalGPT (Chen, Zaharia and Zou, 2023) learned cascades over commercial APIs that matched the best single model's performance at a fraction of its cost on the tasks studied. Whatever you pick, decide with your own eval suite on your own data: a leaderboard measures someone else's task. A sensible default path is to prototype on a strong API to learn what quality is possible, build the eval suite, then test open models against it, and move traffic when one passes and the privacy, control or cost case holds.
Running a model yourself
- llama.cpp runs models stored in its GGUF format on CPUs, GPUs and Apple silicon. When a model is too big for the GPU it can keep some layers in system memory, at system-memory speed.
- Ollama and LM Studio wrap local engines in a model library, an app or command line, and a local OpenAI-compatible server. MLX is Apple's array framework for unified memory; its
mlx-lmpackage runs and fine-tunes models on a Mac. - For many users on GPUs, vLLM (PagedAttention: the KV cache lives in fixed-size blocks, like virtual-memory pages, so little memory is wasted) and SGLang (RadixAttention: requests share cached prompt prefixes) batch requests continuously. Hugging Face's TGI is now in maintenance mode, and its documentation points new users to vLLM, SGLang, llama.cpp or MLX.
# One machine: a 4-bit GGUF file served on localhost:8080
llama-server -m model-Q4_K_M.gguf -c 8192 --port 8080
# A GPU server for many users, with an OpenAI-compatible API on port 8000
vllm serve <org>/<model> --max-model-len 8192 --gpu-memory-utilization 0.90
Both speak the same chat-completions protocol as most hosted APIs, so moving between a provider and your own server can be a configuration change. That portability is your best defence against lock-in.
| Format | How it quantises | Where it runs best |
|---|---|---|
| GGUF (for example Q8_0, Q4_K_M) | Blocks of weights share scales; bit widths can vary by tensor | llama.cpp, Ollama and LM Studio, on CPUs, Apple silicon and GPUs |
| GPTQ | One pass, layer by layer, using approximate second-order information to compensate for rounding | GPU inference at 3–4 bits |
| AWQ | Scales up the weight channels that meet large activations before rounding, protecting the most salient weights | GPU inference at 4 bits |
As a rough guide from the GPTQ and AWQ papers and community testing, 8-bit is close to lossless, 4-bit costs a little quality (more for small models and for demanding work such as maths and code), and 3 bits or fewer degrade quickly. Treat that as a prior and measure on your eval.
Memory: weights, cache and quantisation
A model with parameters at bytes each needs bytes for weights. A 70B model takes 140 GB at 16-bit (), 70 GB at 8-bit, and about 39 GB in a typical 4-bit format, which spends about 4.5 bits per weight once block scales are counted.
The Glimpse shows absmax quantisation. For a block , choose the step , store as int8 with one scale per block, and reconstruct . Since , no value overflows, and rounding moves each value by at most half a step: . For , , and the worst error is . The bound also shows why outliers hurt: one huge value inflates for everything sharing its scale. Hence small blocks (tens of weights in GGUF's formats) and LLM.int8()'s trick of keeping rare outlier features in 16-bit.
Weights are not the only tenant. As Next-Token Prediction showed, decoding keeps a key and a value per layer and KV head for every token seen, so for sequences of tokens
An 8B model with 32 layers, 8 KV heads and stores bytes per token at 16-bit. At 64k tokens that is 8.6 GB, nearly twice the 4-bit weights. Add an allowance for activations and buffers (the lab uses 10%) and compare with the device.
A model has 32 layers, 8 KV heads and head dimension 128, with a 16-bit KV cache. How many GB ( bytes) does the cache need for one sequence of 32,768 tokens?
Speed: decoding is memory-bound
To produce the next token of every sequence, a decoding step must read each active weight once, plus the KV cache. Memory delivers at most bytes per second, so a step takes at least (bytes read) seconds, whatever the arithmetic units do. Each step yields one token per sequence, so
An 8B model at 4.5 bits (4.5 GB) on a 1,008 GB/s card is capped near 224 tokens/s for one user with a short context; on a 288 GB/s card, near 64. A 70B model at 4.5 bits on a 546 GB/s laptop: about 14. The arithmetic is not the bottleneck. A token costs about 2 FLOPs per active weight, while an 80 GB datacentre card performs roughly 990 dense 16-bit TFLOP/s against 3.35 TB/s of bandwidth, about 300 FLOPs per byte read.
Batching closes that gap: one read of the weights serves all sequences, so throughput grows almost linearly in until compute or the growing KV cache becomes the limit. That is why busy servers make tokens far more cheaply than a laptop serving one person, and why vLLM and SGLang fight so hard for KV-cache memory. Prefill, which processes the whole prompt at once, already reuses each weight across many tokens and is usually compute-bound.
A mixture of experts splits the two numbers. Memory follows the total parameters, because every expert must be resident; single-stream speed follows the active ones. A model with 47B total and 13B active parameters needs the memory of a 47B model and decodes like a 13B one. (At larger batches, different tokens choose different experts, so more weights are read per step.)
A dense 8B model is stored at 8 bits per weight on a GPU with 1,000 GB/s of memory bandwidth. Ignoring the KV cache, what is the upper bound on single-sequence decoding speed, in tokens per second?
Cost: the break-even volume
At tokens a month, an API costs . Running locally costs : the fixed part is the hardware price divided by its useful life in months, plus hosting, maintenance and people's time; the marginal part is mostly electricity, power draw times price divided by tokens per hour. Local is cheaper exactly when , that is, when . If , divide by the positive : owning wins exactly above the break-even volume
If , the right-hand side is never positive, so there is no break-even at any volume.
A worked example, every number illustrative. A USD 8,000 machine over 36 months gives USD a month. Drawing 450 W at 0.25 USD per kWh while generating 200 tokens per second gives USD per million tokens. Against an API at 1.00 USD per million, million tokens a month, about 100 tokens every second around the clock. The machine makes at most 518 million tokens in a 30-day month, so it breaks even at about 51% utilisation. At 0.30 USD per million, would be about 1.5 billion: more than this machine can produce at all.
Be honest about and . One engineer-day a month of upkeep can easily triple here. Uptime may need a second machine. Idle power is real. If the local model fails more often, count the cost of those failures. And moves every time API prices fall. Low or uncertain volume favours the API; high, steady volume with a good-enough open model and the skills to run it favours owning.
Running locally costs a fixed USD 300 a month plus USD 0.50 per million tokens. An API charges USD 2.00 per million tokens. What is the break-even volume, in millions of tokens per month?
Read beyond
Article · free online · ~15 min
The Open Source AI Definition 1.0Open Source Initiative · The four freedoms and the preferred form for modifications
List what it requires beyond the weights, then check one open-weights model you use against it.
Article · free online · ~30 min
Transformer Inference Arithmetickipply · KV cache, capacity, latency calculations and batch sizes
Redo its KV-cache and memory-bound latency estimates for the 70B shape in this chamber's lab.
Book · free online · ~45 min
How to Scale Your ModelAustin, Douglas, Pope et al. · All About Transformer Inference
Find the batch size at which it says generation stops being memory-bound, and redo that estimate for one of the lab's GPUs.
Article · free online · ~20 min
A Gentle Introduction to 8-bit Matrix Multiplication for transformers at scaleYounes Belkada & Tim Dettmers · Absmax and zero-point quantisation, then LLM.int8()
Quantise the post's example vector by hand and check every error against the half-step bound.
Tool · free online · ~30 min
llama.cppggml-org · README: quick start and supported backends
Run a small quantised model and compare your measured tokens per second with the bandwidth bound for your machine.
Read the equation in context
LLM.int8(): 8-bit Matrix Multiplication for Transformers at ScaleTim Dettmers, Mike Lewis, Younes Belkada & Luke Zettlemoyer · NeurIPS 2022, 2022Section 2.1 defines absmax quantisation with one constant for a whole tensor. Section 3.1 gives every row of the hidden states and every column of the weights its own constant (vector-wise quantisation). That is not enough at scale: beyond about 6.7B parameters, a few feature dimensions with huge magnitudes appear in every layer. Section 3.2 keeps those, about 0.1% of values, in 16-bit and multiplies the rest in 8-bit, which preserved 16-bit performance up to 175B parameters while roughly halving memory.
Decode the paper · Section 2.1: absmax quantisation (unnumbered equation)
LLM.int8(): 8-bit Matrix Multiplication for Transformers at ScaleTim Dettmers, Mike Lewis, Younes Belkada & Luke Zettlemoyer · NeurIPS 2022, 2022
Section 2.1 defines absmax quantisation with one constant for a whole tensor. Section 3.1 gives every row of the hidden states and every column of the weights its own constant (vector-wise quantisation, Eq. 7), and Section 3.2 keeps the rare large-magnitude feature dimensions, about 0.1% of values, in 16-bit. Together these make LLM.int8().
Options
The quantisation and serving methods you will meet in every local stack come from four more papers:
GPTQ: Accurate Post-Training Quantization for Generative Pre-trained TransformersElias Frantar, Saleh Ashkboos, Torsten Hoefler & Dan Alistarh · ICLR 2023, 2023 AWQ: Activation-aware Weight Quantization for LLM Compression and AccelerationJi Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan & Song Han · MLSys 2024, 2024 Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang & Ion Stoica · SOSP 2023, 2023 Efficiently Scaling Transformer InferenceReiner Pope, Sholto Douglas, Aakanksha Chowdhery, Jacob Devlin, James Bradbury, Anselm Levskaya, Jonathan Heek, Kefan Xiao, Shivani Agrawal & Jeff Dean · MLSys 2023, 2022GPTQ compensates each rounding error using second-order information; AWQ protects the roughly 1% of salient weight channels identified from activations. vLLM's PagedAttention cut KV-cache waste and raised throughput 2–4× over earlier systems at the same latency. Pope et al. model inference as compute time, memory time and communication between chips: every decoding step loads the parameters and KV cache from memory, small batches leave compute idle, and int8 weights save memory time in exactly that regime. They report 29 ms per token for small-batch generation on PaLM 540B.
Your turn
Plan three deployments. Fit each model onto its device by choosing a precision, watching the weights, KV cache and overhead fill the memory bar, and read off how fast bandwidth lets it decode. Then price the result: slide the monthly volume until owning costs the same as paying per token.
Interactive lab
Deployment planner
1 · Fit the model
Model
Shaped like Llama 3 8B: 32 layers, 8 KV heads, head dimension 128.
Device
24 GB at 1008 GB/s, as published for the RTX 4090.
● Weights 16.0 GB● KV cache 1.1 GB● Overhead 1.7 GB| Device memory
Needed
18.8 GB
Available
24 GB
Per sequence ≤
59 tok/s
Whole batch ≤
59 tok/s
2 · Pay for it
API bill
USD 100
Local bill
USD 238
Fixed F
USD 222
Local c per M
USD 0.16
Max per month
518M
Utilisation
19%
Match · Situation ↔ Sensible starting point
Which way to lean
Options
Match · Quantity ↔ Formula
The arithmetic
Options
Match · Tool ↔ What it is
The local toolbox
Options
Proof puzzle
When owning pays
Claim
An API costs for tokens a month; running locally costs with . Show that if , local is cheaper exactly when , and that if it is never cheaper.
Tap lines in the order they should appear. Not every line belongs. Tap a line in your proof to send it back.
Your proof
- Pick the first line below.
Available lines
Prove it yourself
Half a step at most
Claim
For a block with , let , and . Prove that every lies in and that .
Your typeset proof appears here.
Coding problems
Problem 25·Warm-up
How many users fit?
A dense model with parameters is stored at 4.5 bits per weight. It has 80 layers, 8 KV heads and head dimension 128, and keeps its KV cache at 16-bit (2 bytes per number). Each sequence holds 8,192 tokens. The memory needed is (weights + KV cache), the extra 10% covering activations and buffers. What is the largest number of concurrent sequences that fits in bytes?
Problem 26·Standard
Outliers and block scales
Generate 256 weights with the MINSTD generator: , , and for . Then set the first weight to , an outlier. Quantise to int8 with absmax in two ways: (a) one step for the whole vector; (b) one step for each block of 32 consecutive weights. Each step is over the weights sharing it, with (no ties occur) and . Print the mean squared error of (a) divided by the mean squared error of (b), to 2 decimal places.
Problem 27·Challenge
A frugal cascade
1,000 queries arrive. Using MINSTD with ( and after each step), draw three numbers per query, in order: , , . A small local model answers query with confidence and is correct when . A large API model is correct when . A cascade with threshold always runs the local model (cost 1 unit) and, when , also calls the API (10 more units) and uses its answer instead. Among for , consider the cascades with at least 900 correct answers. Print the lowest total cost.
Key takeaways
- Closed APIs buy capability and zero operations; open weights buy control, privacy, reproducibility and freedom from forced migrations. Open weights are not the same as open source AI.
- Memory is weights () plus KV cache () plus overhead; quantisation shrinks the first term with an error of at most half a step.
- Single-stream decoding is bounded by bandwidth over bytes read per token; batching shares the reads, and a mixture of experts needs memory for its total parameters but reads only the active ones.
- Owning breaks even at only when the local marginal cost is below the API price and one machine can actually produce that volume.
- Route or cascade between models, and decide with your own evals rather than leaderboards.
Checkpoint
Prove it to the labyrinth
Answer every question to clear this chamber. First-try answers earn the most XP.
A mixture-of-experts model has 47B parameters in total and 13B active per token. Compared with a dense 13B model at the same precision, it needs…
How many GB ( bytes) of weight memory does a 70B-parameter model need at 4.5 bits per weight?
You must be able to reproduce last year's outputs exactly. Which setup makes that most achievable?
Why does batching raise decoding throughput on a GPU?
The local marginal cost per token is above the API price per token. What is the break-even volume?
A cascade answers every query with a cheap model first and escalates to a stronger one when a scorer is unsure. When does it save money without losing much accuracy?
Absmax int8 quantisation is applied to a block whose largest magnitude is 2.54. After dequantising, what is the largest possible rounding error for any element?
Which file format packages quantised weights for llama.cpp and tools such as Ollama and LM Studio?
End of the chamber
Clear this chamber
- Questions in this chamber (0/12 solved)Next unsolved
- Bonus: Deployment planner (+50 XP)
- Bonus: Problem 25: How many users fit? (+20 XP)
- Bonus: Problem 26: Outliers and block scales (+35 XP)
- Bonus: Problem 27: A frugal cascade (+50 XP)
- Bonus: Proof: When owning pays (+25 XP)
- Bonus: Proof: Half a step at most (+35 XP)
- Bonus: Decode the paper (+25 XP)
- Bonus: Match: Which way to lean (+20 XP)
- Bonus: Match: The arithmetic (+20 XP)
- Bonus: Match: The local toolbox (+20 XP)