Skip to content
AriadneTechnology

The Inner Ring · Chamber 9 of 9

Open Weights or API? Choosing and Running Models

When to call a closed model, when to run open weights yourself, and the memory, speed and cost arithmetic behind the choice.

60 min 60 XP + 12 questions + 1 challengeMathVideoPapersProofsCodeLab

In this chamber you will

  • Weigh closed APIs against open weights on capability, privacy, cost, control and licence
  • Compute the memory a model needs: weights, KV cache and quantisation
  • Estimate decoding speed from memory bandwidth
  • Find the break-even volume between paying per token and owning hardware, and design a hybrid router
DiscoverLearnRead beyondPapers & lecturesYour turn

Every language model runs on somebody's hardware. Call a closed model through an API and the weights, the GPUs and the per-token price belong to the provider. Download open weights and the model becomes a file of numbers that you must fit into memory and stream through a chip once for every token you generate. The choice is partly about trust, control and capability, and partly plain arithmetic: bytes, bytes per second and money per month. The arithmetic starts with a rounding trick that stores each weight as a small integer plus a shared scale.

Spotted in the wild

Xi8=⌊127⋅Xf16max⁡ij(∣Xf16ij∣)⌉=⌊127∥Xf16∥∞Xf16⌉=⌊sxf16Xf16⌉\mathbf{X}_{i8}=\left\lfloor\frac{127\cdot\mathbf{X}_{f16}}{\max_{ij}\left(|\mathbf{X}_{f16_{ij}}|\right)}\right\rceil=\left\lfloor\frac{127}{\|\mathbf{X}_{f16}\|_\infty}\mathbf{X}_{f16}\right\rceil=\left\lfloor s_{x_{f16}}\mathbf{X}_{f16}\right\rceil
LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
DiscoverLearnRead beyondPapers & lecturesYour turn
Symbols for this chamber
  • NN“N”
    Total parameters. Every one must be held in memory, including experts a token does not use.
  • NactN_{\text{act}}“N active”
    Parameters read per generated token: NN for a dense model, far fewer for a mixture of experts.
  • bb“b”
    Bytes per stored parameter: 2 at 16-bit, 1 at 8-bit, about 0.56 at 4.5 bits per weight.
    MW=NbM_W=Nb
  • MKVM_{KV}“KV-cache memory”
    Bytes holding a key and a value for every layer, KV head, token and sequence.
    MKV=2LhkvdhTBbkvM_{KV}=2Lh_{kv}d_hTBb_{kv}
  • Δ\Delta“delta, the step”
    Quantisation step: the gap between neighbouring representable values.
    Δ=∥x∥∞/127\Delta=\|\mathbf x\|_\infty/127
  • ⌊y⌉\lfloor y\rceil“y rounded to the nearest integer”
    Round to nearest; it moves a number by at most one half.
  • β\beta“beta”
    Memory bandwidth, in bytes per second.
  • FF“F”
    Fixed monthly cost of running locally: amortised hardware, hosting, upkeep and people's time.
  • p, cp,\ c“p and c”
    API price per token, and the local marginal cost per token (mostly electricity).
  • V∗V^*“V star”
    Break-even monthly token volume.
    V∗=F/(p−c)V^*=F/(p-c)

Closed, open weights and open source

A closed (proprietary) model is reached only through its provider's API and apps. The weights never leave the provider, and its terms decide what you may send and do. Examples include OpenAI's GPT models, Anthropic's Claude, Google's Gemini and Amazon's Nova. An open-weights model publishes its trained parameters for download, so you can run it wherever you like: examples include Meta's Llama, Alibaba's Qwen, DeepSeek, Mistral's open models, Google's Gemma, Microsoft's Phi and OpenAI's gpt-oss. Licences range from permissive (Apache-2.0, MIT) to custom licences with acceptable-use rules or conditions for very large companies. Read the licence before you build on a model.

Open weights are not the same as open source AI. The Open Source Initiative's Open Source AI Definition 1.0 (October 2024) asks for the freedoms to use, study, modify and share, and for the preferred form for making modifications: the parameters, the complete code used to train and run the system, and enough information about the training data for a skilled person to build a substantially equivalent system. Most popular open-weights releases stop at the weights and inference code; fully open projects such as Ai2's OLMo go much further. For deployment the sharper question is simpler: can you hold the weights yourself?

Quick check +20 XP

Which release meets the OSI's Open Source AI Definition 1.0?

Which one, when

Neither side wins everywhere. Score your product on each row:

FactorLeans towards a closed APILeans towards open weights you run
CapabilityYour eval shows you need the strongest model available, and it is closedYour eval shows an open model is good enough, as it often is for extraction, classification and routine code
Data sensitivity and residencyThe provider's terms, regions and retention options pass your legal reviewData may not leave your network or jurisdiction
Latency and offline useA network round trip is fineYou need on-device, air-gapped or offline use, or latency you control
Cost at your volumeLow or spiky traffic: you pay only for tokens usedHigh, steady traffic above the break-even volume (below)
CustomisationPrompts, retrieval and the provider's fine-tuning and structured outputs sufficeYou need full fine-tuning, any sampler, raw logits or grammar-constrained decoding
ReproducibilityDated snapshots and careful logging are enoughYou hold the exact weights and runtime: no silent updates, so a pinned release stays pinned
Lock-in and deprecationsYou accept migrating when a model is retiredYou must keep a model for years, or switch vendors freely
Operational burdenNo one on your team should run GPUs, drivers, scaling and on-callYou have, or will hire, the skills to keep a serving stack healthy

There are middle grounds. Hosted inference providers serve open-weights models per token, so you can start without hardware and self-host later with the same model. Several clouds offer closed models inside your own cloud tenancy, under your region and enterprise terms. Dedicated deployments (reserved capacity at a fixed price) exist on both sides.

Many systems do not choose at all. A router sends each request to the cheapest model likely to handle it; a cascade tries a cheap model first and escalates when a scorer judges the answer unreliable. FrugalGPT (Chen, Zaharia and Zou, 2023) learned cascades over commercial APIs that matched the best single model's performance at a fraction of its cost on the tasks studied. Whatever you pick, decide with your own eval suite on your own data: a leaderboard measures someone else's task. A sensible default path is to prototype on a strong API to learn what quality is possible, build the eval suite, then test open models against it, and move traffic when one passes and the privacy, control or cost case holds.

Running a model yourself

  • llama.cpp runs models stored in its GGUF format on CPUs, GPUs and Apple silicon. When a model is too big for the GPU it can keep some layers in system memory, at system-memory speed.
  • Ollama and LM Studio wrap local engines in a model library, an app or command line, and a local OpenAI-compatible server. MLX is Apple's array framework for unified memory; its mlx-lm package runs and fine-tunes models on a Mac.
  • For many users on GPUs, vLLM (PagedAttention: the KV cache lives in fixed-size blocks, like virtual-memory pages, so little memory is wasted) and SGLang (RadixAttention: requests share cached prompt prefixes) batch requests continuously. Hugging Face's TGI is now in maintenance mode, and its documentation points new users to vLLM, SGLang, llama.cpp or MLX.
Shell
# One machine: a 4-bit GGUF file served on localhost:8080
llama-server -m model-Q4_K_M.gguf -c 8192 --port 8080

# A GPU server for many users, with an OpenAI-compatible API on port 8000
vllm serve <org>/<model> --max-model-len 8192 --gpu-memory-utilization 0.90

Both speak the same chat-completions protocol as most hosted APIs, so moving between a provider and your own server can be a configuration change. That portability is your best defence against lock-in.

FormatHow it quantisesWhere it runs best
GGUF (for example Q8_0, Q4_K_M)Blocks of weights share scales; bit widths can vary by tensorllama.cpp, Ollama and LM Studio, on CPUs, Apple silicon and GPUs
GPTQOne pass, layer by layer, using approximate second-order information to compensate for roundingGPU inference at 3–4 bits
AWQScales up the weight channels that meet large activations before rounding, protecting the most salient weightsGPU inference at 4 bits

As a rough guide from the GPTQ and AWQ papers and community testing, 8-bit is close to lossless, 4-bit costs a little quality (more for small models and for demanding work such as maths and code), and 3 bits or fewer degrade quickly. Treat that as a prior and measure on your eval.

host ALL your AI locallyNetworkChuck · 24 min

Memory: weights, cache and quantisation

A model with NN parameters at bb bytes each needs MW=NbM_W=Nb bytes for weights. A 70B model takes 140 GB at 16-bit (b=2b=2), 70 GB at 8-bit, and about 39 GB in a typical 4-bit format, which spends about 4.5 bits per weight once block scales are counted.

The Glimpse shows absmax quantisation. For a block x\mathbf x, choose the step Δ=∥x∥∞/127\Delta=\|\mathbf x\|_\infty/127, store qi=⌊xi/Δ⌉q_i=\lfloor x_i/\Delta\rceil as int8 with one scale per block, and reconstruct x^i=Δqi\hat x_i=\Delta q_i. Since ∣xi/Δ∣≤127|x_i/\Delta|\le127, no value overflows, and rounding moves each value by at most half a step: ∣xi−x^i∣≤Δ/2|x_i-\hat x_i|\le\Delta/2. For x=(0.3,−2.54,0.017)\mathbf x=(0.3,-2.54,0.017), Δ=0.02\Delta=0.02, q=(15,−127,1)q=(15,-127,1) and the worst error is 0.003≤0.010.003\le0.01. The bound also shows why outliers hurt: one huge value inflates Δ\Delta for everything sharing its scale. Hence small blocks (tens of weights in GGUF's formats) and LLM.int8()'s trick of keeping rare outlier features in 16-bit.

Weights are not the only tenant. As Next-Token Prediction showed, decoding keeps a key and a value per layer and KV head for every token seen, so for BB sequences of TT tokens

MKV=2 L hkv dh T B bkv.M_{KV}=2\,L\,h_{kv}\,d_h\,T\,B\,b_{kv}.

An 8B model with 32 layers, 8 KV heads and dh=128d_h=128 stores 2⋅32⋅8⋅128⋅2=131,0722\cdot32\cdot8\cdot128\cdot2=131{,}072 bytes per token at 16-bit. At 64k tokens that is 8.6 GB, nearly twice the 4-bit weights. Add an allowance for activations and buffers (the lab uses 10%) and compare with the device.

Quick check +20 XP

A model has 32 layers, 8 KV heads and head dimension 128, with a 16-bit KV cache. How many GB (10910^9 bytes) does the cache need for one sequence of 32,768 tokens?

Speed: decoding is memory-bound

To produce the next token of every sequence, a decoding step must read each active weight once, plus the KV cache. Memory delivers at most β\beta bytes per second, so a step takes at least (bytes read)/β/\beta seconds, whatever the arithmetic units do. Each step yields one token per sequence, so

tokens/s per sequence≤βNact b+MKV.\text{tokens/s per sequence}\le\frac{\beta}{N_{\text{act}}\,b+M_{KV}}.

An 8B model at 4.5 bits (4.5 GB) on a 1,008 GB/s card is capped near 224 tokens/s for one user with a short context; on a 288 GB/s card, near 64. A 70B model at 4.5 bits on a 546 GB/s laptop: about 14. The arithmetic is not the bottleneck. A token costs about 2 FLOPs per active weight, while an 80 GB datacentre card performs roughly 990 dense 16-bit TFLOP/s against 3.35 TB/s of bandwidth, about 300 FLOPs per byte read.

Batching closes that gap: one read of the weights serves all BB sequences, so throughput grows almost linearly in BB until compute or the growing KV cache becomes the limit. That is why busy servers make tokens far more cheaply than a laptop serving one person, and why vLLM and SGLang fight so hard for KV-cache memory. Prefill, which processes the whole prompt at once, already reuses each weight across many tokens and is usually compute-bound.

A mixture of experts splits the two numbers. Memory follows the total parameters, because every expert must be resident; single-stream speed follows the active ones. A model with 47B total and 13B active parameters needs the memory of a 47B model and decodes like a 13B one. (At larger batches, different tokens choose different experts, so more weights are read per step.)

Quick check +20 XP

A dense 8B model is stored at 8 bits per weight on a GPU with 1,000 GB/s of memory bandwidth. Ignoring the KV cache, what is the upper bound on single-sequence decoding speed, in tokens per second?

Cost: the break-even volume

At VV tokens a month, an API costs pVpV. Running locally costs F+cVF+cV: the fixed part FF is the hardware price divided by its useful life in months, plus hosting, maintenance and people's time; the marginal part cc is mostly electricity, power draw times price divided by tokens per hour. Local is cheaper exactly when F+cV<pVF+cV\lt pV, that is, when F<(p−c)VF\lt(p-c)V. If p>cp\gt c, divide by the positive p−cp-c: owning wins exactly above the break-even volume

V∗=Fp−c.V^*=\frac{F}{p-c}.

If c≥pc\ge p, the right-hand side (p−c)V(p-c)V is never positive, so there is no break-even at any volume.

A worked example, every number illustrative. A USD 8,000 machine over 36 months gives F≈222F\approx222 USD a month. Drawing 450 W at 0.25 USD per kWh while generating 200 tokens per second gives c≈0.16c\approx0.16 USD per million tokens. Against an API at 1.00 USD per million, V∗=222.2/(0.844×10−6)≈263V^*=222.2/(0.844\times10^{-6})\approx263 million tokens a month, about 100 tokens every second around the clock. The machine makes at most 518 million tokens in a 30-day month, so it breaks even at about 51% utilisation. At 0.30 USD per million, V∗V^* would be about 1.5 billion: more than this machine can produce at all.

Be honest about FF and cc. One engineer-day a month of upkeep can easily triple FF here. Uptime may need a second machine. Idle power is real. If the local model fails more often, count the cost of those failures. And V∗V^* moves every time API prices fall. Low or uncertain volume favours the API; high, steady volume with a good-enough open model and the skills to run it favours owning.

Quick check +20 XP

Running locally costs a fixed USD 300 a month plus USD 0.50 per million tokens. An API charges USD 2.00 per million tokens. What is the break-even volume, in millions of tokens per month?

DiscoverLearnRead beyondPapers & lecturesYour turn

Read beyond

Article · free online · ~15 min

The Open Source AI Definition 1.0

Open Source Initiative · The four freedoms and the preferred form for modifications

List what it requires beyond the weights, then check one open-weights model you use against it.

Article · free online · ~30 min

Transformer Inference Arithmetic

kipply · KV cache, capacity, latency calculations and batch sizes

Redo its KV-cache and memory-bound latency estimates for the 70B shape in this chamber's lab.

Book · free online · ~45 min

How to Scale Your Model

Austin, Douglas, Pope et al. · All About Transformer Inference

Find the batch size at which it says generation stops being memory-bound, and redo that estimate for one of the lab's GPUs.

Article · free online · ~20 min

A Gentle Introduction to 8-bit Matrix Multiplication for transformers at scale

Younes Belkada & Tim Dettmers · Absmax and zero-point quantisation, then LLM.int8()

Quantise the post's example vector by hand and check every error against the half-step bound.

Tool · free online · ~30 min

llama.cpp

ggml-org · README: quick start and supported backends

Run a small quantised model and compare your measured tokens per second with the bandwidth bound for your machine.

DiscoverLearnRead beyondPapers & lecturesYour turn

Read the equation in context

LLM.int8(): 8-bit Matrix Multiplication for Transformers at ScaleTim Dettmers, Mike Lewis, Younes Belkada & Luke Zettlemoyer · NeurIPS 2022, 2022

Section 2.1 defines absmax quantisation with one constant for a whole tensor. Section 3.1 gives every row of the hidden states and every column of the weights its own constant (vector-wise quantisation). That is not enough at scale: beyond about 6.7B parameters, a few feature dimensions with huge magnitudes appear in every layer. Section 3.2 keeps those, about 0.1% of values, in 16-bit and multiplies the rest in 8-bit, which preserved 16-bit performance up to 175B parameters while roughly halving memory.

Decode the paper · Section 2.1: absmax quantisation (unnumbered equation)

LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale

Tim Dettmers, Mike Lewis, Younes Belkada & Luke Zettlemoyer · NeurIPS 2022, 2022

+25 XP
Xi8=⌊127⋅Xf16max⁡ij(∣Xf16ij∣)⌉=⌊127∥Xf16∥∞Xf16⌉=⌊sxf16Xf16⌉\mathbf{X}_{i8}=\left\lfloor\frac{127\cdot\mathbf{X}_{f16}}{\max_{ij}\left(|\mathbf{X}_{f16_{ij}}|\right)}\right\rceil=\left\lfloor\frac{127}{\|\mathbf{X}_{f16}\|_\infty}\mathbf{X}_{f16}\right\rceil=\left\lfloor s_{x_{f16}}\mathbf{X}_{f16}\right\rceil

Section 2.1 defines absmax quantisation with one constant for a whole tensor. Section 3.1 gives every row of the hidden states and every column of the weights its own constant (vector-wise quantisation, Eq. 7), and Section 3.2 keeps the rare large-magnitude feature dimensions, about 0.1% of values, in 16-bit. Together these make LLM.int8().

Xf16\mathbf{X}_{f16}
Xi8\mathbf{X}_{i8}
127127
∥Xf16∥∞\|\mathbf{X}_{f16}\|_\infty
sxf16s_{x_{f16}}
⌊⋅⌉\lfloor\cdot\rceil

Options

The quantisation and serving methods you will meet in every local stack come from four more papers:

GPTQ: Accurate Post-Training Quantization for Generative Pre-trained TransformersElias Frantar, Saleh Ashkboos, Torsten Hoefler & Dan Alistarh · ICLR 2023, 2023 AWQ: Activation-aware Weight Quantization for LLM Compression and AccelerationJi Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan & Song Han · MLSys 2024, 2024 Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang & Ion Stoica · SOSP 2023, 2023 Efficiently Scaling Transformer InferenceReiner Pope, Sholto Douglas, Aakanksha Chowdhery, Jacob Devlin, James Bradbury, Anselm Levskaya, Jonathan Heek, Kefan Xiao, Shivani Agrawal & Jeff Dean · MLSys 2023, 2022

GPTQ compensates each rounding error using second-order information; AWQ protects the roughly 1% of salient weight channels identified from activations. vLLM's PagedAttention cut KV-cache waste and raised throughput 2–4× over earlier systems at the same latency. Pope et al. model inference as compute time, memory time and communication between chips: every decoding step loads the parameters and KV cache from memory, small batches leave compute idle, and int8 weights save memory time in exactly that regime. They report 29 ms per token for small-batch generation on PaLM 540B.

Which Quantization Method is Right for You? (GPTQ vs. GGUF vs. AWQ)Maarten Grootendorst · 16 min
Fast LLM Serving with vLLM and PagedAttentionAnyscale · 32 min
DiscoverLearnRead beyondPapers & lecturesYour turn

Your turn

Plan three deployments. Fit each model onto its device by choosing a precision, watching the weights, KV cache and overhead fill the memory bar, and read off how fast bandwidth lets it decode. Then price the result: slide the monthly volume until owning costs the same as paying per token.

Interactive lab

Deployment planner

First fit each task's model into its device's memory and read off the bandwidth bound on decoding speed. Then price the setup: find the monthly token volume at which owning it costs the same as paying per token. Sizes use 1 GB = 10⁹ bytes and the KV cache is kept at 16-bit.

1 · Fit the model

Model

Shaped like Llama 3 8B: 32 layers, 8 KV heads, head dimension 128.

Device

24 GB at 1008 GB/s, as published for the RTX 4090.

Needed

18.8 GB

Available

24 GB

Per sequence ≤

59 tok/s

Whole batch ≤

59 tok/s

2 · Pay for it

Monthly cost of the API and of running locally against monthly token volume00130136259272389408518544Million tokens per monthUSD per month
● API: pV● Local: F + cV● One machine's capacity● Your volume

API bill

USD 100

Local bill

USD 238

Fixed F

USD 222

Local c per M

USD 0.16

Max per month

518M

Utilisation

19%

Challenge: Deployment plannerFit three models onto their GPUs, then find the monthly volume at which running locally breaks even.+50 XP

Match · Situation ↔ Sensible starting point

Which way to lean

+20 XP
Patient records that may not leave your network
A hard, low-volume task where only the strongest model passes your eval
Millions of easy classification calls a day, steady all year
Spiky traffic, no one to run servers, and an open model that passes your eval
Most queries are easy but a few are hard

Options

Match · Quantity ↔ Formula

The arithmetic

+20 XP
Weight memory
KV-cache memory
Decode bound per sequence
Absmax step
Break-even volume

Options

Match · Tool ↔ What it is

The local toolbox

+20 XP
llama.cpp
Ollama
MLX
vLLM
SGLang
AWQ

Options

Proof puzzle

When owning pays

+25 XP

Claim

An API costs pVpV for VV tokens a month; running locally costs F+cVF+cV with F>0F>0. Show that if p>cp>c, local is cheaper exactly when V>V∗=F/(p−c)V>V^*=F/(p-c), and that if p≤cp\le c it is never cheaper.

Tap lines in the order they should appear. Not every line belongs. Tap a line in your proof to send it back.

Your proof

  1. Pick the first line below.

Available lines

Prove it yourself

Half a step at most

+35 XP

Claim

For a block x\mathbf x with ∥x∥∞>0\|\mathbf x\|_\infty>0, let Δ=∥x∥∞/127\Delta=\|\mathbf x\|_\infty/127, qi=⌊xi/Δ⌉q_i=\lfloor x_i/\Delta\rceil and x^i=Δqi\hat x_i=\Delta q_i. Prove that every qiq_i lies in [−127,127][-127,127] and that ∣xi−x^i∣≤Δ/2|x_i-\hat x_i|\le\Delta/2.

Preview

Your typeset proof appears here.

Coding problems

Problem 25·Warm-up

How many users fit?

+20 XP

A dense model with 70×10970\times10^9 parameters is stored at 4.5 bits per weight. It has 80 layers, 8 KV heads and head dimension 128, and keeps its KV cache at 16-bit (2 bytes per number). Each sequence holds 8,192 tokens. The memory needed is 1.1×1.1\times (weights + KV cache), the extra 10% covering activations and buffers. What is the largest number of concurrent sequences that fits in 80×10980\times10^9 bytes?

An exact integer (or a fraction like 7/12)

Problem 26·Standard

Outliers and block scales

+35 XP

Generate 256 weights with the MINSTD generator: x0=42x_0=42, xn+1=48271 xn mod (231−1)x_{n+1}=48271\,x_n \bmod (2^{31}-1), un=xn/(231−1)u_n=x_n/(2^{31}-1) and wn=2un−1w_n=2u_n-1 for n=1,…,256n=1,\dots,256. Then set the first weight to w1=20w_1=20, an outlier. Quantise to int8 with absmax in two ways: (a) one step for the whole vector; (b) one step for each block of 32 consecutive weights. Each step is Δ=max⁡∣w∣/127\Delta=\max|w|/127 over the weights sharing it, with q=round⁡(w/Δ)q=\operatorname{round}(w/\Delta) (no ties occur) and w^=Δq\hat w=\Delta q. Print the mean squared error of (a) divided by the mean squared error of (b), to 2 decimal places.

A number, rounded to 2 decimal places

Problem 27·Challenge

A frugal cascade

+50 XP

1,000 queries arrive. Using MINSTD with x0=7x_0=7 (xn+1=48271 xn mod (231−1)x_{n+1}=48271\,x_n \bmod (2^{31}-1) and u=x/(231−1)u=x/(2^{31}-1) after each step), draw three numbers per query, in order: aia_i, bib_i, did_i. A small local model answers query ii with confidence si=ais_i=a_i and is correct when bi<sib_i<s_i. A large API model is correct when di<0.92d_i<0.92. A cascade with threshold τ\tau always runs the local model (cost 1 unit) and, when si<τs_i<\tau, also calls the API (10 more units) and uses its answer instead. Among τ=k/100\tau=k/100 for k=0,1,…,100k=0,1,\dots,100, consider the cascades with at least 900 correct answers. Print the lowest total cost.

An exact integer (or a fraction like 7/12)

Key takeaways

  • Closed APIs buy capability and zero operations; open weights buy control, privacy, reproducibility and freedom from forced migrations. Open weights are not the same as open source AI.
  • Memory is weights (NbNb) plus KV cache (2LhkvdhTBbkv2Lh_{kv}d_hTBb_{kv}) plus overhead; quantisation shrinks the first term with an error of at most half a step.
  • Single-stream decoding is bounded by bandwidth over bytes read per token; batching shares the reads, and a mixture of experts needs memory for its total parameters but reads only the active ones.
  • Owning breaks even at V∗=F/(p−c)V^*=F/(p-c) only when the local marginal cost is below the API price and one machine can actually produce that volume.
  • Route or cascade between models, and decide with your own evals rather than leaderboards.

Checkpoint

Prove it to the labyrinth

Answer every question to clear this chamber. First-try answers earn the most XP.

0/8
Question 1 of 8 +20 XP

A mixture-of-experts model has 47B parameters in total and 13B active per token. Compared with a dense 13B model at the same precision, it needs…

Question 2 of 8 +20 XP

How many GB (10910^9 bytes) of weight memory does a 70B-parameter model need at 4.5 bits per weight?

Question 3 of 8 +20 XP

You must be able to reproduce last year's outputs exactly. Which setup makes that most achievable?

Question 4 of 8 +20 XP

Why does batching raise decoding throughput on a GPU?

Question 5 of 8 +20 XP

The local marginal cost per token is above the API price per token. What is the break-even volume?

Question 6 of 8 +20 XP

A cascade answers every query with a cheap model first and escalates to a stronger one when a scorer is unsure. When does it save money without losing much accuracy?

Question 7 of 8 +20 XP

Absmax int8 quantisation is applied to a block whose largest magnitude is 2.54. After dequantising, what is the largest possible rounding error for any element?

Question 8 of 8 +20 XP

Which file format packages quantised weights for llama.cpp and tools such as Ollama and LM Studio?

End of the chamber

Clear this chamber

+60 XPOpen WeightsQuantisationMemory-Bandwidth BoundBreak-Even Volume