The Pipeline Is the Model
Eighteen notes from data to serving, each stage pricing the next.
One overview plus 17 focused notes on the full path: data, tokenization, transformers, training, systems, evaluation, inference, and alignment. The goal is a working mental model.
Eighteen notes from data to serving, each stage pricing the next.
BPE, byte-level fallbacks, and why the tokenizer is the first line on the bill.
The two numbers that price any run on paper, before anyone grants a cluster.
The decoder recipe every open model converges on, and the measurements behind each choice.
What mixture-of-experts buys per FLOP, and what routing charges back.
Bytes moved set the kernel ceiling, and FlashAttention is the proof that traffic beats math.
When to fuse, when to write a kernel, and the benchmark traps that lie first.
Splits answer the bottleneck, and cluster topology decides the cost.
Measured collectives against datasheet peaks, and the scheduling that closes the gap.
Fitting the loss line small so the expensive choices stop being blind.
Serving economics: the KV cache, the latency budget, and the scheduler between them.
What happened to Chinchilla's twenty tokens when production teams measured again.
Contamination and length bias, the two choices that decide leaderboards unseen.
The same tokens, points apart: how extraction, filtering, and mixture set the ceiling.
The one cheap pattern behind language ID, quality, toxicity, and deduplication.
How feedback shortcuts become model behavior, one rushed rater at a time.
Proxy rewards, their turn-down curve, and the move to rewards a program checks.
The baseline that makes sparse rewards trainable, and what it stabilizes.
Building a language model, data to deployment If you only have an hour, read kernels, scaling laws, inference, and alignment. That path shows the full loop: make training efficient, predict scale, serve the model, then shape behavior.