Read The Stack Sideways
Six parts to any language-model claim, and where the benchmark stops.
How language models are built, measured and misused
Six parts to any language-model claim, and where the benchmark stops.
Probabilities over strings, and the decoding dials that turn them into a writer.
Why a capability claim is incomplete without its prompt, context, and failures.
Three kinds of harm, each needing evidence an aggregate score hides.
Moderation thresholds as chosen errors, with both failure directions priced.
WebText's Reddit proxy, and every corpus choice that lands before training.
Extraction attacks, memorized secrets, and why deletion stops being clean.
Four legal questions a dataset must clear, answered before the takedown.
The tokenizer as first suspect when code, math, or names go strange.
Objectives as curricula, and the lower loss that hides worse behavior.
Naming the wall, then the parallel strategy, in that order.
Probe runs that make frontier compute legible, within the setup that produced them.
MoE and retrieval buy capacity, and selection becomes the failure point.
The least invasive adaptation that clears the bar, from prompts to tuning.
The serving bill that reopens on every call and outgrows the training run.