Harms Start With Measurement
Representational, allocative, and quality-of-service harm each need different evidence, and one aggregate score hides the subgroup it fails.
Representational harm, allocative harm, and quality-of-service harm each need their own evidence. One aggregate score supplies none of the three, and it hides the subgroup it fails.
The capabilities note measured what a model can do. Harm analysis names who pays when that model fails, and what the failure costs them. Capability without harm analysis is power without accounting.
The overview note asked which layer did the work. Harm answers to the same question. It enters through the corpus, model behavior, the interface, and the deployment setting. Name the layer and a bad output turns into a decision someone owns.
Models Inherit The Patterns In Their Documents
A language model inherits the patterns of the societies and documents that produced its data. Those patterns come back four ways: worse performance for some groups, biased associations, stereotyped completions, and hierarchy encoded in language.
The data note supplies the mechanism. Filtering decides which text disappears from the corpus. Mixture weights decide which text becomes normal. A dialect that gets filtered out or under-weighted arrives later as an error rate nobody planned.
Performance disparity is the measurable form. Error rates differ across groups, domains, dialects, languages, and contexts. Social categories are contextual and datasets are imperfect, so one group label covers different people in different settings.
Three Harms, Three Kinds Of Evidence
Split the harm before you measure it. The three kinds behave differently, and each one answers to its own evidence.
- Representational harm changes how people are portrayed. Read completions and review the failures by hand.
- Allocative harm changes who gets resources or opportunities. Compare outcomes by group against base rates.
- Quality-of-service harm changes who receives worse behavior from the same system. Slice the error rate before you report it.
The three do not trade off against each other. A model can portray a group badly while its error rates look even across groups. It can allocate resources evenly and still write demeaning completions. The type you measure decides the failure you can see.
Every threshold picks an error. Tighten it and the system produces false positives. Loosen it and the system produces false negatives. The number decides which group eats which mistake, so name the group before you set the number.
The Aggregate Hides The Subgroup
The same aggregate metric hides different harm profiles, depending on how you slice the population. Aggregate numbers are blunt by construction. They average over the people the system treats worst.
Two versions of a system can report the same overall error rate and earn different verdicts. Slice by dialect and one holds flat, while the other fails on the dialect its corpus carried least. The average passes both. The slice passes one.
Hand review of failures catches what no slice defines. A completion can read as demeaning while the error rate stays flat. The dialect slice returns in the moderation note ahead.
Measurement Decides Which Harms Become Visible
Measurement choices define the groups, the prompts, the labels, the raters, the metrics, and the thresholds. Those choices decide which groups, mistakes, contexts, and trade-offs a benchmark can see. Everything outside that frame stays easy to ignore.
Prompt choice decides which behavior a benchmark ever elicits. A behavior that no prompt triggers reads as absent. The same benchmark can reward surface-level mitigation and leave the deeper behavior intact.
The person who writes the rubric decides what counts as harm. The person outside that rubric pays when the measurement misses.
A harm name gives you a category. The number behind it needs a definition, a dataset, a threshold, and an error analysis. Measurement is the first hard step.
The shortcut names harm categories and never tests whether the measurement can see them. That buys moral clarity on the page and operational blindness in the system.
You can act on a harm only after a measurement choice makes it visible.
Start The Harm Analysis Before Launch
A fairness dashboard after launch reports what already shipped. Run the harm analysis on the use case first, and write down four things.
- Name who the data represents, and who it leaves thin.
- Name who a wrong output lands on.
- Write down who can appeal, and the path that appeal takes.
- Price the false positive and the false negative, group by group.
The appeal path is where a person the measurement missed corrects the record. A system without one makes every wrong output final.
Draw the system boundary wide. The model, the prompt, the product interface, and the deployment setting each create harm on their own. The interface is part of the claim, so an analysis that stops at the weights stops early.
Early has a test. The analysis must land while product decisions still change. After launch it becomes a report.
The Builder Test
Take one aggregate metric from a system you run this week. Slice it by group or by dialect, and name the subgroup the average was carrying. Then name the error that costs that subgroup more, and the signal that tells you it happened.
If the slice does not exist yet, that absence is the finding. Build the slice before the next release.
What Carries
A harm definition that cannot name who is affected is too abstract to act on. Carry the affected user into every measurement decision, and the definition stays usable.
Bias in what a model reflects is one problem. What it produces on demand is the next one.