Failure AtlasCriminal Justice · 2016

COMPAS: When Fairness Has More Than One Definition

A risk score used in courtrooms was calibrated correctly and still wrong more often about Black defendants who would not reoffend. Both sides of the argument were right, which is the whole lesson.

Failure typeProbabilistic (a validity failure: which fairness, exactly)
Where it failedProcess · Technology
FiledJun 2026
The missing control

A risk score drove consequential decisions without anyone deciding which definition of fairness it was required to meet. The control absent: choose the fairness criterion that fits the stakes, in the open, and accept that you cannot satisfy them all at once.

NIST RMFDefine and measure fairness explicitly. Pick the metric that matches the harm.
ISO 42001Document the objective, the trade-offs, and the human's role in the decision.
EU AI ActRisk assessment in justice is high-risk, with transparency and oversight duties.
STAMPThe hazard was an unstated objective. No one specified the fairness the system owed.
PlainYou cannot be fair in every sense at once. Decide which sense you mean.

In 2016, ProPublica investigated COMPAS, a risk-assessment tool sold by a company then called Northpointe and used in US courts to estimate how likely a defendant was to reoffend. The scores informed real decisions, about bail, sentencing, and supervision.

ProPublica found that among defendants who did not go on to reoffend, Black defendants were labeled high-risk almost twice as often as white ones. The false-positive rate was about 45 percent for Black defendants and about 23 percent for white ones. That looks like a clear case of a biased algorithm.

Northpointe answered, and their answer was also correct. The tool was calibrated: a given score meant the same probability of reoffending regardless of race. A seven meant the same thing for everyone. By that definition, the standard one for a risk score, COMPAS was fair.

Both were right. That is not a paradox to be resolved. It is the point.

Where it failed

COMPAS breaks the assumption that fairness is a single thing you can test for and pass. When two groups reoffend at different base rates, it is mathematically impossible for a risk score to be calibrated and to have equal false-positive rates at the same time. You get one or the other, not both. So the question is never simply whether it is fair. The question is fair in which sense, and that is a choice someone has to make.

In COMPAS, no one made that choice in the open. A probabilistic score was built to one reasonable definition of fairness and dropped into courtrooms where a different definition, the one about who gets wrongly flagged, mattered enormously to the people being scored. The failure was not a bug in the model. It was a validity question, what fairness does this system owe, that was never asked or answered.

The score was calibrated, and it was wrong more often about Black defendants who never reoffended. Both were true at once.

How it could have been caught

The control is to make the choice explicit, before deployment, with the people who bear the consequences in the room. Decide which fairness criterion the system is required to meet, given what it is used for, and accept in writing what you are trading away to get it. For a score that can keep someone in jail, the rate of wrongly flagging people who will not reoffend is not a footnote. State the objective, measure against it across groups, and keep a human making the decision who understands that the number is an estimate built on a contested definition, not a fact.

What it means for AI

Every model that scores or ranks people inherits this problem, and that covers most of the consequential ones. Hiring, lending, fraud, content moderation: every one embeds a definition of fairness, whether or not anyone stated it. The math that bit COMPAS bites all of them. You cannot optimize every fairness metric at once, and the ones you ignore do not disappear, they land on whichever group the base rates disfavor.

So fairness has to be treated as a design decision rather than a box checked at the end. Decide which definition the system owes, given what it is used for, and say so where the people affected can see it. Then measure the outcome across groups, including the metric you chose not to optimize, so you at least know who is absorbing the cost of the choice. A model that is fair by one honest definition can be unjust by another, and the people being scored are the ones who feel the difference.


Back to the Atlas