The Bias Audit Passed. The Tool Still Discriminated
A clean bias audit can hide real discrimination. That’s the uncomfortable thing in a new Stanford Digital Economy Lab study that got inside an actual hiring algorithm. The researchers analyzed 4 million applications from about 3 million applicants, all screened by the vendor pymetrics, because pymetrics actually handed over the data.
Pymetrics had published aggregated audits showing no adverse impact under the four-fifths rule. The researchers re-ran the analysis position by position instead of lumping everything together, and the disparities appeared, clearest for Black and Asian applicants. 10.62% of positions showed adverse impact against Black applicants. Nearly a third of Black applicants applied to at least one of them.
The aggregation is the problem. Average every role together and a disparity in one job gets cancelled out by the pattern in another, leaving a company-wide number that looks clean. US employment law doesn’t work that way. The four-fifths rule is applied one position at a time, comparing the people who applied to that role against each other. An aggregate audit skips the test the law actually requires. Measuring per position is the study’s own first recommendation and the authors point out that current guidance for NYC’s Local Law 144 pushes employers toward merging the data, which is exactly backwards.
So why does a system that never sees race still produce racial disparities? Applicants give no name, gender, race, or age. The mechanism the paper states as fact: each model is trained on the gameplay of at least 50 of the employer’s current employees in that role. The model learns to favor people who resemble the existing workforce, and if that workforce isn’t diverse, the model reproduces it. The authors also suspect the game features themselves track race indirectly, but they’re clear there’s no causal proof of that, only the outcome.
The bigger problem is the monoculture. When many employers screen through the same vendor, applying to more jobs doesn’t get you more chances, it gets you the same judgment again. The study found 10% of applicants to four positions were rejected by all four, well above what independent decisions would produce. Run the same method on an older study of 83,000 applications across 108 firms and the rejections there look independent. The algorithmic data doesn’t. A person can be shut out of a whole slice of the labor market by one model, and no single employer can see it, because each only sees its own funnel.
This is where ForHumanity’s AEDT governance standard helps. It breaks the assessment into per-use-case criteria rather than taking a vendor’s “no bias” summary as the end of the conversation. That per-role review is what would have caught the half of this study an aggregated audit smoothed over. It won’t catch monoculture, since that only shows up across employers and no single company can see it from inside its own hiring.
What gets me is that none of this is visible to the person being screened. They get a rejection, or they get nothing, and they never learn that one model made the call and that the same model may be making it everywhere else they apply.
A passing audit and a fair tool are not the same thing. This study is the proof.
Based on “Algorithmic Monocultures in Hiring,” Stanford Digital Economy Lab. Read the Q&A.
First published on Substack.
← Back to Field Notes