Meta and the Metric That Ate the Mission
Facebook rebuilt its feed around 'meaningful social interactions', then weighted an angry reaction five times a like. The system optimized exactly what it was told to. That was the failure.
Ranking optimized a proxy built from reactions, comments, and reshares, and no control checked whether the proxy still matched the goal it stood for. The control absent: a validity review of the objective itself, with the authority to change the weights when the measurements said the metric had diverged.
In 2018, Facebook rebuilt its News Feed ranking around a metric it called meaningful social interactions, MSI. The stated goal was worthy: less passive scrolling, more interaction between people. The mechanism was a scoring formula. Content earned distribution based on the engagement it generated, and different signals carried different weights. When emoji reactions entered the formula, they were worth five times a like. All of them, including anger.
This entry stays on the mechanism, because the mechanism is the lesson. A reaction is a click. The formula could not ask why someone reacted, only that they did. And the content most likely to make people react, comment, and reshare is not the content the mission statement had in mind. By 2019, Facebook’s own data scientists had confirmed internally that posts drawing heavy angry reactions were disproportionately likely to contain misinformation, toxicity, and low-quality news. The machine was working precisely. It had been told that engagement meant meaning, and it delivered engagement.
The weights tell the story of a slow correction. Anger was cut from five times a like to four in 2018. Internal proposals to go further stalled. In September 2020, the weight on the angry reaction finally went to zero, and the internal measurements improved: less misinformation, less of the content users themselves said they did not want. The world learned the details in 2021, when documents disclosed by Frances Haugen, the Facebook Papers, put the internal research on the record.
Where it failed
This is the proxy failure at the largest scale there is. The ranking system was probabilistic, but it was not malfunctioning, drifting, or biased in the engineering sense. It optimized its objective with enormous competence. The failure was that the objective, a weighted sum of clicks, stood in for a goal, meaningful human interaction, that it measured less and less faithfully as the optimization pressure grew. This is Goodhart’s law with a distribution system attached: when the measure became the target, it stopped being a good measure, and two billion feeds were tuned to the gap.
The break was in process above technology. The weights were a design decision, an editable number in a formula. What was missing was the governance around that number: a standing review of whether the proxy still tracked the mission, run by people with the authority to change it, triggered by the company’s own measurements rather than by a whistleblower and a congressional document dump.
Nothing about the system malfunctioned. It optimized the number it was handed, and that number was wrong.
How it could have been caught
It was caught. That is what makes the case instructive. The internal research existed years before the correction, and the fix, when it came, was a single weight set to zero. The control that failed was the path from measurement to authority: findings about the metric’s divergence had no guaranteed route to the people who owned the metric, and the metric’s owners were measured on growth in the metric. Treat the objective function as a governed artifact. Give it an owner, a review cadence, divergence measurements with teeth, and a change process that does not require a crisis.
What it means for AI
Every recommender, every ranking model, every agent given a number to maximize is running this experiment. The objective is always a proxy. Engagement for value, cost for need, clicks for relevance, task completion for helpfulness. Optimization pressure finds the gap between the proxy and the goal as reliably as water finds a crack, and more capable systems find it faster.
Govern the target, not just the model. Audits of accuracy, bias, and drift all miss the case where the system is perfectly accurate at the wrong thing. So the questions are what the number actually stands for and whether it still stands for that. And the power to change it has to sit with someone whose pay does not depend on it going up. The most consequential line of code in any AI system is the one that says what to maximize.
← Back to the Atlas