Failure AtlasLegal · Professional Services · 2023

Mata v. Avianca: The Cases That Did Not Exist

Two lawyers filed a brief built on cases ChatGPT invented, then asked ChatGPT whether the cases were real and accepted its yes. A federal judge made the lesson permanent.

Failure typeProbabilistic (hallucination, verified against itself)
Where it failedPeople · Process
FiledJul 2026
The missing control

Citations went to court checked against nothing except the tool that invented them. The control absent: verification in an independent source of record before an output becomes a filing, and candor with the court the moment doubt appeared.

NIST RMFKnow the failure modes of the tool. A text generator is not a database.
ISO 42001Human oversight means checking the output, not asking the system to vouch for itself.
EU AI ActProfessional users of AI remain responsible for what they submit.
STAMPThe verification loop closed through the component being verified.
PlainDo not ask the thing that made it up whether it made it up.

Roberto Mata sued the airline Avianca in federal court in New York over a knee injury from a metal serving cart. The case would be a footnote, except for how his lawyers wrote their brief. Facing a motion to dismiss, attorney Steven Schwartz used ChatGPT for legal research, and the brief his colleague Peter LoDuca filed cited case after case supporting their position. Varghese v. China Southern Airlines. Shaboon v. Egyptair. Martinez v. Delta.

None of them existed. ChatGPT had generated them, complete with convincing citations, quotations, and internal reasoning. Avianca’s lawyers told the court they could not find the cases. Neither could Judge P. Kevin Castel, who ordered the plaintiffs’ lawyers to produce them. At that point the process could still have been saved by a single honest sentence. Instead, Schwartz went back to ChatGPT, asked whether the cases were real, and was assured they were and could be found in reputable databases. The lawyers then filed excerpts of the fake opinions themselves, fabrications layered on fabrications.

In June 2023, Judge Castel sanctioned Schwartz, LoDuca, and their firm under Rule 11, fined them 5,000 dollars, and wrote an opinion that has been quoted in courtrooms and boardrooms ever since. The underlying case was dismissed on other grounds. The sanction is what the case is remembered for.

Where it failed

The model did what probabilistic text generators do. It produced plausible text, and legal citations are a genre it can imitate fluently, right down to the pincites. Calling this a technology failure misses the mechanism. The tool was never wrong about what it is. The lawyers were wrong about what it is, treating a text generator as a search engine with a legal database behind it.

The controlling failure was in process and people. Legal practice already had the control this needed: you verify citations in a source of record before you file. That step was skipped. Worse, when doubt arrived, the verification loop was closed through the very system being verified. Asking ChatGPT to confirm its own citations is not a check at all. It only looks like one. And when the court’s questions made the situation plain, the response was to double down rather than disclose, which converted an embarrassing error into sanctionable conduct.

They asked the tool that invented the cases whether the cases were real, and it said yes.

How it could have been caught

One search in Westlaw or LexisNexis, the databases every lawyer already uses, would have caught all six fake cases in minutes. That is the whole control, and it existed the entire time. Behind it sit two habits worth writing into any AI policy: verification must be independent of the system that produced the output, and the moment doubt appears, the move is disclosure, because in any regulated setting the cover-up costs more than the error.

What it means for AI

Mata is the cleanest demonstration that a language model’s confidence carries no information about its accuracy. The system will assert fabrications and, if asked, vouch for them, in the same tone it uses when it is right. Every organization deploying these tools inherits that property, in legal research, in compliance summaries, in anything an agent asserts about the world.

The lesson transfers as one rule: no output that matters goes forward on the model’s own authority. Verification runs through an independent source, and the person who signs remains responsible for what they sign. The lawyers in Mata were not sanctioned for using AI. The judge said as much. They were sanctioned for abandoning the verification their profession already required, and for letting the tool vouch for itself when it mattered most.


Back to the Atlas