AI Accuracy

What the AI hallucination cases mean for any firm that signs AI-drafted work

Courts and tribunals on both sides of the Atlantic have now ruled on AI-generated errors. The pattern is consistent: the tool is never the defendant. Whoever put their name to the output answers for it.

For a while, an AI system inventing something was treated as a curiosity. That period is over. Courts and tribunals have now ruled on what happens when AI-generated content is wrong and someone relies on it, and the answer is the same every time.

What happened in Mata v Avianca?

In 2023, lawyers in a personal injury claim against the airline Avianca filed a brief in a New York federal court citing cases that did not exist. They had been produced by ChatGPT. When the court and the other side could not find them, the lawyers stood by them for weeks. In June 2023 Judge P. Kevin Castel sanctioned the lawyers and their firm, fining them $5,000, in an opinion that is now cited in almost every discussion of the subject.

The fine was small. The point of the opinion was not. The judge was clear that using a new tool was not the problem; failing to check it, and then standing behind what it produced, was.

What did the High Court say in June 2025?

On 6 June 2025 the Divisional Court of the King's Bench Division, presided over by the President of the King's Bench Division, gave judgment in two referred cases, Ayinde v London Borough of Haringey and Al-Haroun v Qatar National Bank, in both of which material put before the court contained authorities that did not exist (judgment, Courts and Tribunals Judiciary).

The court used the judgment to warn the profession that freely available AI tools are not capable of conducting reliable legal research, that lawyers who use them remain responsible for checking the results against authoritative sources, and that those who put false material before a court can expect serious consequences, up to and including referral to their regulator.

Why does a chatbot case matter to firms that are not law firms?

Because the same principle applies outside the courtroom. In Moffatt v. Air Canada (2024 BCCRT 149), a customer relied on the airline's website chatbot, which told him he could claim a bereavement fare refund after travelling. That was not the airline's policy. Air Canada argued, in effect, that the chatbot was responsible for its own statements. The British Columbia Civil Resolution Tribunal rejected that outright: the chatbot was part of Air Canada's website, and Air Canada was responsible for all the information on it (Moffatt v. Air Canada, CanLII).

No regulated firm would put its name to a client letter, a suitability report or a clinical note and then argue the drafting tool should carry the blame. These cases simply confirm that the argument would fail.

The rule that runs through all of them

Take the cases together and the rule is simple. The AI tool is never the party that answers. The person or organisation that put the output into the world does, and the question they will be asked is what they did to check it.

That shifts the problem. It is no longer enough that someone looked at the output. The organisation needs to be able to show, after the event, what was checked, against what, and what was found.

What fabrication cases miss

Invented citations are the failure everyone talks about because they are easy to spot once someone looks. The harder failure is the output that is accurate in everything it says and leaves out the one fact that mattered: the exclusion missing from a policy summary, the allergy missing from a clinical note, the condition missing from a credit memo. No fabrication check will find an omission, because nothing in the output is false.

Research on clinical note summarisation has found omissions to be more common than outright fabrications. In one study, doctors reviewing AI-drafted notes found 1,712 omitted sentences against 191 hallucinated ones (Asgari et al., npj Digital Medicine, 2025). An evidence process that looks only for false statements is looking for the rarer problem.

What good evidence looks like

  • Each claim in the output checked against the source it came from, with the passage it rests on recorded
  • Claims the source does not mention reported as unsupported, which is not the same as wrong
  • Material facts from the source checked for in the output, with omissions reported separately
  • The exact versions of the source and output recorded, so the check can be repeated
  • Disagreements left for a person, not settled by whichever checker spoke last

Key takeaways

  • Courts and tribunals have ruled consistently: whoever signs AI-drafted work answers for it
  • The question after an error is what you did to check, and whether you can show it
  • Omissions are a larger risk than fabrications and need their own check
  • Evidence has to be recorded at the time, not reconstructed when someone asks

The Accuracy Evidence File produces exactly this record for a set of your AI outputs: every claim checked against its source, every omission listed, in three weeks at a fixed price.

AI AccuracyAccuracy Evidence FileAI Governance

Want to apply this to a specific process?

Bring the workflow you had in mind. We will talk through whether these ideas apply to it, and what it would take to find out.

Around four minutes. Indicative guidance based on your answers.

Region & currency

Changes spelling, terminology, the data-protection regime named in our notices, and the currency used in indicative figures. It does not change where we are or quote you a price in your currency. ETT is headquartered in Dallas, with offices in London and Vancouver.