Case 003 · on the docket
Case 003Source under trialWIRED

Is AI wrong half the time? Sometimes. That number needs a trial of its own.

The rulingMostly upheld

The warning is right; a single “AI error rate” is not.

A professional WIRED fact-checker argues that today's AI systems remain unreliable for serious verification work. She cites studies with high error rates and describes a hands-on test in which ChatGPT, Claude, Gemini and Grok handed her fact-checking plans rather than actually completing the verification.

The professional judgement is sound and we largely uphold it. What does not survive is the thing readers take away: a single portable number for how often AI is wrong.

Accuracy & trust3 cited sources8 min read

AI-generated analysis. Written by AI in conversation with a user. Not an official statement, position or publication of OpenAI or any other AI vendor.

Your ruling

How would you rule?

Anonymous · no account required.

Step 1
The accusation

AI can't be trusted to check facts

Stated at its strongest: verification is not text generation. A fact-checker interrogates primary and offline sources, calls people, reads documents that were never digitised, weighs context and legal exposure, and decides what "true" even means for an ambiguous claim. Systems that answer from statistical patterns over text, plus whatever a search index returns, are not doing that job — and in tests they produce error rates that would end a human checker's career.

Step 2
The evidence offered

What the article actually demonstrated

Two distinct things. First, a hands-on trial: asked to verify claims, the assistants mostly produced methodology — here's how you would check this — instead of finishing the task. That is a real and reproducible observation about product behaviour, and it matches the incentives of assistants tuned to be helpful and cautious rather than conclusive.

Second, citations to studies reporting high error rates. Each of those studies is presumably accurate about what it measured. The article's move — and it is the move nearly every downstream summary repeats — is to let heterogeneous results imply one background fact about AI reliability.

Step 3
Right of reply

What the vendors would say, and where it fails

The vendor reply is that these evaluations often disable the tools that matter: no web search, no retrieval, no citations, no iteration. Run with search enabled and a prompt demanding sources, they'd argue, the numbers change substantially.

That reply is partly right and mostly beside the point. It does not touch the article's central claim, which is about professional verification, not about lookup. Retrieval helps with "what did the report say"; it does nothing for "is the person on the phone telling the truth." And per rule two, a vendor's own reliability measurement is testimony, not a ruling.

Step 4
Independent check

Why a single error rate isn't a real quantity

Two independent lines of evidence show how much the number moves with the setup.

Evaluation design changes behaviour. Work published in Nature makes the point that scoring models purely on answer accuracy rewards guessing: a model that says "I don't know" scores zero, a model that guesses sometimes scores. Hallucination rates are therefore partly a property of the exam, not only of the student.

Format and retrieval dominate. A study of commercial chatbots as news intermediaries put roughly 2,100 same-day news questions to six chatbots. Strong systems could exceed 90% on some multiple-choice setups yet fall materially on free-response versions of comparable material — and most errors traced back to retrieval failures rather than reasoning failures.

Exhibit AEvidence exhibit

One headline number, six hidden variables

The claim as circulated
~50%

“AI gets it wrong about half the time” — a single rate, presented as a property of the technology.

What actually moves the number
  • Question difficulty and recency
  • Multiple-choice vs free response
  • Retrieval on or off
  • Whether abstention is rewarded
  • Which model and which product surface
  • Who grades, and against what

A 2% figure, a 45% figure and a 90% figure can each be honestly measured. None of them describes how often AI is wrong for you, on your questions, with your tools switched on.

Nature (s41586-026-10549-w) on abstention scoring; arXiv:2605.22785 on ~2,100 same-day news questions across six chatbots.

That is the whole argument in one place. A 2% figure, a 45% figure, a 60% figure and a 90% figure can each be honestly measured, and none of them describes "how often AI is wrong" for you, on your questions, in your product, with your tools switched on.

Step 5
The ruling

Verdict: mostly upheld

The warning is right; a single "AI error rate" is not. AI is not a substitute for professional fact-checking — especially where offline sources, human interviews, legal or ethical judgement, or genuinely ambiguous claims are involved. We concede that fully and without hedging.

We challenge only the compression: collapsing incompatible tests into an implied universal reliability figure. Error rates vary dramatically with task difficulty, evaluation format, web and retrieval access, prompt design, and whether abstention is rewarded or punished.

Step 6
For a real user

Use it as an accelerator, not an authority

Treat an assistant as a research accelerator that surfaces sources and candidate explanations, then verify anything high-stakes against the primary material yourself. Some concrete habits:

  • Ask for sources and open them. An uninspected citation is not evidence.
  • Prefer free-response answers you can check over confident summaries you cannot.
  • Say explicitly that "I don't know" is an acceptable answer — you are removing the incentive to guess.
  • For anything legal, medical, financial or reputational, the AI's job is to find the document, not to be the document.

Retrieval is doing more of the work than most people assume — which is exactly why Case 004, on poisoning what AI retrieves, is the natural next read.

References

Sources

Still your ruling

Now you've read the evidence — anonymous, no account required.

Share the ruling

No ads, no signup — sharing is the only distribution we have.

Is AI wrong half the time? Sometimes. That number needs a trial of its own. — verdict: Mostly upheld. The warning is right; a single “AI error rate” is not. Evidence-led, model-neutral. https://chadgpt-response.lovable.app/cases/ai-wrong-half-the-time

Next case · Case 004Can propaganda poison AI answers? Yes — and the weak point may be retrieval, not intelligence.Demos / The Times · Upheld
Related cases
Help shape the docket

What brings you to AI Rebuttal?

One click, nothing else. Anonymous · no account required.