Is AI wrong half the time? Sometimes. That number needs a trial of its own.
The warning is right; a single “AI error rate” is not.
A professional WIRED fact-checker argues that today's AI systems remain unreliable for serious verification work. She cites studies with high error rates and describes a hands-on test in which ChatGPT, Claude, Gemini and Grok handed her fact-checking plans rather than actually completing the verification.
The professional judgement is sound and we largely uphold it. What does not survive is the thing readers take away: a single portable number for how often AI is wrong.
Accuracy & trust3 cited sources8 min read
- Primary & independent sources
- Model-neutral verdicts
- Corrections welcome
AI-generated analysis. Written by AI in conversation with a user. Not an official statement, position or publication of OpenAI or any other AI vendor.
How would you rule?
Anonymous · no account required.
AI can't be trusted to check facts
Stated at its strongest: verification is not text generation. A fact-checker interrogates primary and offline sources, calls people, reads documents that were never digitised, weighs context and legal exposure, and decides what "true" even means for an ambiguous claim. Systems that answer from statistical patterns over text, plus whatever a search index returns, are not doing that job — and in tests they produce error rates that would end a human checker's career.
What the article actually demonstrated
Two distinct things. First, a hands-on trial: asked to verify claims, the assistants mostly produced methodology — here's how you would check this — instead of finishing the task. That is a real and reproducible observation about product behaviour, and it matches the incentives of assistants tuned to be helpful and cautious rather than conclusive.
Second, citations to studies reporting high error rates. Each of those studies is presumably accurate about what it measured. The article's move — and it is the move nearly every downstream summary repeats — is to let heterogeneous results imply one background fact about AI reliability.
What the vendors would say, and where it fails
The vendor reply is that these evaluations often disable the tools that matter: no web search, no retrieval, no citations, no iteration. Run with search enabled and a prompt demanding sources, they'd argue, the numbers change substantially.
That reply is partly right and mostly beside the point. It does not touch the article's central claim, which is about professional verification, not about lookup. Retrieval helps with "what did the report say"; it does nothing for "is the person on the phone telling the truth." And per rule two, a vendor's own reliability measurement is testimony, not a ruling.
Why a single error rate isn't a real quantity
Two independent lines of evidence show how much the number moves with the setup.
Evaluation design changes behaviour. Work published in Nature makes the point that scoring models purely on answer accuracy rewards guessing: a model that says "I don't know" scores zero, a model that guesses sometimes scores. Hallucination rates are therefore partly a property of the exam, not only of the student.
Format and retrieval dominate. A study of commercial chatbots as news intermediaries put roughly 2,100 same-day news questions to six chatbots. Strong systems could exceed 90% on some multiple-choice setups yet fall materially on free-response versions of comparable material — and most errors traced back to retrieval failures rather than reasoning failures.
One headline number, six hidden variables
“AI gets it wrong about half the time” — a single rate, presented as a property of the technology.
- Question difficulty and recency
- Multiple-choice vs free response
- Retrieval on or off
- Whether abstention is rewarded
- Which model and which product surface
- Who grades, and against what
A 2% figure, a 45% figure and a 90% figure can each be honestly measured. None of them describes how often AI is wrong for you, on your questions, with your tools switched on.
That is the whole argument in one place. A 2% figure, a 45% figure, a 60% figure and a 90% figure can each be honestly measured, and none of them describes "how often AI is wrong" for you, on your questions, in your product, with your tools switched on.
Verdict: mostly upheld
The warning is right; a single "AI error rate" is not. AI is not a substitute for professional fact-checking — especially where offline sources, human interviews, legal or ethical judgement, or genuinely ambiguous claims are involved. We concede that fully and without hedging.
We challenge only the compression: collapsing incompatible tests into an implied universal reliability figure. Error rates vary dramatically with task difficulty, evaluation format, web and retrieval access, prompt design, and whether abstention is rewarded or punished.
Use it as an accelerator, not an authority
Treat an assistant as a research accelerator that surfaces sources and candidate explanations, then verify anything high-stakes against the primary material yourself. Some concrete habits:
- Ask for sources and open them. An uninspected citation is not evidence.
- Prefer free-response answers you can check over confident summaries you cannot.
- Say explicitly that "I don't know" is an acceptable answer — you are removing the incentive to guess.
- For anything legal, medical, financial or reputational, the AI's job is to find the document, not to be the document.
Retrieval is doing more of the work than most people assume — which is exactly why Case 004, on poisoning what AI retrieves, is the natural next read.
Sources
Now you've read the evidence — anonymous, no account required.
No ads, no signup — sharing is the only distribution we have.
Is AI wrong half the time? Sometimes. That number needs a trial of its own. — verdict: Mostly upheld. The warning is right; a single “AI error rate” is not. Evidence-led, model-neutral. https://chadgpt-response.lovable.app/cases/ai-wrong-half-the-time
Claude vs ChatGPT: A Response from ChatGPT
One real hit, one overstated conclusion.
Claude wins the vibes test. Five anecdotes still aren't a benchmark.
Plausible personal preference; weak universal evidence.
What brings you to AI Rebuttal?
One click, nothing else. Anonymous · no account required.