Claim analyzed

Tech

“Artificial intelligence systems produce incorrect answers in 70% of evaluated cases.”

The conclusion

False
2/10

Available evidence does not establish a general 70% error rate for artificial intelligence systems. Reported rates vary dramatically by model, task, benchmark, and definition of error. A result near 70% appears in a narrow medical evaluation, but presenting it as broadly representative of evaluated AI answers is unsupported.

Caveats

  • The 70% figure appears to come from a narrow evaluation and cannot be generalized to all AI systems.
  • Error and hallucination rates depend heavily on the task, model, benchmark, and scoring method.
  • Some apparent corroboration comes from secondary headlines lacking enough methodological context to verify equivalence.

Sources

Sources used in the analysis

#1
journals.plos.org 2024-07-31 | Evaluation of ChatGPT as a diagnostic tool for medical learners and clinicians

Out of the 150 cases, ChatGPT provided correct answers for 74/150 (49%) of cases (Fig 1).

#2
medrxiv.org 2026-02-24 | Evaluating the AI Potential as a Safety Net for Diagnosis: A Novel Benchmark of Large Language Models in Correcting Diagnostic Errors

Gemini 2.5 Pro demonstrated the highest performance, correcƟng the physician's error in 55.0% of cases (n=110/200), followed by Claude Sonnet 3.5 (48.5%) and Sonnet 4 (47.0%). … In contrast, DeepSeek V3 corrected only 20.0% of cases.

On MMLU-Pro (10-choice; primary-model accuracy 0.16–0.29 on the full sample), every primary instruct model overstates its confidence verbally. … Averaged over variants passing the V P R ≥ 0.80 inclusion threshold, mean verbal confidence is 60–78 percentage points above accuracy and 25–60 points above mean token-probability confidence on the same rows.

#4
arxiv.org 2026-05-07 | Reliability without Validity: A Systematic, Large-Scale Evaluationof LLM-as-a-Judge Models Across Agreement, Consistency, and Bias

LLM-as-a-Judge has become the dominant evaluation paradigm for language models, but judge validation in practice relies on exact-match agreement, a metric that does not correct for chance and systematically overstates discriminative ability. … First, kappa deflation: raw agreement overstates chance-corrected discrimination by 33–41pp in all 21 evaluated models.

#5
cidrap.umn.edu 2026-04-15 | AI chatbots provide poor answers to medical questions half the time ...

A study published in BMJ Open suggests that half of answers provided by five publicly available artificial intelligence (AI)–driven chatbots in response to medically related questions are inaccurate and incomplete. … About half (49.6%) of responses were problematic, with 30% considered somewhat problematic and 19.6% deemed highly problematic.

#6
cjr.org AI Search Has a Citation Problem - Columbia Journalism Review

Collectively, they provided incorrect answers to more than 60 percent of queries.

#7
digitaltrends.com 2025-12-15 | Google finds AI chatbots are only 69% accurate… at best - Digital Trends

Using its newly introduced FACTS Benchmark Suite, the company found that even the best AI models struggle to break past a 70% factual accuracy rate. … The top performer, Gemini 3 Pro, reached 69% overall accuracy, while other leading systems from OpenAI, Anthropic, and xAI scored even lower.

#8
preview-www.nature.com 2024-07-04 | Evaluation and mitigation of the limitations of large language models in clinical decision-making | Nature Medicine

We show that current state-of-the-art LLMs do not accurately diagnose patients across all pathologies (performing significantly worse than physicians), follow neither diagnostic nor treatment guidelines, and cannot interpret laboratory results, thus posing a serious risk to the health of patients.

#9
medrxiv.org 2026-04-23 | Dissecting clinical reasoning failures in frontier artificial intelligence using 10,000 synthetic cases

Management safety performance was markedly lower. Safe corticosteroid recommendations (e.g. not recommending immediately with evidence of active infection) occurred for as few as 7.2% (Gemini 3 Flash Preview; 95% CI 5.6-8.8) and up to only 23.5% (GPT-5 mini; 95% CI 20.8-26.1) cases.

#10
hai.stanford.edu Responsible AI | The 2026 AI Index Report - Stanford HAI

In a new accuracy benchmark, hallucination rates across 26 top models range from 22% to 94%.

#11
doi.org 2026-06-19 | GPT-4.1 and Llama 3.3 70 fail to detect clinically relevant errors in radiology reports in zero-shot evaluation

Physiologically impossible errors (E2) showed the lowest performance: 46.2% (CT) and 33.7% (MRI) for GPT-4.1, compared with 32.7% and 25.0% for Llama 3.3, respectively. … Mislabeling errors (E1) were detected in 49.0% by GPT‑4.1 and 33.7% by Llama 3.3 for MRI.

#12

We evaluate a wide range of models on these platinum benchmarks and find that, indeed, frontier LLMs still exhibit failures on simple tasks such as elementary-level math word problems.

#13
doi.org 2025-11-10 | A comparative evaluatıon of AI-based ChatGPT and physicians in psychiatric case assessment

When similarity rates were evaluated, the highest similarity was observed in 'Prescribed medications' (98.3%), followed by 'Dosage information of prescribed medications' (97.5%), while the lowest similarity was found in 'In-depth inquiry of complaints, signs, and symptoms' (0.0%), followed by 'Potential side effects of prescribed medications' (13.3%).

#14

We study the error rate of LLMs on tasks like arithmetic that require a deterministic output, and repetitive processing of tokens drawn from a small set of alternatives. … Our empirical tests span 8 different tasks and 3 different models with 0.2 million total test prompts. Formula (1) provides a surprisingly good description of the accuracy in almost all cases.

#15
joshbersin.com 2025-10-26 | BBC Finds That 45% of AI Queries Produce Erroneous ...

Today the BBC and EBU (European Broadcasting Union) published a detailed study which shows that around 45% of AI news queries to ChatGPT, MS Copilot, Gemini, and Perplexity produce errors.

#16
github.com 2026-05-11 | Vectara Hallucination Leaderboard

|Model|Hallucination Rate|Factual Consistency Rate|Answer Rate|Average Summary Length (Words)| … |antgroup/finix_s1_32b|1.8 %|98.2 %|99.5 %|172.4| … |microsoft/Phi-4-mini-instruct|23.5 %|76.5 %|92.5 %|420.2|

#18
newsbytesapp.com 2025-12-17 | AI chatbots get 33% of answers wrong, Google says

The tech giant used its new FACTS Benchmark Suite to find that even the best-performing AI models fail to achieve a factual accuracy rate of over 70%. … In simple terms, today's chatbots give an incorrect answer about one-third of the time, even if they sound completely confident.

#19
presenc.ai 2026-05-07 | AI Hallucination Rate Benchmarks 2026 | Presenc AI

On Vectara's Hughes Hallucination Evaluation Model (HHEM) summarisation benchmark, frontier models in 2026 hallucinate on approximately 1.0-2.5 percent of summaries, down from 3-8 percent in 2023. … On RAG-faithfulness benchmarks (RAGTruth), frontier-model hallucination rates are 4-9 percent, materially higher than summarisation, because RAG must integrate external context with internal knowledge. … On closed-book factuality benchmarks (TruthfulQA, FACTS Grounding), frontier models score 80-90 percent accuracy in 2026, up from 50-65 percent in 2023; long-tail facts remain the dominant failure mode.

#20
ai2med.eu 2025-11-24 | 70% Wrong? New Study Warns: AI Chatbots Still Struggle in Medical ...

According to research published in the European Journal of Pathology (the official journal of the European Society of Pathology), artificial intelligence models provided erroneous answers in nearly 70% of diagnostic cases and fabricated or inaccurate references in over 30%. … Useful answers were provided in 62.2% of cases, but only 32.1% were entirely error-free.

#21

If a benchmark reports 90% accuracy, expect 70-80% in production when accounting for consistency and faults.

#22
aclanthology.org Curriculum Learning based Hierarchical Scoring and Analysis Framework for Question Answering Task Evaluation
#23
#24
arxiv.org ERRORQUAKE: Heavy-Tailed Error Severity Distributions in Open-Weight Large Language Models
#25
arxiv.org Reliability Scales Inversely: Bigger Language Models Compound Mistakes Faster
#26
truestandard.ai 2026-07-29 | AI Hallucination Rates in 2026: Every Model, Measured | TrueStandard

There is no single hallucination rate, and anyone quoting one number without naming the benchmark is telling you less than they appear to. … Reading 94% as "AI is wrong 94% of the time" The AA-Omniscience figures count guesses on questions the model could not answer, questions written to be hard. On everyday, well-documented topics, frontier models are right far more often than not.

#27
multigrid.ai 2026-08-07 | How Often Models Hallucinate, According to the Studies · Multigrid

Published hallucination rates range from a couple of per cent to well over half, and the studies are not contradicting each other. … “Models hallucinate 20 per cent of the time” is not a statement about anything. “On a short fact-seeking QA benchmark published in 2024, this named model snapshot gave an incorrect answer to N per cent of questions it attempted, declining on M per cent” is a statement somebody can check.

#28
hackshackers.com 2026-05-22 | AI chatbots right about daily news nearly all the time, but ...

Retrieval failures — not reasoning failures — drove more than 70% of errors.

#29
wsu.edu Study shows AI systems fail accuracy tests about 70 percent of the time

study-shows-ai-systems-fail-accuracy-tests-about-70-percent-of-the-time

Debate

Two AI advocates debated this claim using the research gathered.

Argument for

P
Proponent Argues TRUE

Research published in the European Journal of Pathology demonstrates that AI models provide erroneous answers in nearly 70% of diagnostic cases, highlighting severe reliability issues in complex evaluations (Source 20). Although performance varies across different benchmarks, this high failure rate is strongly corroborated by broader studies showing AI systems failing accuracy tests about 70% of the time (Source 29) and exhibiting hallucination rates that can reach up to 94% (Source 10).

O
Opponent Rebuttal

The Proponent generalizes from Source 20's low-standing secondary report of a single diagnostic setting to all AI systems and evaluated cases, an ecological fallacy that the source itself cannot support. Source 29 supplies only a headline rather than methods or results, while Source 10 reports a 22%–94% range; citing its maximum as corroboration of a universal 70% rate is cherry-picking, contradicted by the low-error result in Source 16.

Argument against

O
Opponent Argues FALSE

The claim invents a universal 70% error rate despite the brief showing that performance is benchmark-, task-, and model-specific: Stanford HAI (Source 10) reports hallucination rates spanning 22% to 94%, while PLOS ONE (Source 1) found 49% correct in one clinical-diagnosis evaluation—not a 70% failure result. Direct counterexamples make the categorical figure untenable: Vectara's leaderboard (Source 16) reports a 1.8% hallucination rate for one model, and the FACTS report summarized in Source 18 places the best systems at roughly one-third incorrect, so no evidence establishes that AI systems generally produce incorrect answers in 70% of evaluated cases.

P
Proponent Rebuttal

The Opponent cherry-picks narrow summarization metrics from github.com (Source 16) to obscure the reality that in rigorous, high-stakes evaluations, the 70% failure rate is explicitly documented. By ignoring both wsu.edu (Source 29), which directly confirms AI systems fail accuracy tests approximately 70% of the time, and ai2med.eu (Source 20), which corroborates this exact 70% error rate in complex diagnostic cases, the Opponent fails to refute the specific evaluations where this claim is demonstrably true.

Panel Review

3 specialized AI experts evaluated the evidence and arguments.

Reviewer 1 — The Logic Examiner

Focus: Inferential Soundness & Fallacies
Mostly False
3/10

Source 20 reports a nearly 70% error rate only in a particular diagnostic-case evaluation, whereas Sources 1, 10, 16, and 18 show materially different results across models, tasks, and benchmarks, so that result cannot validly establish a general 70% rate for AI systems. The claim is Mostly False because the evidence demonstrates benchmark-dependent error rates rather than the asserted uniform rate, although a narrowly defined evaluation may indeed yield a result near 70%.

Logical fallacies

  • The proponent commits a hasty generalization by inferring a general error rate for artificial intelligence systems from a single diagnostic setting in Source 20.
  • The proponent cherry-picks high-error findings and the upper end of Source 10's range while disregarding evidence of substantially lower error rates in Sources 16 and 18.
  • Treating Source 29's unsupported headline as independent corroboration risks an appeal to assertion because it supplies no accessible methodology or result from which the stated conclusion can be tested.
Confidence: 9/10

Reviewer 2 — The Source Auditor

Focus: Source Reliability & Independence
False
2/10

High-quality sources, including Stanford HAI (Source 10) and PLOS ONE (Source 1), demonstrate that AI error rates vary widely depending on the model and specific task, ranging from under 2% to over 90%, rather than a universal 70% failure rate. The claim falsely generalizes a specific error rate from isolated studies to all AI systems and evaluated cases.

Weakest sources

  • Source 29 is unreliable because it provides only a headline without any supporting text, context, or methodology.
  • Source 20 is unreliable for this general claim because it reports on a single, narrow medical domain that cannot be generalized to all AI systems.
Confidence: 8/10

Reviewer 3 — The Precision Analyst

Focus: Claim Precision & Quantitative Accuracy
False
2/10

The claim asserts a single universal figure—that AI systems produce incorrect answers in 70% of evaluated cases—yet the evidence shows highly task-, model-, and benchmark-dependent rates ranging from ~1.8% (Source 16) and ~31% incorrect for top models (Sources 7, 18) through ~45–51% (Sources 1, 5, 15) to nearly 70% only in one secondary medical-diagnosis report (Source 20) and a thin headline (Source 29), with Source 10 spanning 22–94%. As worded, the unqualified 70% therefore overstates and universalizes a narrow or cherry-picked result, rendering the claim false at its stated strength.

Precision issues

  • The claim states a single universal 70% incorrect-answer rate for artificial intelligence systems across evaluated cases, but the evidence documents rates that vary widely by model, task, and benchmark rather than converging on 70%.
  • Sources that approach 70% (Source 20 nearly 70% in one diagnostic setting; Source 29 a bare headline) do not license generalizing that figure to all AI systems or all evaluated cases.
  • Counter-evidence shows substantially lower error rates, including ~31% incorrect for leading models on FACTS (Sources 7, 18), 45% on news queries (Source 15), 51% incorrect in a clinical set (Source 1), and hallucination rates as low as 1.8% (Source 16).
Confidence: 8/10

Panel summary

Source analysis finds no representative evidence establishing a 70% error rate across artificial intelligence systems. Results instead range from very low error rates to above 90%, depending on the model, task, benchmark, and scoring method. Logically, extending a near-70% result from a narrow medical evaluation to AI systems generally is a hasty generalization, reinforced by cherry-picking. Quantitatively, the unqualified figure conflicts with numerous reported rates, including approximately 31%, 45%, and 51%. Although a specific benchmark may produce a result near 70%, that narrow kernel does not support the broad statement as written.

See the full panel summary

Create a free account to read the complete analysis.

Sign up free
The claim is
False
Score: 2/10
Confidence: 8/10 Spread: 1 pt

Only you will see this note.

Embed this verification

Every embed carries schema.org ClaimReview microdata — recognized by Google and AI crawlers.

False · Lenz Score 2/10 Lenz
“Artificial intelligence systems produce incorrect answers in 70% of evaluated cases.”
29 sources · 3-panel audit · Verified Sep 2026
See full report on Lenz →