Claim analyzed

Tech

“When asked to state their confidence, frontier large language models hallucinate in more than half of their answers.”

The conclusion

Mostly False
3/10

Available evidence does not establish that frontier models hallucinate in most answers when asked for confidence. Some models exceed 50% on particular benchmarks, but others—including reported evaluations of GPT-4o and Llama 3.1 405B—remain below it. Many cited measurements also omit confidence elicitation or examine only high-confidence subsets, so they cannot support the claim's broad causal framing.

Caveats

  • The claim generalizes from selected models and benchmarks to frontier models as a class.
  • Several cited hallucination rates were measured without asking models to state confidence.
  • Rates are not directly comparable because studies handle refusals and define hallucinations differently.

Sources

Sources used in the analysis

#1
arxiv.org 2025-11-17 | AA-Omniscience: Evaluating Cross-Domain Knowledge Reliability in Large Language Models

High hallucination is the dominant factor driving these low scores. For instance, although Grok 4 and GPT-5 (high) record the highest accuracy at 39%, their hallucination rates of 64% and 81% result in substantial penalties on the Omniscience Index (Figure 7).

#2
ar5iv.labs.arxiv.org 2025-06-02 | [2506.07309] ConfQA: Answer Only If You Are Confident

Whereas higher confidence generally correlates with higher answer accuracy, the confidence is poorly calibrated: when Llama-3.1-70B predicts a confidence score above $80%80\%$ on CRAG (Yang et al., 2024), the real accuracy is only $33%33\%$.

#3
arxiv.org 2026-06-05 | A Unified Benchmark for Hallucination Detection across ...

Within the black-box family, direct verbalized confidence is generally weaker than sample-consistency signals. For example, on Llama-3.2-3B-Instruct, Verbalized Confidence obtains an overall average of 57.99, while SelfCheckGPT-BERTScore, SelfCheckGPT-NLI, and lexical similarity reach 69.25, 67.86, and 69.16, respectively. This suggests that comparing stochastic generations is generally more informative than relying on the model’s stated confidence, although the usefulness of such consistency signals still depends on the task and comparison metric.

Despite its high refusal rate, the Llama3.1-8B model is less prone to hallucination (48.37%) compared to similar-sized models like Qwen2.5 7B (85.22%) and Mistral 7B (81.19%). … We further examine the hallucination rate when models do not refuse to answer. The Llama 3.1 405B Instruct model achieves the lowest hallucination rate at 26.84% but falsely refuses to answer 56.77% of the time. … GPT-4o, with a 45.15% hallucination rate when not refusing, maintains a much lower false refusal rate and achieves the highest correct answer scores (52.59%), indicating a trade-off between precision and recall.

#5
arxiv.org 2026-04-03 | : A Decision-Theoretic Approach to Evaluating Large Language Model Confidence

Our results reveal substantial variation in confidence reliability, and while larger and more accurate models tend to also achieve better reliability metrics, even frontier models remain prone to severe overconfidence in challenging settings.

Large language models (LLMs) frequently produce confident but incorrect answers, partly because common binary scoring conventions reward answering over honestly expressing uncer tainty. … In this setting, we operationalize hallucination as giving a false answer to such a question and measure it with the false-answer rate (FAR). … As a baseline, directly prompting GPT-5 mini without reward framing or verbal confidence (Pure Eval) yields FARanswered = 52.3%.

#7
p.rst.im 2023-12-06 | The Curious Case of Hallucinatory (Un)answerability: Finding Truths in the Hidden States of Over-Confident Large Language Models

We found ample evidence for LM’s ability to en code the (un)answerability of questions, despite the fact that models tend to be over-confident and gen erate hallucinatory answers when presented with (un)answerable questions.

#8
proceedings.iclr.cc 2024-01-01 | UNCERTAINTY?

LLMs, when verbalizing their confidence, tend to be overconfident, potentially imitating human patterns of expressing confidence. … While GPT-4 displays lower ECE, its AUROC and AUPRC-Negative scores remain suboptimal, with an average AUROC of merely 62.7%—close to the 50% random guess threshold—highlighting challenges in distinguishing correct from incorrect predictions.

#9
arxiv.org 2024-06-19 | Factual Confidence of LLMs: on Reliability and Robustness of Current Estimators

Our experiments across a series of LLMs in dicate that trained hidden-state probes provide the most reliable confidence estimates, albeit at the expense of requiring access to weights and training data. … In the verbalization (verbalized confidence) meth ods (Lin et al., 2022a; Xiong et al., 2024; Yin et al., 2023; Tian et al., 2023; Kadavath et al., 2022), the model is directly prompted to report its confidence level (e.g., “How confident are you that the answer is correct?”). … With the exception of Falcon-40B instruct, the other methods perform close to or below chance (depending on the model’s label distribution, chance level varies between 0.11 and 0.27). This indicates that non-trained estima tors are generally not reliable for P(IK) despite being frequently used in the literature.

#10
arxiv.org 2023-11-24 | Calibrated Language Models Must Hallucinate

In the case of 5W, one could imagine that half of the posts occur exactly once. This would suggest that a calibrated-factoid model would hallucinate on about half of its generations on 5W factoids.

#11
arxiv.org 2024-11-07 | Measuring short-form factuality in large language models

However, the fact that performance is well below the line y = x means that models consistently overstate their confidence. Hence, there is a lot of room to improve the calibration of large language models in terms of stated confidence.

#12

It is challenging for LLMs to accurately express their internal confidence in natural language. … When expressing confidence in words, LLMs are not well-calibrated and tend to be overconfident which is consistent with the previous findings [17].

#13
nature.com 2026-04-22 | Evaluating large language models for accuracy ...

For example, on the SimpleQA evaluation 14, accuracy slightly favours OpenAI’s o4-mini, which answers almost all questions (with over 3/4 error rate) 15 over GPT-5-mini, even though GPT-5-mini makes many fewer errors (owing to abstentions) 6.

#14
aclanthology.org 2025-11-04 | Trust Me, I’m Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer

A non-negligible amount of hallucinations despite knowledge (16–43%) occur with high certainty, demonstrating the existence of CHOKE examples across all combinations of certainty methods, models, and prompt settings.

#15
aclanthology.org 2026-07-02 | Beyond Output Confidence: Epistemic-Aware Hallucination Detection with Answer-Level Signals

This fluency, however, coexists with a systematic tendency to hallucinate: LLMs often produce statements that contradict ver ifiable facts while maintaining high predictive con fidence (Ji et al., 2023; Liu et al., 2025).

#16
arxiv.org 2026-05-04 | HalluScan: A Systematic Benchmark for Detecting and Mitigating Hallucinations in Instruction-Following LLMs

Self-Evaluation leverages the model's introspective capabilities by prompting it to assess the confidence and factual correctness of its own output using Chain-of-Thought (CoT) reasoning [49]. … Self-Evaluation (SE) achieves the highest pooled AUROC of 0.688 across all 600 model-response pairs.

#17

Even recent models, such as OpenAI’s GPT-4.5, have been found to hallucinate as often as 37.1% on certain benchmarks (OpenAI, 2025), underscoring the ongoing challenge of ensuring reliability in LLM outputs.

#18
arxiv.org 2026-11-07 | Do LLM Recommenders Know When They’re Hallucinating? Auditing Confidence Calibration in Catalog Faithfulness

Each model holds a near-constant confidence level barely responsive to the catalog, while the catalog-hit rate swings 60 points, so the sign of the error is set by where a model’s constant lands against a catalog’s accuracy: 7 of the twelve cells are under-confident and 5 over-confident, all four under-confident on MovieLens, all four over-confident on Amazon Toys. … We read this as an elicitation mismatch: “Just Ask” elicits a generic quality rating, not a catalog-membership probability.

#19
doi.org 2026-02-18 | Know When You're Wrong: Aligning Confidence with Correctness for LLM Error Detection

Large language models frequently generate plausible but incorrect outputs with unwarranted confidence—a phe nomenon commonly termed “hallucinations” (Ji et al., 2023; Maynez et al., 2020; Zhang et al., 2025).

#20
arxiv.org 2026-08-11 | Calibrating Verbal Uncertainty as a Linear Feature to Reduce Hallucinations

We find that “verbal uncertainty” is governed by a single linear feature in the representation space of LLMs, and show that this has only moderate correlation with the actual “semantic uncertainty” of the model. … The miscalibration of semantic and verbal uncertainty triggers overconfident hallucinations.

#21
arxiv.org 2025-09-01 | Trusted Uncertainty in Large Language Models: A Unified Framework for Confidence Calibration and Risk-Controlled Refusal

Yet modern neural predictors are notoriously miscalibrated, often assigning high confidence to wrong answers, a phenomenon documented from classic discriminative models through contempo rary deep networks [1–6].

#22
doi.org 2026-05-02 | Hallucinations Undermine Trust; Metacognition is a Way Forward

While LLMs are currently poor at faithfully conveying their uncertainty, we believe this problem offers tangible headroom. … Yona et al. (2024) have demonstrated that current state-of-the-art models are far from satisfying this desideratum; they typically express high linguistic confidence even when their intrinsic uncertainty is low.

#23
ar5iv.labs.arxiv.org 2025-03-06 | [2505.24858] MetaFaith: Faithful Natural Language Uncertainty Expression in LLMs

When no uncertainty prompt is used (none), all models perform poorly with cMFG scores close to or less than 0.5, indicating a tendency toward worse faithfulness than when a random level of decisiveness is exhibited. … Models often did not generate any expressions of uncertainty, instead producing highly decisive answers with mean decisiveness near 1.0 even when very uncertain, indicating baseline uncertainty expressions are highly unreliable.

#24
arxiv.org 2026-08-09 | Claim-Level Confidence Calibration for Reliable Decision Making with Large Language Models

Large Language Models (LLMs) increasingly support decision-making in high-stakes domains, but they often hallucinate and express confidence that is misaligned with factual correctness.

#25
arxiv.org 2025-09-04 | Layer-0 Suppressors Ground Hallucination Inevitability: A Mechanistic Account of How Transformers Trade Factuality for Hedging

Under binary grading, abstaining is strictly sub-optimal. IDK-type responses are maximally penalized while an overconfident “best guess” is optimal. … Observation 1. Let c be a prompt. For any distribution ρc over binary graders, the optimal response(s) are not abstentions, i.e., A c ∩ arg max r∈R c Egc∼ ρc[gc(r)] = ∅ .

#26
arxiv.org 2025-10-05 | Large Language Models Hallucination: A Comprehensive Survey

Uncertainty-based detection addresses the challenge of data dependency using model confidence rather than external labeled datasets. However, its effectiveness is highly sensitive to the calibration of uncertainty thresholds, and it often fails to detect hallucinations when the model shows high confidence in an incorrect response.

#27
p.rst.im 2024-08-11 | Factual Confidence of LLMs: on Reliability and Robustness of Current Estimators

Our experiments across a series of LLMs in dicate that trained hidden-state probes provide the most reliable confidence estimates, albeit at the expense of requiring access to weights and training data. … With the exception of Falcon-40B instruct, the other methods perform close to or below chance (depending on the model’s label distribution, chance level varies between 0.11 and 0.27). This indicates that non-trained estima tors are generally not reliable for P(IK) despite being frequently used in the literature.

#28
aclanthology.org 2024-08-11 | Factual Confidence of LLMs: on Reliability and Robustness of Current Estimators

Our experiments across a series of LLMs in dicate that trained hidden-state probes provide the most reliable confidence estimates, albeit at the expense of requiring access to weights and training data. … We find that the confidence of LLMs is often unstable across semantically equivalent inputs, suggesting that there is much room for improvement of the sta bility of models’ parametric knowledge.

#29
doi.org 2024-10-06 | Confidence Elicitation Improves Selective Generation in Black-Box Large Language Models

An important research topic is how to make LLMs accurately express the confidence to their answers, so they can refrain from outputting or regenerate output in cases of low-confidence predictions. … Currently, research on eliciting calibrated confidence from LLMs is still insufficient.

#30
doi.org 2025-12-22 | Mitigating LLM Hallucination via Behaviorally Calibrated Reinforcement Learning

The phenomenon of hallucination—where models confabulate facts with high confidence—remains a stubborn artifact of the current post-training paradigm [16].

#31
aclanthology.org 2025-11-04 | Calibrating Verbal Uncertainty as a Linear Feature to Reduce Hallucinations

We find that “verbal 1 uncertainty” is governed by a single linear feature in the representation space of LLMs, and show that this has only moder ate correlation with the actual “semantic uncer tainty” of the model. … We propose a novel hallucination detection method by incorporating VU and SU, outperforming detection methods that rely solely on SU. Next, we propose a mitiga tion method, Mechanistic Uncertainty Calibration (MUC), that steers LLM activations along VUF to calibrate VU with SU.

#32

The empirical results suggest that ChatGPT is likely to generate hallucinated content in specific topics by fabricating unverifiable information (i.e., about $11.4\%$ user queries).

#33
proceedings.neurips.cc Reasoning Models Better Express Their Confidence

A persistent weakness of large language models (LLMs) is their tendency to sound confident even when they are wrong [13, 45]. … Generally, we observe that reasoning models still prefer to express high confidence, assigning values below 55% infrequently, as shown in Figure 2 (right).

#34
arxiv.org 2025-08-25 | LLMs Hallucinate with Certainty Despite Knowing the Answer
#35
openai.com 2025-09-05 | Why language models hallucinate | OpenAI

By this we mean instances where a model confidently generates an answer that isn’t true. Our new research paper⁠(opens in a new window) argues that language models hallucinate because standard training and evaluation procedures reward guessing over acknowledging uncertainty.

#36
openai.com 2024-10-30 | Introducing SimpleQA | OpenAI

One way to measure calibration is to directly ask the language model to state its confidence in its answer using the prompt: “Please give your best guess, along with your confidence as a percentage that that is the correct answer.” Then we can plot the correlation between the stated confidence of the model, and how accurate the model actually was. … However, the fact that performance is well below the line y=x means that models consistently overstate their confidence. Hence, there is a lot of room to improve the calibration of large language models in terms of stated confidence.

#37
forbes.com 2024-11-05 | OpenAI Newly Released SimpleQA Helps Reveal That Generative AI Blatantly And Alarmingly Overstates What It Knows

OpenAI Newly Released SimpleQA Helps Reveal That Generative AI Blatantly And Alarmingly Overstates What It Knows

#38
arxiv.org Trust Me, I’m Wrong: High-Certainty Hallucinations in LLMs
#39
#40
ar5iv.labs.arxiv.org [2502.11028] Mind the Confidence Gap: Overconfidence, Calibration, and Distractor Effects in Large Language Models
#41
arxiv.org PARALLAX: Separating Genuine Progress from Benchmark Artifacts in Hallucination Detection
#42
#43
academic.oup.com Profiling Legal Hallucinations in Large Language Models
#44
arxiv.org Uncertainty Quantification and Confidence Calibration in Large Language Models: A Survey
#45
arxiv.org Frontier Models Can Take Actions at Low Probabilities

When target rates are given, larger models achieve good calibration at lower rates. Yet, when models must derive the optimal target rate themselves, all models fail to achieve calibration without entropy or hint to generate it.

Debate

Two AI advocates debated this claim using the research gathered.

Argument for

P
Proponent Argues TRUE

Direct empirical measurements confirm the claim: Source 1 shows frontier models Grok 4 and GPT-5 (high) hallucinate at 64% and 81% rates respectively even when reporting confidence, while Source 6 finds GPT-5 mini's false-answer rate under direct confidence prompting reaches 52.3%, and Source 2 shows Llama-3.1-70B's high-confidence claims (>80%) are correct only 33% of the time—all squarely exceeding the 'more than half' threshold. This pattern is corroborated across independent studies including Source 4 (Qwen2.5 7B at 85.22% hallucination, Mistral 7B at 81.19%), Source 10's theoretical proof that calibrated models 'hallucinate on about half of its generations,' and Source 11/36's SimpleQA findings that models 'consistently overstate their confidence,' forming a robust consensus that frontier LLMs' verbalized confidence fails to track truth in a majority of answers.

O
Opponent Rebuttal

The Proponent misstates Source 6: its 52.3% false-answer rate is explicitly the "Pure Eval" baseline without verbal confidence, while Source 1 reports benchmark-specific rates rather than establishing that confidence elicitation makes frontier models hallucinate in a majority of all answers. The Proponent also commits a hasty generalization by treating smaller Qwen and Mistral results and Source 10's conditional theoretical illustration as proof about frontier models, despite Source 4 reporting sub-50% among-answer hallucination rates for frontier Llama 3.1 405B Instruct (26.84%) and GPT-4o (45.15%).

Argument against

O
Opponent Argues FALSE

The claim's universal “frontier models” and “more than half” threshold is contradicted by Source 4 (arxiv.org), which reports hallucination rates below 50% for Llama 3.1 405B Instruct (26.84%) and GPT-4o (45.15%) among answers, demonstrating that a majority-hallucination rate is not a general frontier-model property. Moreover, Source 6's 52.3% false-answer result for GPT-5 mini explicitly comes from a baseline without verbal confidence, while Source 1's high rates are model- and benchmark-specific, so neither establishes that asking models to state confidence causes more than half of their answers to hallucinate.

P
Proponent Rebuttal

The Opponent's reliance on Source 4's sub-50% figures ignores that those numbers describe only non-refused answers under a specialized refusal-tuned setup, not the general behavior of frontier models under direct confidence elicitation—cherry-picking favorable outliers while ignoring Source 1's Grok 4 (64%) and GPT-5 high (81%), Source 4's own Qwen2.5 (85.22%) and Mistral (81.19%), and Source 2's 33% accuracy at >80% stated confidence, all of which converge on majority-hallucination. The Opponent's claim that Source 6's 52.3% is inapplicable is a distinction without difference, since the paper explicitly frames this as the baseline confidence-elicitation condition against which calibration interventions are measured, meaning it directly represents the 'stating confidence' scenario the claim addresses.”}}}}}}}}}} 🚨 fix formatting.,

Panel Review

3 specialized AI experts evaluated the evidence and arguments.

Reviewer 1 — The Logic Examiner

Focus: Inferential Soundness & Fallacies
Mostly False
3/10

Sources 1 and 4 report benchmark-specific hallucination rates without establishing that the models were asked to state confidence, Source 6 explicitly uses a baseline without verbal confidence, and Source 2 concerns only the subset of answers rated above 80% confidence rather than all answers. Thus, although the evidence supports that verbal confidence is often miscalibrated, it does not validly prove the claim that frontier models hallucinate in more than half of answers when confidence is requested, and Source 4 supplies relevant counterinstances to any general majority-rate inference.

Logical fallacies

  • The proponent commits a hasty generalization by extrapolating selected models, benchmarks, and high-confidence subsets to frontier models generally.
  • The proponent conflates hallucination rates measured without confidence elicitation with rates obtained when models are asked to state confidence.
  • The proponent treats overconfidence and poor calibration as proof that more than half of all answers are hallucinations.
Confidence: 8/10

Reviewer 2 — The Source Auditor

Focus: Source Reliability & Independence
Mostly False
4/10

Reliable sources consistently show that while LLMs are overconfident and poorly calibrated when stating their confidence, the hallucination rate is not universally 'more than half' across all frontier models. Source 4 explicitly notes that GPT-4o and Llama 3.1 405B have hallucination rates below 50% when not refusing, contradicting the claim's absolute threshold.

Confidence: 8/10

Reviewer 3 — The Precision Analyst

Focus: Claim Precision & Quantitative Accuracy
Mixed
5/10

The claim asserts a universal, precise quantitative threshold ('more than half') tied to a specific elicitation condition ('when asked to state their confidence') for 'frontier LLMs' as a class. The evidence is genuinely mixed on this exact operationalization: Source 1 shows Grok 4/GPT-5(high) hallucinate well above 50%, and Source 6's Pure Eval baseline is 52.3%, but Source 4 shows frontier-tier models (Llama 3.1 405B at 26.84%, GPT-4o at 45.15%) falling below 50% among non-refused answers, and most other sources (2,3,5,7-33) discuss overconfidence/miscalibration generally without giving a clean 'more than half' figure specifically tied to confidence-elicitation prompts across a defined set of 'frontier' models. The claim's scope ('frontier LLMs' generically) and its specific numeric threshold are not uniformly supported—some frontier models exceed 50%, others don't, and the benchmarks/definitions of hallucination vary widely (accuracy-conditioned, refusal-adjusted, FAR baselines), so the claim overgeneralizes a real but inconsistent pattern into a settled universal quantitative fact.

Precision issues

  • The claim uses an unqualified 'frontier large language models' scope, but Source 4 shows frontier-tier models like GPT-4o (45.15%) and Llama 3.1 405B (26.84%) falling below the 50% threshold, contradicting a universal claim.
  • The 'more than half' threshold is benchmark- and definition-dependent; sources measure hallucination differently (among all answers vs. among non-refused answers vs. false-answer-rate baselines), so the same numeric threshold cannot be applied uniformly across studies.
  • Source 6's 52.3% figure is explicitly a 'Pure Eval' baseline without verbal confidence elicitation, which weakens its direct relevance to the claim's specific framing of 'when asked to state their confidence.','Source 10's theoretical illustration ('about half') is a conditional, hypothetical calculation for a specific factoid distribution, not an empirical measurement of frontier LLM behavior, yet it is cited as corroborating evidence for the majority-hallucination claim.
Confidence: 6/10

Panel summary

Source analysis supports widespread overconfidence and poor calibration but does not support a general majority hallucination rate under confidence elicitation. Some benchmark results exceed 50%, while reported rates for GPT-4o and Llama 3.1 405B fall below that threshold. The central inference is unsound because results obtained without verbal confidence prompts, or only from high-confidence subsets, cannot establish what happens across all answers when confidence is requested. Quantitative comparisons are further weakened by differing benchmarks, refusal handling, and hallucination definitions. The claim retains a kernel of truth because some frontier models exceed 50% in specific evaluations, but it overgeneralizes those results to frontier models as a class and ties them to an elicitation condition that much of the evidence did not test.

See the full panel summary

Create a free account to read the complete analysis.

Sign up free
The claim is
Mostly False
Score: 3/10
Confidence: 7/10 Spread: 2 pts

Only you will see this note.

Embed this verification

Every embed carries schema.org ClaimReview microdata — recognized by Google and AI crawlers.

Mostly False · Lenz Score 3/10 Lenz
“When asked to state their confidence, frontier large language models hallucinate in more than half of their answers.”
45 sources · 3-panel audit
See full report on Lenz →