63%
of real-world fact-checks, top AI models don't agree on the answer
1,000 claims, rated by 5 frontier models · version v1.1, published
Beyond Benchmarks: Disagreement Among Frontier LLMs on Real-World Fact-Checks
Abstract
Frontier large language models achieve comparable accuracy on public benchmarks, a parity often read as evidence that they are interchangeable as assessors of factual claims. We test that assumption directly, on claims not drawn from any public benchmark. Five frontier models were each asked to adjudicate 1,000 real-world claims submitted by users to a fact-checking platform. They were tasked with assigning every claim a verdict on a five-point scale from True to False, as well as reporting their confidence in each answer. On the 997 claims where all five models returned a usable verdict, they fail to reach consensus on 63%, and on 23% the two furthest-apart verdicts differ by at least two verdict categories. The ordinal Krippendorff’s α of 0.77 reflects structured but far from interchangeable judgement. We find that disagreement concentrates in the intermediate verdicts, where claims resist clean adjudication: the definitive poles are unanimous about half the time, against roughly one in ten for the intermediate verdicts. In relation to disagreement, our results show that model confidence is not a reliable indicator of whether the panel will agree. The models are highly confident almost everywhere, rating 76% of answers 9 or 10 on a 1–10 confidence scale, yet they agree with one another on confidence markedly less than on the verdicts themselves (α = 0.44). The verdict that is given to an assertion will then depend heavily on which model one consults, which matters a great deal as more people turn to such models to verify information.
Versions
| Version | Published | Status | What changed | Files |
|---|---|---|---|---|
| v1.1 | Current | Five-point verdict scale, a refreshed five-model panel, and every model run with retrieval and thinking capabilities. Rates are not comparable with v1.0. | Page PDF Source CSV DOI | |
| v1.0 | Superseded | First release. Four-point verdict scale, single-label prompt. | Page PDF CSV DOI |