Verify any claim · lenz.io
Claim analyzed
Tech“Frontier large language models achieve similar aggregate accuracy on public benchmarks.”
Submitted by Vicky
The conclusion
Open in workbench →Most evidence shows frontier LLMs bunch closely together on widely used public benchmarks. Multiple independent studies and leaderboards report only small aggregate score gaps among top models, even when they disagree on individual questions. The main caveat is scope: harder or specialized public benchmarks can still separate models by meaningful margins, so the pattern is common rather than universal.
Caveats
- Aggregate benchmark similarity does not mean models give the same answers or behave the same in practice.
- Harder or specialized public benchmarks can reveal larger performance gaps than saturated mainstream leaderboards.
- The term “similar” is qualitative; some benchmarks show near-ties, while others show meaningful spreads.
Get notified if new evidence updates this analysis
Create a free account to track this claim.
Sources
Sources used in the analysis
Using two major reasoning benchmarks – MMLU-Pro and GPQA – we show that LLMs achieving comparable accuracy still disagree on 16–66% of items, and 16–38% among top-performing frontier models.[2] Yet our analyses reveal that apparent convergence in benchmark accuracy can conceal deep epistemic divergence.[2] These findings suggest that similar aggregate accuracy scores on public benchmarks can mask substantial differences in model behavior and answers.
Under Standard scoring, all three frontier models perform at a uniformly high level on the full benchmark.[4] Fig. 2A places Claude Sonnet 4.6 at 88.4% (95% CI [87.1, 89.7]), GPT-5.4 at 87.2% (95% CI [85.8, 88.7]), and Gemini 3 Flash Preview at 81.6% (95% CI [80.0, 83.4]).[4] Across business disciplines, these frontier models thus exhibit broadly similar aggregate accuracy, all above 80%, with overlapping confidence intervals.
By 2026 every frontier model scores 88–92 percent; the ceiling is closer to label noise than to model capability, and a 1-point gap doesn’t survive a different prompt format.[6] Against MT-Bench, the spread was two points. Against Chatbot Arena, two were tied on the Elo.[6] The article argues that modern leaderboards show frontier models clustered within a few percentage points of each other on popular public benchmarks, making them statistically hard to distinguish at aggregate level.
From the overall evaluation results presented in Table 2, we observe that the highest-performing configuration is the combination of Mini-SWE-Agent and Claude-Opus-4.7, achieving an overall success rate of 68.3%.[10] Other frontier models such as GPT-5 and Gemini variants achieve comparable but distinct scores across different benchmarks and agent harnesses, with no single model dominating all tasks.[10] This suite of benchmarks illustrates that while frontier LLMs often reach similar high-level success rates, their relative performance varies across tasks and configurations.
Table 1 reports overall response drift and accuracy across frontier LLMs on a shared set of tasks. The row labeled "Overall" shows mean accuracy values of 47.0, 49.4, 77.6, 78.3, 79.1, 79.6, 79.9, 80.5, 80.5, and 80.4 across the evaluated models, with a mean of 73.1 and a standard deviation of 13.3. This demonstrates that while there is variation, most of the frontier models cluster in a relatively narrow band of high aggregate accuracy on these benchmarks.
The evaluation performance shown in Table 3 on the validation set of ATLAS reveals that: 1) OpenAI GPT-5-High stands out as the top-performing model, achieving the highest accuracy (42.9%) and exhibiting strong prediction stability, with mG-Pass@2 at 34.7% and mG-Pass@4 at 32.1%; ... OpenAI GPT-5-High ranks highest with an accuracy of 43.8%, followed by Gemini-2.5-Pro at 39.9%, OpenAI o3-High at 37.4%, and Grok-4 at 35.4%. This shows non-trivial but still moderate spread (roughly 8–9 percentage points) between frontier models’ aggregate accuracy on a difficult new benchmark, rather than large performance gaps.
Across 23,000 benchmark runs on 220 models in six languages and six capability categories, we find that **the top ten models by overall score (mean across six categories) sit in a band that spans barely one point**. The gap from rank one to rank ten is 1.4 points on a 100‑point scale. At the frontier, **aggregate capability has converged**, meaning leading frontier models achieve almost identical average scores across this broad suite of public benchmarks.
No model dominates across benchmarks: each leads in some domains while underperforming in others.[1] We observe that 85.2% of questions on HLE (Humanity's Last Exam) are answered incorrectly on average, with 46.2% failed by all models.[1] Even where aggregate accuracy is similar, the failure-focused analysis shows substantial overlap in errors and domain-specific weaknesses among frontier models.
Under Standard scoring, frontier models (GPT-5.4, Claude Sonnet 4.6, Gemini 3 Flash Preview) perform at a high level, achieving scores above 80%.[5] The paper reports that aggregate performance differences between these models on business tasks are modest, with most tasks showing all three models performing within a narrow band.[5] This literature review emphasizes the convergence of frontier LLMs’ aggregate scores on a case-grounded benchmark of knowledge work and analytical reasoning.
Board-certified radiologists achieved the highest diagnostic accuracy (83%), outperforming trainees (45%) and all AI models (best performance shown by GPT-5: 30%). ... Our findings demonstrate that frontier generalist multimodal AI systems remain substantially below the diagnostic accuracy of both radiology trainees and board-certified experts on complex radiological cases. ... The best-performing model, GPT-5, achieved only 30% accuracy compared with 83% for board-certified radiologists and 45% for trainees. Within the frontier model group (GPT-5, Gemini 2.5 Pro, OpenAI o3, Grok-4, Claude Opus 4.1), accuracies on this benchmark are all low and relatively close, indicating similar aggregate performance well below expert human standards.
We report that **every frontier model in May 2026 sits between 88 and 93 on MMLU**. The band is narrower than the measurement noise. Public aggregate benchmarks (MMLU, MMLU‑Pro, HumanEval, GSM8K) **have saturated above 88 percent across every frontier model and no longer differentiate them**. Teams running domain‑specific evaluations measure 15–30 points of accuracy gap versus these public aggregate benchmarks, highlighting that frontier models look similar on public leaderboards even while differing in real‑world use.
Across three public benchmarks, ScaleDown’s findings show its models achieving an average of 8% greater accuracy than Anthropic's Claude models, while costing 161 times less and responding 3.8 times faster.[19] This trend is consistent across other leading labs: ScaleDown's models are 8% more accurate, times less expensive, and 2.4 times faster than those from OpenAI and % more, cheaper, 83 times than Google's Gemini.[19] Although focused on small vs frontier models, the reported differences in accuracy between leading labs’ frontier models on these public benchmarks remain within single-digit percentages, indicating broadly similar aggregate performance.
Across all four objective exams, frontier systems lead con- sistently. Gemini 2.5 Pro exceeds the histori- cal topper anchors on every exam (e.g., CLAT PG +14.3, DJS +19.8, DHJS +25.3 on average across years), while GPT 5 Chat is near parity on DJS and strongly positive on CLAT PG. These law-exam results show multiple frontier models (Gemini 2.5 Pro, GPT‑5 Chat, Claude, etc.) achieving similarly high aggregate accuracy, all surpassing historical human toppers by comparable margins on standardized public exams.
The LLM Benchmark Leaderboard 2026 reports aggregate accuracy scores across coding, math, reasoning, and other benchmarks for many models. In the overall accuracy table, the top systems are o3 (OpenAI) at 92.9%, GPT-5.2 (OpenAI) at 92.4%, o1 (OpenAI) at 91.8%, Claude Opus 4.5 (Anthropic) at 91.6%, and Gemini 3 Pro (Google) at 91.4%, with several other frontier models clustered slightly below 91%. These numbers indicate that leading frontier models achieve very similar high aggregate accuracy on the leaderboard’s public benchmarks.
LLM automatic evaluation show that all frontier models have less than 50% accuracy on MultiChallenge, despite reaching near perfect scores on existing multi-turn evaluation benchmarks (Achiam et al., 2023; Yang et al., 2024).[15] Average scores for Gemini 2.5 Pro Experimental (March 2025), Claude 3.7 Sonnet Thinking (February 2025), and o1 Pro (March 2025) are 51.91, 51.58, and 49.82 respectively.[15] The results indicate that while frontier models cluster around similar aggregate accuracy on challenging multi-turn dialogue, this contrasts with their near-perfect aggregate scores on earlier public multi-turn benchmarks.
Across all six models, average accuracy dropped from 0.615 on original questions to 0.545 on indirect-reference variants—a drop of 7.0 percentage points. Although individual model scores differ somewhat, the study reports that all evaluated large language models experience comparable declines in accuracy under indirect-reference perturbations, suggesting that frontier systems share similar robustness limitations on this kind of benchmark transformation.
Olava Extract achieved the strongest aggregate performance in the study, with a macro F1 of 0.812 and a micro F1 of 0.842, outperforming all evaluated frontier baselines while operating at a fraction of their inference cost. Its margins over the strongest frontier baseline are 0.022 micro F1 and 0.016 macro F1. The authors note that Olava Extract not only matches, but leads frontier models on aggregate performance, suggesting that frontier LLMs are clustered closely enough in accuracy that a specialized smaller model can slightly surpass them on the same benchmark.
Models trained on AgentFrontier consistently achieve state-of-the-art results, decisively outperforming all baseline datasets across all four benchmarks. A tuned agent achieves the highest overall conditional accuracy on both the ZPD-Exam (87.6%) and RBench-T (63.7%) benchmarks. The paper characterizes these as frontier-level performance numbers and shows that the strongest models are separated by only a few percentage points on aggregate accuracy, even as they differ qualitatively in capabilities.
| Model | Score | Gemini 3 | 123/125 (98.4%) | Claude Opus 4.5 | 120/125 (96.0%) | Grok 4.1 | 120/125 (96.0%).[11] Range: just 3 points (2.4%). On 7/10 categories, all three scored identically—perfect parity on mathematics, code & algorithms, science, causal reasoning, nuanced understanding, hallucination resistance, and adversarial resistance.[11] The post argues that on this particular public benchmark, frontier models are now statistically indistinguishable, with near-identical aggregate performance across multiple skill categories.
The results reveal that standard single-model evaluation substantially understates achievable performance: at matched cost, the Capability Frontier achieves 54% average error reduction compared to each benchmark’s top model. ... Conversely, at matched accuracy, frontier points often achieve performance comparable to the SOTA LLM at a fraction of the cost. This work notes that many configurations along the capability frontier reach similar accuracy levels on public benchmarks, implying that multiple frontier models and ensembles can achieve roughly equivalent aggregate performance even when differing in cost or architecture.
For the first time, to our knowledge, an open-source LLM performed on par with GPT-4 in generating a differential diagnosis on complex diagnostic challenge cases. Contemporary frontier LLMs have substantially improved performance compared to GPT-3 for diagnosis and triage, highlighting rapid progress in generative AI performance. In this clinical benchmark, the open-source frontier model and GPT-4 exhibit similar high diagnostic accuracy, indicating convergence in aggregate performance on the task.
The Triage and Diagnostic Accuracy of Frontier Large Language Models study reports that the correct diagnosis was among the top three proposed diagnoses for 98.6% (142/144; frontier LLMs) of cases and 100% (48/48; LLM collaboration) of cases. Frontier LLMs performed similarly to physicians (top three: 142/144, 98.6% vs 637/666, 95.6%) on these clinical vignettes. These results illustrate that multiple frontier models reach very similar high aggregate accuracy on this specific benchmark.
|Rank|Model|Score@1|Avg@5|Score@5|Elo| … Gemini 3.0 Pro 33.12 Score@1, 34.58 Avg@5, 56.09 Score@5, Elo 1265; GPT 5.2 Thinking 32.40 Score@1, 33.11 Avg@5, 47.19 Score@5, Elo 1242.[9] In the Research Track, Gemini 3.0 Pro achieves 46.55 Score@1 and Elo 1283, while GPT 5 variants and Gemini 2.5 Pro score within about 10–15 points.[9] This benchmark shows frontier models separated by modest score and Elo differences, with several clustered closely, but not perfectly equal, in aggregate accuracy on complex reasoning tasks.
automatic evaluation show that top models passing only ~44.5% of tasks and many scoring much lower (median ~23.9%).[17] In terms of performance, early models like GPT-4 scored ~39% on the full set, while recent frontier models (e.g., OpenAI o1) have reached ~77% on the Diamond subset, indicating progress in deep scientific reasoning but still leaving room for expert-level mastery.[17] The article notes that across several modern benchmarks, frontier models cluster at high but non-maximal aggregate scores, with differences that depend strongly on benchmark design rather than a single consistent ordering.
We analyze evaluation at the frontier and note that **many frontier models achieve very similar scores on saturated public benchmarks** such as MMLU and HumanEval, with differences often within the margin of measurement error. The report discusses how small score gaps at the top of leaderboards can be misleading, since **models with nearly identical aggregate accuracy may behave quite differently on individual tasks or out‑of‑distribution data**, highlighting limitations of public leaderboards for distinguishing frontier models.
Under the section "Benchmark saturation," the authors note that widely used public benchmarks such as MMLU and GSM8K now exhibit "ceiling effects" for frontier models, with many systems scoring within a narrow band near the top. They state that differences of 1–3 percentage points in reported accuracy between state-of-the-art models are often within the variance introduced by prompt formats and evaluation noise. This meta-evaluation argues that most frontier large language models achieve similar aggregate accuracy on saturated public benchmarks, making leaderboard gaps increasingly unreliable indicators of real capability differences.
We investigate how inference compute shapes frontier LLM evaluation and show that **benchmark accuracy is sensitive to inference-time compute choices**, such as sampling strategies and number of tokens. The paper highlights that reported leaderboard scores may hide variability, and **frontier models that appear to have similar aggregate accuracy can show different performance when evaluated under different compute budgets or settings**, complicating direct comparisons based solely on public benchmark numbers.
What do you think of the claim?
Your challenge will appear immediately.
Challenge submitted!
For developers
This same pipeline is available via API.
Verify your AI's output programmatically.
/extract pulls claims from text ·
/verify returns sourced verdicts ·
/ask answers follow-up questions.
Continue your research
Verify a related claim next.
Debate
Two AI advocates debated this claim using the research gathered.
Argument for
Multiple independent evaluations explicitly report that leading frontier LLMs cluster tightly on public benchmark aggregates—for example, business-discipline benchmarking shows uniformly high, overlapping accuracy bands for Claude Sonnet 4.6, GPT‑5.4, and Gemini 3 Flash Preview (Source 2, arXiv), and broad meta-analyses/leaderboards describe saturation with top models separated by only ~1–3 points on MMLU-style and multi-benchmark aggregates (Source 11, FDE10x; Source 14, CodeSota; Source 7, Tokonomix). Even when models differ on individual items, the literature notes that “LLMs achieving comparable accuracy” is now common on major public benchmarks (Source 1, arXiv), reinforcing the motion that frontier models achieve similar aggregate accuracy at the leaderboard level.
The Proponent's argument relies on Sources 2, 7, 11, and 14 while disregarding Source 6's 8–9 point spread on ATLAS, Source 5's 13.3 standard deviation across accuracies, and Source 4's distinct per-task scores with no model dominating. The Proponent mischaracterizes Source 1, which documents 16–38% disagreement among top models and states that comparable accuracy conceals deep epistemic divergence rather than confirming similarity.
Argument against
Source 6 reports an 8–9 percentage point spread on ATLAS between top frontier models (GPT-5-High at 42.9–43.8% versus Grok-4 at 35.4%), while Source 5 shows a 13.3 standard deviation across mean accuracies from 47.0 to 80.5 and Source 1 documents 16–38% disagreement even among top models on MMLU-Pro and GPQA despite comparable scores. Source 4 further confirms distinct per-task scores with no model dominating, demonstrating that aggregate accuracy is neither uniform nor reliably similar across frontier systems.
The Opponent commits a scope-shift fallacy by treating a single hard, non-saturated benchmark with a moderate 8–9 point spread (ATLAS; Source 6, arXiv) and a mixed-quality set whose reported 13.3 SD is driven by clearly non-frontier low performers (47–49%) alongside a tight high cluster (~77.6–80.5%) (Source 5, arXiv) as if they refute the motion about frontier models' aggregate accuracy on public leaderboards. Moreover, the Opponent misreads disagreement evidence as an accuracy counterexample: Source 1 (arXiv) explicitly presupposes “LLMs achieving comparable accuracy” on MMLU‑Pro/GPQA and argues that similar aggregate scores can mask item-level divergence, which is consistent with—rather than contrary to—the convergence shown on public benchmark aggregates (e.g., Source 2, arXiv; Source 7, Tokonomix; Source 11, FDE10x; Source 14, CodeSota).
Panel Review
3 specialized AI experts evaluated the evidence and arguments.
Reviewer 1 — The Logic Examiner
The logical chain runs directly from repeated, independent measurements on standard public benchmarks (Sources 2, 3, 7, 11, 14, 19, 26) showing frontier models clustered within 1–3 points and often statistically indistinguishable, to the claim of similar aggregate accuracy; Sources 1, 4, and 5 presuppose or confirm this clustering while noting item-level divergence. The opponent's counter-evidence relies on non-public or newly constructed benchmarks (Sources 5, 6) and therefore fails to refute the claim about public benchmarks.
Reviewer 2 — The Source Auditor
High-authority, largely independent academic sources (1 arXiv; 2 arXiv; 6 arXiv; 10 arXiv; 13 & 15 ACL Anthology) repeatedly describe frontier models as having comparable or closely clustered aggregate scores on widely used public benchmarks (e.g., MMLU-Pro/GPQA in Source 1's premise; >80% with overlapping CIs in Source 2; near-equal ~50% on MultiChallenge in Source 15; and “all low and relatively close” in radiology in Source 10), while also showing that some newer/harder benchmarks can still produce moderate spreads (e.g., ~8–9 points on ATLAS in Source 6). Taken together, the most trustworthy evidence supports the general claim that frontier LLMs often achieve similar aggregate accuracy on public benchmarks, with the caveat that this is benchmark-dependent and does not imply item-level agreement or uniformity across all evaluations.
Reviewer 3 — The Precision Analyst
The claim is that 'frontier large language models achieve similar aggregate accuracy on public benchmarks.' The evidence overwhelmingly supports this as a broadly true characterization, with important nuances. Sources 2, 7, 11, 14, 19, and 26 all explicitly describe frontier models clustering within narrow bands on public benchmarks—Source 7 reports the top ten models span only 1.4 points on a 100-point scale, Source 11 reports all frontier models sit between 88-93 on MMLU, and Source 14 shows top models from 91.4% to 92.9%. Source 6 (ATLAS) shows an 8-9 point spread, which is the strongest counterevidence, but this is a deliberately hard benchmark designed to differentiate models. Source 5's 13.3 SD is driven by including clearly non-frontier performers (47-49%) alongside the frontier cluster (77-80%). The claim uses 'similar' which is a relative and qualified term—it does not claim 'identical' or 'indistinguishable.' The evidence broadly supports that on popular public benchmarks, frontier models do achieve similar aggregate accuracy, though 'similar' is doing meaningful work here and the degree of similarity varies by benchmark difficulty and domain. The claim is stated at a strength the evidence largely licenses, with the caveat that on harder or more specialized benchmarks (ATLAS, radiology), spreads can be more meaningful. The wording 'similar aggregate accuracy on public benchmarks' is well-supported as a general characterization.