Library

6 published verifications about large language model large language model ×

“Frontier large language models achieve similar aggregate accuracy on public benchmarks.”

Mostly True

Most evidence shows frontier LLMs bunch closely together on widely used public benchmarks. Multiple independent studies and leaderboards report only small aggregate score gaps among top models, even when they disagree on individual questions. The main caveat is scope: harder or specialized public benchmarks can still separate models by meaningful margins, so the pattern is common rather than universal.

“ChatGPT is free to use for everyone.”

Mostly False

ChatGPT does offer a free plan, but it is not available to everyone. OpenAI limits access to supported countries and may block accounts outside those regions, so the claim's universal wording is materially wrong. Free users also face usage and feature limits, though those are less important than the geographic restriction.

“Using ChatGPT causes a person's brain to deteriorate.”

False

The evidence does not show that ChatGPT use causes brain deterioration. Existing studies mainly examine short-term cognitive offloading or reduced engagement during specific tasks, not lasting damage or clinical decline. Some reports also rely on media amplification of preliminary findings, while peer-reviewed evidence does not establish a general causal harm to the brain.

“Large language model hallucinations are produced by the same underlying mechanism that generates correct outputs.”

Mostly True

Both hallucinations and correct outputs do emerge from the same autoregressive next-token prediction process — no separate "hallucination engine" exists within large language models. Multiple peer-reviewed sources confirm this shared generative pipeline. However, the claim omits critical nuance: hallucinations have distinct causal drivers — such as training procedures that reward guessing over expressing uncertainty, data distribution gaps, and prompting effects — that do not equally govern correct outputs. The generation channel is shared, but the upstream conditions that produce errors are separable and require distinct mitigation strategies.

“A technology executive used ChatGPT to help develop a personalized cancer vaccine for his dog, which had been diagnosed with cancer.”

Mostly True

The core claim is accurate: Sydney-based tech professional Paul Conyngham used ChatGPT — alongside other AI tools — to help plan and develop a personalized mRNA cancer vaccine for his dog Rosie after her cancer diagnosis. However, "technology executive" is a loose description (sources call him a tech entrepreneur, AI consultant, or data engineer), and ChatGPT's role was primarily as a research and planning assistant — human scientists at UNSW performed the actual genome sequencing, vaccine synthesis, and treatment.

“AI chatbots, such as ChatGPT, provide medical advice that is consistently reliable and safe for users.”

False

The claim that AI chatbots like ChatGPT provide "consistently reliable and safe" medical advice is not supported by the evidence. Multiple high-quality studies from 2024–2026 show ChatGPT gave incorrect advice in over 51% of medical emergencies, exhibited hallucination rates of 50–82%, and correctly identified conditions in fewer than 34.5% of real-world cases. ECRI designated AI chatbot misuse as the top health technology hazard for 2026. While chatbots show promise in narrow, controlled tasks, their performance is neither consistent nor safe for general medical advice.