10 verifications about Large Language Models Large Language Models ×
“Large language models generate text by predicting likely word sequences from patterns learned during training rather than inherently retrieving verified facts from a database.”
The description accurately captures how standard large language models generate text. They predict tokens from learned statistical patterns and may encode factual associations in their parameters, but they do not inherently consult a verified factual database. External retrieval systems can add database or document access, while “word sequences” is a reasonable simplification of token prediction.
“OpenAI's ChatGPT reached one million users within five hours of its launch.”
ChatGPT reached one million users in about five days, not five hours. Contemporaneous statements and subsequent reliable reporting consistently give the five-day timeframe, while none of the cited evidence supports five hours. The incorrect time unit creates a 24-fold error in the claim’s central quantitative assertion.
“ChatGPT's growth to 100 million monthly users within two months was the fastest adoption of any consumer application in history.”
ChatGPT did reach an estimated 100 million monthly active users roughly two months after launch and was considered the fastest-growing consumer application at that time. The enduring historical superlative is outdated because Threads later reached 100 million registered users in five days. Those measurements are not perfectly equivalent, but the claim supplies no qualification preserving ChatGPT’s former record.
“ChatGPT reached 100 million monthly users within two months of its launch.”
The milestone is well supported as a widely accepted estimate. UBS analysis based on Similarweb data placed ChatGPT at roughly 100 million monthly active users in January 2023, about two months after its November 30, 2022 launch. However, the figure was not an official OpenAI count, and extensive media repetition largely traces back to that same underlying analysis.
“Frontier large language models achieve similar aggregate accuracy on public benchmarks.”
Most evidence shows frontier LLMs bunch closely together on widely used public benchmarks. Multiple independent studies and leaderboards report only small aggregate score gaps among top models, even when they disagree on individual questions. The main caveat is scope: harder or specialized public benchmarks can still separate models by meaningful margins, so the pattern is common rather than universal.
“ChatGPT is free to use for everyone.”
ChatGPT does offer a free plan, but it is not available to everyone. OpenAI limits access to supported countries and may block accounts outside those regions, so the claim's universal wording is materially wrong. Free users also face usage and feature limits, though those are less important than the geographic restriction.
“Using ChatGPT causes a person's brain to deteriorate.”
The evidence does not show that ChatGPT use causes brain deterioration. Existing studies mainly examine short-term cognitive offloading or reduced engagement during specific tasks, not lasting damage or clinical decline. Some reports also rely on media amplification of preliminary findings, while peer-reviewed evidence does not establish a general causal harm to the brain.
“Large language model hallucinations are produced by the same underlying mechanism that generates correct outputs.”
Both hallucinations and correct outputs do emerge from the same autoregressive next-token prediction process — no separate "hallucination engine" exists within large language models. Multiple peer-reviewed sources confirm this shared generative pipeline. However, the claim omits critical nuance: hallucinations have distinct causal drivers — such as training procedures that reward guessing over expressing uncertainty, data distribution gaps, and prompting effects — that do not equally govern correct outputs. The generation channel is shared, but the upstream conditions that produce errors are separable and require distinct mitigation strategies.
“A technology executive used ChatGPT to help develop a personalized cancer vaccine for his dog, which had been diagnosed with cancer.”
The core claim is accurate: Sydney-based tech professional Paul Conyngham used ChatGPT — alongside other AI tools — to help plan and develop a personalized mRNA cancer vaccine for his dog Rosie after her cancer diagnosis. However, "technology executive" is a loose description (sources call him a tech entrepreneur, AI consultant, or data engineer), and ChatGPT's role was primarily as a research and planning assistant — human scientists at UNSW performed the actual genome sequencing, vaccine synthesis, and treatment.
“AI chatbots, such as ChatGPT, provide medical advice that is consistently reliable and safe for users.”
The claim that AI chatbots like ChatGPT provide "consistently reliable and safe" medical advice is not supported by the evidence. Multiple high-quality studies from 2024–2026 show ChatGPT gave incorrect advice in over 51% of medical emergencies, exhibited hallucination rates of 50–82%, and correctly identified conditions in fewer than 34.5% of real-world cases. ECRI designated AI chatbot misuse as the top health technology hazard for 2026. While chatbots show promise in narrow, controlled tasks, their performance is neither consistent nor safe for general medical advice.