Skip to content

Claim analyzed

Tech

“Generative models predict likely word sequences rather than retrieve facts, which can lead to factual inaccuracies.”

The conclusion

True
9/10

The claim accurately describes a source of factual errors in text-generating language models. NIST guidance and peer-reviewed research show that predicting likely next tokens can produce factual inaccuracies; generation does not inherently involve retrieving facts from an external source. Retrieval-augmented systems can add that capability, but the claim's wording correctly says inaccuracies can result.

Caveats

  • The evidence primarily addresses text-generating language models, not every type of generative model.
  • Retrieval-augmented systems can consult external sources; retrieval is not excluded from every generative system.
  • The two NIST links refer to the same document and should not be counted as independent sources.

Sources

Ranked by source quality and relevance

#1
tsapps.nist.gov 2024-07 | Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile

Confabulations are a natural result of the way generative models are designed: they generate outputs that approximate the statistical distribution of their training data; for example, LLMs predict the next token or word in a sentence or phrase. While such statistical prediction can produce factually accurate and consistent outputs, it can also produce outputs that are factually inaccurate or internally inconsistent.2.2. Confabulation “Confabulation” refers to a phenomenon in which GAI systems generate and confidently present erroneous or false content in response to prompts. Confabulations also include generated outputs that diverge from the prompts or other input or that contradict previously generated statements in the same context. These phenomena are colloquially also referred to as “hallucinations” or “fabrications.” Confabulations can occur across GAI outputs and contexts. 9,10 Confabulations are a natural result of the way generative models are designed: they generate outputs that approximate the statistical distribution of their training data; for example, LLMs predict the next token or word in a sentence or phrase. While such statistical prediction can produce factually accurate and consistent outputs, it can also produce outputs that are factually inaccurate or internally inconsistent. This dynamic is particularly relevant when it comes to open-ended prompts for long-form responses and in domains which require highly contextual and/or domain expertise. Risks from confabulations may arise when users believe false content – often due to the confident nature of the response – leading users to act upon or promote the false information. This poses a challenge for many real-world applications, such as in healthcare, where a confabulated summary of patient …

#2
nature.com 2026-04-22 | Evaluating large language models for accuracy incentivizes hallucinations

Here we show how next-word prediction and accuracy-based evaluations inadvertently reward unwarranted guessing. Initially, next-word pretraining creates statistical pressure towards hallucination even with idealized error-free data: using learning theory 8, 9, we show that facts lacking repeated support in training data (such as one-off details) yield unavoidable errors, whereas recurring regularities (such as grammar) do not. … Initially, LLMs are trained to optimize next-word (or next-token) prediction, during pretraining. Although falsehoods in training data can cause hallucinations, the phenomenon is not purely garbage-in garbage-out. We show that this objective creates statistical tendency towards hallucination even with ideal error-free training data.Abstract Large language models sometimes produce confident, plausible falsehoods (‘hallucinations’), limiting their reliability 1, 2. Previous work has offered numerous explanations and effective mitigations such as retrieval and tool use 3, consistency-based self-verification 4 and reinforcement learning from human feedback 5. Nonetheless, the problem persists even in state-of-the-art language models 6, 7. Here we show how next-word prediction and accuracy-based evaluations inadvertently reward unwarranted guessing. Initially, next-word pretraining creates statistical pressure towards hallucination even with idealized error-free data: using learning theory 8, 9, we show that facts lacking repeated support in training data (such as one-off details) yield unavoidable errors, whereas recurring regularities (such as grammar) do not. Subsequent training stages aim to correct such errors. However, dominant headline metrics such as accuracy systematically reward guessing over admitting uncertainty. To align incentives, we suggest two additions to the classic approach of adding error penalties to evaluations to control abstention 10, 11. … To make this lens precise, we first address a practical obstacle: defining and measuring hallucination is complicated by issues such as responses that contain multiple vague claims 13. We side-step these definitional tangles by analysing an abstract set of errors, following learning theory, which applies to binary classification (for example, dog versus cat images) without adjudicating edge cases (for example, images containing both). Initially, LLMs are trained to optimize next-word (or next-token) prediction, during pretraining. Although falsehoods in training data can cause hallucinations, the phenomenon is not purely garbage-in garbage-out. We show that this objective creates statistical tendency towards hallucination even with ideal error-free training data. The key insight is a formal reduction to binary classification: any language model implicitly answers ‘Is this response valid?’ for each candidate generation. This connection establishes lower bounds on hallucination rates and illuminates which error types to expect, as causes of misclassification are well understood 9. …

#3
papers.nips.cc 2022-11-28 | Factuality Enhanced Language Models for Open-Ended Text Generation

However, the generative LMs (e.g., GPT-3) are solely trained to model the statistical correlations between subword tokens [5], and have limited capability to generate factually accurate text as illustrated in Table 1. … Both error types can be viewed as wrong associations of entities that appear at different parts of the training corpus with similar context. Such behavior is unsurprising because these LMs are uniformly trained with the next subword prediction objective instead of a fact-related objective.1 Introduction Large-scale pre-trained language models (LMs) have demonstrated impressive natural language generation results [1–4]. However, the generative LMs (e.g., GPT-3) are solely trained to model the statistical correlations between subword tokens [5], and have limited capability to generate factually accurate text as illustrated in Table 1. As a result, there are increasing concerns about the nonfactual generations from large-scale pre-trained LMs [e.g., 6–8], which needs to be adequately addressed for their safe deployment in real-world applications, e.g., content creation [9] and dialogue [10]. … about a film called “The Best of Me”. However, the correct author’s name is “Nicholas Sparks”, not “Gayle Forman”. Note that Gayle Forman is also an American young adult fiction author who writes similar type of novels as Nicholas Sparks. • Fabricated Fact: Fabricating some random facts. For example, “Samuel Witwer’s father is a Lutheran minister.” Note that, the pretraining corpus contains non-factual or fictional information, which can also contribute to such fabricated facts. Both error types can be viewed as wrong associations of entities that appear at different parts of the training corpus with similar context. Such behavior is unsurprising because these LMs are uniformly trained with the next subword prediction objective instead of a fact-related objective. (a) Diversity vs. NEER (b) Repetition vs. NEER Figure 2: Comparison between nucleus sampling (blue line) and factual-nucleus sampling (orange line). The x-axis is named entity error NEER. The y-axes are diversity and repetition in (a) and (b) respectively. The lower the repetition, the better. It is evident that factual-nucleus sampling has better trade-offs between factuality and diversity/repetition. For a reference, the diversity score of randomly sampled 5000 Wikipedia documents is 0.767.

#4
link.springer.com 2026-06-01 | Retrieval-augmented generation for natural language processing: a survey

One of the most prominent challenges is the hallucination problem (Ji et al. 2023a; Dale et al. 2023; Ji et al. 2023b; Siino and Tinnirello 2024), which refers to the tendency of LLMs to generate responses that are coherent and fluent but factually incorrect. … By supplying the model with relevant, factually grounded information at inference time, RAG enables generation that is informed by verifiable sources rather than relying solely on parametric memory. This approach mitigates the hallucination problem by anchoring outputs in retrieved information. … Task characteristics. Language modeling refers to the capability of modeling the probability distribution over sequences of tokens. In practice, this capability is instantiated as the next-token prediction task: given such a sequence of tokens \(x_1, \ldots , x_n\) called Prefix, the language modeling task aims to model its probability via next-token prediction… 2022; Brown et al. 2020; Kaplan et al. 2020; Hoffmann et al. 2022). Although existing LLMs have become a foundational component in modern NLP systems, several challenges still hinder the development of LLMs. One of the most prominent challenges is the hallucination problem (Ji et al. 2023a; Dale et al. 2023; Ji et al. 2023b; Siino and Tinnirello 2024), which refers to the tendency of LLMs to generate responses that are coherent and fluent but factually incorrect. Furthermore, the knowledge update issue poses a significant obstacle. To update the knowledge stored in the LLMs’ internal memory (Meng et al. 2022a; Wang et al. 2025; Zhang et al. 2024), it is necessary to retrain/fine-tune LLMs with new data, which is a costly process. … 2024). Building a domain-specific LLM demands substantial effort in data curation and model adaptation. To address these challenges, recent works (Lewis et al. 2020; Borgeaud et al. 2022; Guu et al. 2020) have proposed augmenting LLMs with an external knowledge base (KB), known as Retrieval-Augmented Generation (RAG). By supplying the model with relevant, factually grounded information at inference time, RAG enables generation that is informed by verifiable sources rather than relying solely on parametric memory. This approach mitigates the hallucination problem by anchoring outputs in retrieved information. Furthermore, RAG addresses the issue of knowledge staleness by updating the external database allows the model to leverage current information without costly retraining. … ### 6.1 Language modeling Task characteristics. Language modeling refers to the capability of modeling the probability distribution over sequences of tokens. In practice, this capability is instantiated as the next-token prediction task: given such a sequence of tokens \(x_1, \ldots , x_n\) called Prefix, the language modeling task aims to model its probability via next-token prediction, $$\begin{aligned} p(x_1, \ldots , x_n)=p(x_1)\cdot \prod ^n_{i=2} p(x_i|x_1, \ldots , x_{i-1}), \end{aligned}$$

#5
proceedings.iclr.cc TOWARDS UNDERSTANDING FACTUAL KNOWLEDGE OF LARGE LANGUAGE MODELS

Unlike conventional Knowledge Bases (KBs) that explicitly store factual knowl edge, LLMs implicitly store facts in their parameters. Content generated by the LLMs can often exhibit inaccuracies or deviations from the truth, due to facts that can be incorrectly induced or become obsolete over time. … Despite advancements in LLMs, they still struggle with generating content that exhibits inaccuracies or deviations from the facts and making reasoning errors (Lin et al., 2022; Bubeck et al., 2023). These factual errors can be difficult to identify since LLMs implicitly memorize facts through their parameters rather than explicitly store factual knowledge as traditional Knowledge Bases. … current LLMs are primarily trained on unstructured text using next word prediction loss (Brown et al., 2020; Touvron et al., 2023a).ABSTRACT Large language models (LLMs) have recently driven striking performance im provements across a range of natural language processing tasks. The factual knowledge acquired during pretraining and instruction tuning can be useful in various downstream tasks, such as question answering, and language generation. Unlike conventional Knowledge Bases (KBs) that explicitly store factual knowl edge, LLMs implicitly store facts in their parameters. Content generated by the LLMs can often exhibit inaccuracies or deviations from the truth, due to facts that can be incorrectly induced or become obsolete over time. To this end, we aim to explore the extent and scope of factual knowledge within LLMs by designing the benchmark Pinocchio. Pinocchio contains 20K diverse factual questions that span different sources, timelines, domains, regions, and languages. … bases (Petroni et al., 2019; Jiang et al., 2020c). Factual knowledge in language models acquired during pretraining can benefit knowledge-intensive downstream tasks such as question answering and fact checking (Roberts et al., 2020; Yu et al., 2023a; Pan et al., 2023). Despite advancements in LLMs, they still struggle with generating content that exhibits inaccuracies or deviations from the facts and making reasoning errors (Lin et al., 2022; Bubeck et al., 2023). These factual errors can be difficult to identify since LLMs implicitly memorize facts through their parameters rather than explicitly store factual knowledge as traditional Knowledge Bases. Accessing and interpreting the computations and memories of these models can be challenging (Ribeiro et al., 2016; Belinkov & Glass, 2019), especially when APIs are the only means of interaction and many interpretation methods rely on weights and representations (Cao et al., 2021b). The presence of errors … | Real-World | Contain | factual statements spread | online | PolitiFact | 986 | 1,987 | 609 | 3,582 | Domain-Specific · Contain facts · from health and science · domains PubHealth, · SciFact · 1,156 · 715 · 737 · 2,608 Multi-Lingual · Contain · facts in different · languages XFact, · CHEF · 820 · 848 · 547 · 2,215 current LLMs are primarily trained on unstructured text using next word prediction loss (Brown et al., 2020; Touvron et al., 2023a). In order to process structured data, it is often converted into text strings using various methods, such as linearizing tables. This raises the question of whether LLMs are capable of effectively memorizing and reasoning over facts from structured sources, similar to their performance with unstructured text. …

#6
cambridge.org This peer-reviewed article has been accepted for publication but not yet copyedited or typeset, and so may be subject to change during the production process. The article is considered published and may be cited using its DOI. This is an Open Access article, distributed under the terms of the Creative Commons Attribution NonCommercial-NoDerivatives licence (http://creativecommons.org/licenses/by-nc-nd/4.0/), which permits non-commercial re-use, distribution, and reproduction in any medium, provided the original work is unaltered and is properly cited. The written permission of Cambridge University Press must be obtained for commercial re-use or in order to create a derivative work.

In AI systems, such falsehoods arise from inadequate source verification. This can be caused by the probabilistic nature of text generation, weak grounding in external evidence, failure to retrieve or use relevant sources, misleading prompt context, and optimisation for fluency or user satisfaction rather than truth. … The crucial difference is that human confabulation is shaped by memory, subjectivity and emotional salience, whereas AI confabulation is shaped by training data, statistical prediction, prompt context, system design and optimisation procedures.… Confabulated accounts may also show positive or wish-fulfilling bias, with narratives that appear to support self-enhancement or psychological coherence (10). These features make confabulation a closer analogy to AI-generated falsehoods than hallucination. Like human confabulations, false AI outputs may be coherent, plausible and assembled from fragments of available information. In AI systems, such falsehoods arise from inadequate source verification. This can be caused by the probabilistic nature of text generation, weak grounding in external evidence, failure to retrieve or use relevant sources, misleading prompt context, and optimisation for fluency or user satisfaction rather than truth. AI systems outputs may show biases that have a limited resemblance to positive bias in human confabulation. One example is sycophancy, where models tend to agree with, reassure or please users even when this compromises accuracy (11). The crucial difference is that human confabulation is shaped by memory, subjectivity and emotional salience, whereas AI confabulation is shaped by training data, statistical prediction, prompt context, system design and optimisation procedures. In both cases, however, the core problem is the production of a coherent narrative without adequate verification against external evidence.

However, the generative LMs (e.g., GPT-3) are solely trained to model the statistical correlations between subword tokens [5], and have limited capability to generate factually accurate text as illustrated in Table 1. … Augmenting LM with an information retrieval (IR) system is one possible solution to leverage textual facts, however, at the cost of additional complexity and resource overhead to the model [10, 26, 22, 27, 28]. Therefore, we explore an IR-free method that enhances the innate factuality of LMs by continued training on a factually rich plain-text corpus. … Both error types can be viewed as wrong associations of entities that appear at different parts of the training corpus with similar context. Such behavior is unsurprising because these LMs are uniformly trained with the next subword prediction objective instead of a fact-related objective.… the factual errors. We release our code and FACTUALITYPROMPTS benchmark at: https://github.com/nayeon7lee/FactualityPrompt. 1 Introduction Large-scale pre-trained language models (LMs) have demonstrated impressive natural language generation results [1–4]. However, the generative LMs (e.g., GPT-3) are solely trained to model the statistical correlations between subword tokens [5], and have limited capability to generate factually accurate text as illustrated in Table 1. As a result, there are increasing concerns about the nonfactual generations from large-scale pre-trained LMs [e.g., 6–8], which needs to be adequately addressed for their safe deployment in real-world applications, e.g., content creation [9] and dialogue [10]. … human annotations for high-quality construction. A method that can directly leverage plain text knowledge (e.g., Wikipedia, encyclopedia books, peer-reviewed publications) would be desirable for factuality enhancement as it can remove the human annotation bottleneck and easily scale up the amount of injected knowledge. Augmenting LM with an information retrieval (IR) system is one possible solution to leverage textual facts, however, at the cost of additional complexity and resource overhead to the model [10, 26, 22, 27, 28]. Therefore, we explore an IR-free method that enhances the innate factuality of LMs by continued training on a factually rich plain-text corpus. In this work, we focus on measuring and improving the factuality of large-scale pre-trained language models (LMs) for open-ended text generation. Specifically, we make the following contributions: 1. We build the benchmark and metrics to measure the factual accuracy of pre-trained LM for open-ended text generation. … similar type of novels as Nicholas Sparks. • Fabricated Fact: Fabricating some random facts. For example, “Samuel Witwer’s father is a Lutheran minister.” Note that, the pretraining corpus contains non-factual or fictional information, which can also contribute to such fabricated facts. Both error types can be viewed as wrong associations of entities that appear at different parts of the training corpus with similar context. Such behavior is unsurprising because these LMs are uniformly trained with the next subword prediction objective instead of a fact-related objective. (a) Diversity vs. NEER (b) Repetition vs. NEER Figure 2: Comparison between nucleus sampling (blue line) and factual-nucleus sampling (orange line). The x-axis is named entity error NEER. The y-axes are diversity and repetition in (a) and (b) respectively. The lower the repetition, the better. …

#8
ar5iv.labs.arxiv.org 2023-08-26 | [2308.15711] Optimizing Factual Accuracy in Text Generation through Dynamic Knowledge Selection

Despite this, these generative LMs primarily model the statistical relationships between subword tokens [6] and exhibit limited ability in generating factually correct text. Consequently, there is a growing concern regarding the production of nonfactual content (also called hallucination) by these LLMs [7, 8, 9, 10].1 Introduction Large language models (LLMs), such as ChatGPT, have shown remarkable capabilities in natural language generation tasks [1, 2, 3, 4, 5]. Despite this, these generative LMs primarily model the statistical relationships between subword tokens [6] and exhibit limited ability in generating factually correct text. Consequently, there is a growing concern regarding the production of nonfactual content (also called hallucination) by these LLMs [7, 8, 9, 10]. Addressing this issue is crucial for the safe use of such models into real-world applications. Many prior studies have aimed to improve the factuality of text generation [11], and a promising approach among these involves incorporating external knowledge into the text generation process [12, 13, 14, 15, 16, 17]. These techniques generally employ either prepared external knowledge or retrieve knowledge via an information retrieval (IR) system. …

#9
nvlpubs.nist.gov 2024-07-25 | Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile

Confabulations are a natural result of the way generative models are designed: they generate outputs that approximate the statistical distribution of their training data; for example, LLMs predict the next token or word in a sentence or phrase. While such statistical prediction can produce factually accurate and consistent outputs, it can also produce outputs that are factually inaccurate or internally inconsistent.2.2. Confabulation “Confabulation” refers to a phenomenon in which GAI systems generate and confidently present erroneous or false content in response to prompts. Confabulations also include generated outputs that diverge from the prompts or other input or that contradict previously generated statements in the same context. These phenomena are colloquially also referred to as “hallucinations” or “fabrications.” Confabulations can occur across GAI outputs and contexts. 9,10 Confabulations are a natural result of the way generative models are designed: they generate outputs that approximate the statistical distribution of their training data; for example, LLMs predict the next token or word in a sentence or phrase. While such statistical prediction can produce factually accurate and consistent outputs, it can also produce outputs that are factually inaccurate or internally inconsistent. This dynamic is particularly relevant when it comes to open-ended prompts for long-form responses and in domains which require highly contextual and/or domain expertise. Risks from confabulations may arise when users believe false content – often due to the confident nature of the response – leading users to act upon or promote the false information. This poses a challenge for many real-world applications, such as in healthcare, where a confabulated summary of patient …

#10
mdpi.com 2025-03-03 | Hallucination Mitigation for Retrieval-Augmented Large ...

These models learn statistical representations of linguistic patterns and word distributions through complex deep learning architectures and pre-training on vast datasets and adapt to domain-specific tasks through supervised fine-tuning. … However, due to their reliance on fixed parameters, LLMs often produce task-irrelevant outputs or generate factually inconsistent responses when faced with tasks beyond the scope of their training data. This limitation can be understood as the knowledge boundary of LLMs, which refers to the extent of their knowledge derived solely from the patterns and associations in the training data. This phenomenon, where the model generates responses that are inconsistent with facts or appear meaningless, is commonly referred to as hallucination or confabulation. … The capability boundaries of LLMs are mainly reflected in the deficiencies in precise calculation, data retrieval, and complex logic processing. These deficiencies are mainly due to their design goals based on statistical learning and language patterns, rather than being built for mathematical reasoning or symbolic operations.1. Introduction In recent years, large language models (LLMs) such as GPT-4 [1], LLaMA [2], and Gemini [3] have rapidly advanced, achieving significant progress in the field of natural language processing (NLP). These models learn statistical representations of linguistic patterns and word distributions through complex deep learning architectures and pre-training on vast datasets and adapt to domain-specific tasks through supervised fine-tuning. They are widely used for various tasks, including semantic analysis [4], text generation [5], and reasoning [6]. However, due to their reliance on fixed parameters, LLMs often produce task-irrelevant outputs or generate factually inconsistent responses when faced with tasks beyond the scope of their training data. This limitation can be understood as the knowledge boundary of LLMs, which refers to the extent of their knowledge derived solely from the patterns and associations in the training data. This phenomenon, where the model generates responses that are inconsistent with facts or appear meaningless, is commonly referred to as hallucination or confabulation. Hallucinations undermine the reliability and trustworthiness of LLMs, and in scenarios requiring precise information, they can lead to severe consequences. … - Capability Boundary The capability boundaries of LLMs are mainly reflected in the deficiencies in precise calculation, data retrieval, and complex logic processing. These deficiencies are mainly due to their design goals based on statistical learning and language patterns, rather than being built for mathematical reasoning or symbolic operations. LLMs lack built-in mathematical logic, and their training data mainly comes from unstructured natural language texts. The model does not specifically learn mathematical rules, resulting in reliance on language patterns rather than real mathematical operations when dealing with mathematical problems. …

#11
nature.com 2026-04-27 | Hyper-RAG: combating LLM hallucinations using hypergraph-driven retrieval-augmented generation

LLMs are adept at interpreting input data and generating responses based on their training data, often exhibiting high confidence in their outputs. However, this confidence does not inherently guarantee factual correctness, resulting in discrepancies commonly referred to as LLM hallucinations 15. … By leveraging these meticulously curated knowledge points, RAG frameworks empower LLMs to anchor their generative outputs in verified data, thereby mitigating the incidence of hallucinations and enhancing the factual integrity of the responses, as opposed to relying solely on the inherently compressed knowledge acquired during model training.… Despite these advancements, the integration of LLMs within the medical domain has been relatively cautious. This hesitancy primarily stems from concerns regarding the accuracy and reliability of the generated content, which can introduce uncertainty into clinical decision-making processes and potentially lead to adverse medical outcomes 12, 13, 14. LLMs are adept at interpreting input data and generating responses based on their training data, often exhibiting high confidence in their outputs. However, this confidence does not inherently guarantee factual correctness, resulting in discrepancies commonly referred to as LLM hallucinations 15. LLM hallucinations occur when the generated content diverges from established facts, colloquially termed as “bullshit.” For instance, in the diagnosis of neurological disorders, an LLM might incorrectly attribute symptoms to an unrelated condition, potentially misleading healthcare professionals 12, 13, 14. … In contrast, LightRAG introduces a dual-layered knowledge graph architecture, comprising both local and global structures, to effectively organize and index granular details alongside overarching concepts within the original knowledge base. The quintessential attribute of these classical RAG methodologies lies in their ability to structurally encode the knowledge embedded within raw textual data, facilitating rapid retrieval of relevant prior information in response to specific inquiries. By leveraging these meticulously curated knowledge points, RAG frameworks empower LLMs to anchor their generative outputs in verified data, thereby mitigating the incidence of hallucinations and enhancing the factual integrity of the responses, as opposed to relying solely on the inherently compressed knowledge acquired during model training. Structuring raw corpus data can significantly enhance the efficiency of information retrieval; however, existing graph-based approaches to information architecture often result in substantial data loss 25, 26. Specifically, traditional graphs are constrained to representing pairwise correlations between entities, as illustrated in Fig. 1. …

#12
arxiv.org 2024-09-09 | LLMs Will Always Hallucinate, and We Need to Live With This

At the core of large language models lies a deceptively simple principle: the prediction of linguistic patterns. [3, 4]. The fundamental operation of an LLM can be distilled to a single, powerful question: Given a sequence of words, what word is likely to come next? … Impressive as they are, the model has not learned to think and has no concept of truth; it has learned to mimic the products of thought with astonishing fidelity. … Hallucinations in large language models (LLMs) occur when the models generate content that is false, fabricated, or inconsistent with their training data. These happen when the model, in an attempt to produce coherent responses, fills in gaps with plausible-sounding but incorrect information.The Essence of Large Language Models At the core of large language models lies a deceptively simple principle: the prediction of linguistic patterns. [3, 4]. The fundamental operation of an LLM can be distilled to a single, powerful question: Given a sequence of words, what word is likely to come next? Assuming we refer to a transformer-like architecture [5], for a sequence of tokens $x=(x1,x2,…,xn)​where​xi∈Vx=(x_{1},x_{2},...,x_{n})\text{where}x_{i}\in V$, a language model computes [6]: … $P⁡(xn+1|x1,…,xn)=P⁡(x1,x2,…​xn,xn+1)P⁡(x1,x2,…​xn)P(x_{n+1}|x_{1},...,x_{n})=\frac{P(x_{1},x_{2},...x_{n},x_{n+1})}{P(x_{1},x_{2},...x_{n})}$ As we scale these models, we observe a fascinating phenomenon: the emergence of apparently intelligent behaviours. [8] Impressive as they are, the model has not learned to think and has no concept of truth; it has learned to mimic the products of thought with astonishing fidelity. How language model generations work … ### Hallucinations in LLMs: What They Are and How They Happen Hallucinations in large language models (LLMs) occur when the models generate content that is false, fabricated, or inconsistent with their training data. These happen when the model, in an attempt to produce coherent responses, fills in gaps with plausible-sounding but incorrect information. Hallucinations can range from subtle inaccuracies to completely fictional assertions, often presented with high confidence. It is important to note that LLM hallucinations can occur even with the best training, fine-tuning, or the use of techniques like Retrieval-Augmented Generation (RAG). Types of Hallucinations in LLMs

#13
pmc.ncbi.nlm.nih.gov 2025-10-08 | MEGA-RAG: a retrieval-augmented generation framework with ...

As illustrated in Figures 1, 2, conventional large language models (LLMs) generate answers based solely on their internal knowledge, which increases the risk of hallucinations and factual inconsistencies. Traditional retrieval-augmented generation (RAG) frameworks mitigate this issue by retrieving supporting documents; however, they typically rely on a single retrieval pass and lack mechanisms for evaluating answer consistency or handling conflicting evidence. … The inflated false-positive rate in the vanilla LLM arises primarily from its reliance on semantic matching without sufficient factual grounding. The model tends to label passages as relevant whenever strong lexical or contextual overlap is detected, even when they do not precisely align with the query's factual requirements.3.1 MEGA-RAG framework We propose MEGA-RAG, a novel retrieval-augmented question answering (QA) framework specifically tailored for the biomedical and public health domains, where accuracy, interpretability, and evidence grounding are critical. As illustrated in Figures 1, 2, conventional large language models (LLMs) generate answers based solely on their internal knowledge, which increases the risk of hallucinations and factual inconsistencies. Traditional retrieval-augmented generation (RAG) frameworks mitigate this issue by retrieving supporting documents; however, they typically rely on a single retrieval pass and lack mechanisms for evaluating answer consistency or handling conflicting evidence. In contrast, MEGA-RAG integrates a four-stage architecture designed to overcome these limitations. Figure 1. … Confusion matrix comparison of LLM, LLM+RAG, and MEGA-RAG. The inflated false-positive rate in the vanilla LLM arises primarily from its reliance on semantic matching without sufficient factual grounding. The model tends to label passages as relevant whenever strong lexical or contextual overlap is detected, even when they do not precisely align with the query's factual requirements. Standard RAG mitigates this by constraining candidate retrieval to external documents, but its dense retriever's limited filtering capacity can lead to overcorrection, excluding many borderline-relevant passages and increasing false negatives. In contrast, MEGA-RAG introduces a multi-stage evidence aggregation pipeline that combines dense retrieval with cross-encoder re-ranking and weighted entailment scoring. …

#14
aclanthology.org 2025-04-29 | Improve Decoding Factuality by Token-wise Cross Layer Entropy of Large Language Models

Despite their impressive capacities, Large lan guage models (LLMs) often struggle with the hallucination issue of generating inaccurate or fabricated content even when they pos sess correct knowledge. … Large language models typically consist of N stacked transformer layers, followed by an affine layer that maps the internal representations to a next-token probability distribution. … Decoding methods cannot inject additional knowledge into LLMs, they can only amplify the model’s inherent knowledge to im prove next-token predictions and reduce erroneous outputs.Abstract Despite their impressive capacities, Large lan guage models (LLMs) often struggle with the hallucination issue of generating inaccurate or fabricated content even when they pos sess correct knowledge. In this paper, we extend the exploration of the correlation be tween hidden-state prediction changes and out put factuality into a deeper, token-wise level. Based on the insights , we propose cross-layer Entropy eNhanced Decoding (END), a decod ing method that mitigates hallucinations with out requiring extra training. … The cross-layer entropy adjusts the final next-token prediction by suppressing the token A for low factuality and highlighting token D for both high prediction probability and high factuality. 4.1 Cross-Layer Entropy Large language models typically consist of N stacked transformer layers, followed by an affine layer that maps the internal representations to a next-token probability distribution. We denote the hidden state of the l-th layer as h (l) , the classifi cation head as φ(· ), and vt as generation token at step t over the vocabulary set V . The prediction probability from the l-th layer can be expressed as: … Limitation Hallucination Type Decoding methods cannot inject additional knowledge into LLMs, they can only amplify the model’s inherent knowledge to im prove next-token predictions and reduce erroneous outputs. Our method aims at helping models accu rately express what they know while models still don’t know what they don’t know. Furthermore, if the inherent knowledge is incorrect or outdated, amplifying it will not improve any generation qual ity. …

#15
ar5iv.labs.arxiv.org 2021-05-14 | [2105.06597] RetGen: A Joint framework for Retrieval and Grounded Text Generation Modeling

Recent advances in large-scale pre-training such as GPT-3 allow seemingly high quality text to be generated from a given prompt. However, such generation systems often suffer from problems of hallucinated facts, and are not inherently designed to incorporate useful external information. … These models, however, are not designed to leverage external information to enhance or to verify the predicted text. Gao et al. (2020), for example, demonstrate that they fail to reliably generate responses grounded in real-world knowledge, and may fall short when generating goal-directed responses that are optimized for information-seeking task completion. … Third, even in scenarios that call only for public information, generation from these LMs may be unfaithful to the facts (e.g., hallucinations about birth dates), especially when the people or entities are less well known and the scenario demands a high degree of fidelity.Abstract Recent advances in large-scale pre-training such as GPT-3 allow seemingly high quality text to be generated from a given prompt. However, such generation systems often suffer from problems of hallucinated facts, and are not inherently designed to incorporate useful external information. Grounded generation models appear to offer remedies, but their training typically relies on rarely-available parallel data where information-relevant documents are provided for context. … ## 1 Introduction Recent large-scale pre-trained language models (LMs) such as BERT (Devlin et al. 2019), GPT-2 (Radford et al. 2019), GPT-3 (Brown et al. 2020), and T5 (Raffel et al. 2019) have brought numerous breakthroughs in natural language generation (NLG) across a variety of tasks. These models, however, are not designed to leverage external information to enhance or to verify the predicted text. Gao et al. (2020), for example, demonstrate that they fail to reliably generate responses grounded in real-world knowledge, and may fall short when generating goal-directed responses that are optimized for information-seeking task completion. These models pose several challenges in information-demanding scenarios: First, they are usually trained offline, rendering the model agnostic to the latest information (e.g., asking a chat-bot trained from 2011-2018 about COVID-19). Second, they are mostly trained on public data, rendering them less suitable in scenarios where customized or personalized information must be processed (e.g., writing suggestions based on private user-data). Third, even in scenarios that call only for public information, generation from these LMs may be unfaithful to the facts (e.g., hallucinations about birth dates), especially when the people or entities are less well known and the scenario demands a high degree of fidelity. As a practical matter, moreover, there remains a fundamental capacity issue in that large LMs cannot effectively represent all the information about every person or entity in the world.

#16
research.google 2025-09-17 | Making LLMs more accurate by using all of their layers

LLMs break sentences into smaller units called "tokens”, which can be individual words, parts of words, or even punctuation marks. When an LLM generates text, it does so one token at a time. At each step, the LLM doesn't just pick the single most likely token. Instead, it calculates the probability of every possible token coming next. This set of probabilities is what’s known as a “distribution”. … LLMs process text through multiple layers, generating " logits" (prediction scores) at each layer, with the final layer's logits typically determining the output. "Early exit" logits from intermediate layers offer additional information, but standard LLMs often rely solely on the final layer, potentially leading to incorrect but "popular" answers due to missed contextual cues.How SLED works LLMs break sentences into smaller units called "tokens”, which can be individual words, parts of words, or even punctuation marks. When an LLM generates text, it does so one token at a time. At each step, the LLM doesn't just pick the single most likely token. Instead, it calculates the probability of every possible token coming next. This set of probabilities is what’s known as a “distribution”. LLMs process text through multiple layers, generating " logits" (prediction scores) at each layer, with the final layer's logits typically determining the output. "Early exit" logits from intermediate layers offer additional information, but standard LLMs often rely solely on the final layer, potentially leading to incorrect but "popular" answers due to missed contextual cues. SLED improves this by using information from all the layers of the LLM, not just the last one. It does this by reusing the final projection matrix in the Transformer architecture on early exit logits to create probability distributions over the same set of possible tokens that the final layer uses. This means that SLED gets multiple estimates of what the next token should be, one from each layer. …

#17
link.springer.com 2026-07-22 | Disentangling Faithfulness Hallucinations in Retrieval-Augmented Generation: A Systematic Benchmark and Analysis

Large Language Models (LLMs) have transformed generative AI and NLP, yet their propensity to produce fluent but factually incorrect outputs so-called hallucinations remains a key barrier to reliability and real-world deployment. … RAG blends the encyclopedic memory of a search engine with generative models and consists of two main modules: the retrieval phase and the generation phase. Unlike traditional models that depend solely on internal knowledge, RAG enhances the generation process by incorporating relevant external documents retrieved from knowledge sources such as databases, search engines, or vector stores.1 Introduction Large Language Models (LLMs) have transformed generative AI and NLP, yet their propensity to produce fluent but factually incorrect outputs so-called hallucinations remains a key barrier to reliability and real-world deployment. Hallucinations are typically categorized as intrinsic (also known as faithfulness hallucinations), when outputs deviate from the provided context, or extrinsic, when they contradict established facts. In this study, we particularly focus on intrinsic, or in other words faithfulness hallucinations. RAG has emerged as a promising framework to mitigate hallucinations by grounding LLM outputs in external, verifiable sources. RAG systems combine a retrieval module which fetches relevant documents from knowledge bases with a generation module that conditions its output on both the user query and the retrieved evidence. While RAG is designed to improve factual alignment, faithfulness hallucinations still occur, as shown in Figure 1, revealing limitations in how models interpret and reason over retrieved content (Islam et al., 2024). RAG blends the encyclopedic memory of a search engine with generative models and consists of two main modules: the retrieval phase and the generation phase. Unlike traditional models that depend solely on internal knowledge, RAG enhances the generation process by incorporating relevant external documents retrieved from knowledge sources such as databases, search engines, or vector stores. RAG systems typically adopt a retrieve-then-read architecture, wherein a retriever first identifies relevant content, and a generator subsequently produces a response conditioned on both the user query and the retrieved documents (Guu et al., 2020; Lewis et al., 2020; Shuster et al., 2021; Karpukhin et al., 2020). Fig. 1

#18
proceedings.iclr.cc SATISFIES: A

Large language models (LLMs) encode substantial knowledge (Petroni et al., 2019; Srivastava et al., 2022), yet they are prone to generating factually incorrect text. For instance, LLMs can generate confident-appearing completions with hallucinations (Zhang et al., 2023; Ji et al., 2023), fabricating entities or factual claims. … The LLM produces a predicted probability distribution for the next token Pˆ(tT +1| t1:T ) using a linear softmax layer on the last layer representation x L T. … Surprisingly, even though LLMs are optimized by maximizing the next token probability, probing attention patterns exclusively on the constraints can match or sometimes exceed this performance—without using hidden states or non-constraint tokens.… We propose SAT Probe, a method probing attention patterns, that can predict factual errors and fine-grained constraint satisfaction, and allow early error identification. The approach and findings take another step towards using the mechanistic understanding of LLMs to enhance their reliability. 1 1 INTRODUCTION Large language models (LLMs) encode substantial knowledge (Petroni et al., 2019; Srivastava et al., 2022), yet they are prone to generating factually incorrect text. For instance, LLMs can generate confident-appearing completions with hallucinations (Zhang et al., 2023; Ji et al., 2023), fabricating entities or factual claims. As LLMs reach wider audiences and are used for safety-critical applications, understanding the factuality of generations rises to paramount importance. However, our understanding of how LLMs process factual queries and produce errors is nascent. … Often, each hidden state vector has the same number of dimensions, i.e., ∀ i, ℓ x ℓ i ∈ R d. The states are obtained by x ℓ i = x ℓ− 1 i + a ℓ i + m ℓ i, (1) where we call m ℓ i the MLP contribution and a ℓ i the attention contribution to a token i at layer ℓ, respectively. The LLM produces a predicted probability distribution for the next token Pˆ(tT +1| t1:T ) using a linear softmax layer on the last layer representation x L T. In this work, we study the interactions among tokens. Since the MLP layers (see Appendix A) in standard Transformers do not capture token interactions, we focus primarily on the attention operation. The attention operation updates each token’s state using the previous states at all positions, i.e., … This further provides insight into why the COMBINED predictor performs better overall. Attention alone is significantly better than the CONSTANT baseline which suggests that the signals relay a nontrivial amount of information, sometimes exceeding CONFIDENCE. Surprisingly, even though LLMs are optimized by maximizing the next token probability, probing attention patterns exclusively on the constraints can match or sometimes exceed this performance—without using hidden states or non-constraint tokens. However, attention alone does not explain all failures (we observe some attention on constraints where the model still fails), there is an opportunity for further investigation. Our findings demonstrate the value of studying the procedure by which a model produces an output, rather than only the output itself. 5.2 EXTENSIONS

#19
aclanthology.org 2024-11-12 | Analysis of Plan-based Retrieval for Grounded Text Generation

In text generation, hallucinations refer to the generation of seemingly coherent text that contradicts established knowledge. One com pelling hypothesis is that hallucinations occur when a language model is given a generation task outside its parametric knowledge (due to rarity, recency, domain, etc.). … Among the errors made by these models, producing generations with factual and/or grounding errors, often referred to as hallucina tions, limit the broader applicability and capability of language models (Gao et al., 2023a; Manakul et al., 2023; Min et al., 2023; Ji et al., 2023a; Peng et al., 2023). Hallucinations differ from other kinds of errors in that the generated text is syntactically correct and semantically plausible.Abstract In text generation, hallucinations refer to the generation of seemingly coherent text that contradicts established knowledge. One com pelling hypothesis is that hallucinations occur when a language model is given a generation task outside its parametric knowledge (due to rarity, recency, domain, etc.). A common strat egy to address this limitation is to infuse the language models with retrieval mechanisms, providing the model with relevant knowledge for the task. In this paper, we leverage the plan ning capabilities of instruction-tuned LLMs and analyze how planning can be used to guide retrieval to further reduce the frequency of hallucinations. We empirically evaluate sev eral variations of our proposed approach on long-form text generation tasks. … ### 1 Introduction Large, parametric language models (LLMs) pro vide highly fluent text for many applications such as summarization, dialogue, and translation (De vlin et al., 2019; Brown et al., 2020; Thoppilan et al., 2022; Chowdhery et al., 2024; Anil et al., 2023, inter alia). Among the errors made by these models, producing generations with factual and/or grounding errors, often referred to as hallucina tions, limit the broader applicability and capability of language models (Gao et al., 2023a; Manakul et al., 2023; Min et al., 2023; Ji et al., 2023a; Peng et al., 2023). Hallucinations differ from other kinds of errors in that the generated text is syntactically correct and semantically plausible. These halluci nations are generations that, were they factually accurate, would be satisfactory model output. †Work done as a Student Researcher at Google. ∗Equal contribution.

#20
dl.acm.org 2026-09-01 | Revealing and Mitigating the Impact of LLM-Generated Content on Retrieval-Augmented Generation Systems

Although LLMs encode rich internal knowledge, the potential inherent social biases [18, 27, 55] and the abuse of generative capabilities, such as AI-driven rewriting, may introduce distorted facts or even create fake news [15, 33, 37], which can significantly impact the performance of RAG systems. … This can be due to that LLM tends to sample higher-probability tokens when generating distorted facts, resulting a lower PPL.… The framework assesses the performance of the RAG system in scenarios where aligned and distorted LLM-generated content are mixed with human-written content, respectively. revealed a phenomenon known as “source bias”, wherein neural retrieval models tend to favor LLM-generated content over semantically related human-written content. This bias poses a direct threat to the quality of retrieved knowledge that RAG systems rely on. Although LLMs encode rich internal knowledge, the potential inherent social biases [18, 27, 55] and the abuse of generative capabilities, such as AI-driven rewriting, may introduce distorted facts or even create fake news [15, 33, 37], which can significantly impact the performance of RAG systems. Therefore, evaluating and mitigating the impact of LLM-generated content is crucial for maintaining the reliability of such systems. Existing work involves detecting [44] and filtering out all LLM-generated content or training retrieval models to reduce the inherent bias [10]. However, these approaches risk overcorrection, causing significant information loss. Moreover, existing research [9, 10] overlooks the distinction between LLM-generated content with different factual … Lower PPL indicates that the LM better understands and ”trusts” the patterns in the text. We compute PPL1for the human-written corpus and two types of the LLM-generated corpus of the NQ dataset, with results shown in Figure 3. LLM-generated content generally has lower PPL than human-written content, and distorted LLM-generated content exhibits lower PPL than aligned LLM-generated content. This demonstrates that neural retrievers better understand and model distorted LLM-generated content. This can be due to that LLM tends to sample higher-probability tokens when generating distorted facts, resulting a lower PPL. 1Since most neural retrievers are based on BERT, we follow [43] to compute PPL using BERT.

#21
papers.nips.cc I Don’t Know: Explicit Modeling of Uncertainty with an Token

Large Language Models are known to capture real-world knowledge, allowing them to excel in many downstream tasks. Despite recent advances, these models are still prone to what are commonly known as hallucinations, causing them to emit unwanted and factually incorrect text. … Despite the popularity of LLMs, they are prone to what is commonly referred to as hallucinations, which severely hinder their performance and reliability [Ji et al., 2023, Manduchi et al., 2024]. Examples of hallucinations include factually incorrect [Maynez et al., 2020, Devaraj et al., 2022, Tam et al., 2023], inconsistent [Elazar et al., 2021, Mündler et al., 2023], self-contradicting [Cohen et al., 2024] or non-attributable text [Bohnet et al., 2022, Rashkin et al., 2023, Yue et al., 2023].Abstract Large Language Models are known to capture real-world knowledge, allowing them to excel in many downstream tasks. Despite recent advances, these models are still prone to what are commonly known as hallucinations, causing them to emit unwanted and factually incorrect text. In this work, we propose a novel calibration method that can be used to combat hallucinations. We add a special [IDK] (“I don’t know”) token to the model’s vocabulary and introduce an objective function that shifts probability mass to the [IDK] token for incorrect predictions. This approach allows the model to express uncertainty in its output explicitly. We evaluate our proposed method across multiple model architectures and factual downstream tasks. … amount of the information seen during pre-training, allowing them to encode real-world knowledge in their parameters and act as knowledge bases [Petroni et al., 2019, Roberts et al., 2020, Cohen et al., 2023a, Pan et al., 2023]. Owing to this phenomenon, LLMs can be used in multiple settings requiring this real-world knowledge, such as closed-book question answering [Brown et al., 2020, Roberts et al., 2020] and information retrieval [Tay et al., 2022]. Despite the popularity of LLMs, they are prone to what is commonly referred to as hallucinations, which severely hinder their performance and reliability [Ji et al., 2023, Manduchi et al., 2024]. Examples of hallucinations include factually incorrect [Maynez et al., 2020, Devaraj et al., 2022, Tam et al., 2023], inconsistent [Elazar et al., 2021, Mündler et al., 2023], self-contradicting [Cohen et al., 2024] or non-attributable text [Bohnet et al., 2022, Rashkin et al., 2023, Yue et al., 2023]. A prominent method employed to combat such hallucinations is model calibration [Guo et al., 2017a, Brundage et al., 2020], which aims to calibrate the confidence of model predictions such that they 1We release our code and IDK-tuned model checkpoints at https://github.com/roi-hpi/ IDK-token-tuning. 38th Conference on Neural Information Processing Systems (NeurIPS 2024).

#22
arxiv.org Retrieval is Accurate Generation

Standard language models generate text by selecting tokens from a fixed, finite, and standalone vocabulary. We introduce a novel method that selects context-aware phrases from a collection of supporting documents. … Our research aims to enhance the interpretability and factuality of language models (LMs) by transitioning from token generation to phrase retrieval. … However, it’s important to note that while our model achieves a high MAUVE score based solely on token prediction, the factuality of the generated text is lower than when phrase retrieval is integrated.Abstract Standard language models generate text by selecting tokens from a fixed, finite, and standalone vocabulary. We introduce a novel method that selects context-aware phrases from a collection of supporting documents. One of the most significant challenges for this paradigm shift is determining the training oracles, because a string of text can be segmented in various ways and each segment can be retrieved from numerous possible documents. … ### 3.1 Overview Our research aims to enhance the interpretability and factuality of language models (LMs) by transitioning from token generation to phrase retrieval. First, the semantics of phrases are enhanced by their surrounding contexts (Mikolov et al., 2013), leading to a more discriminative representation for inference. Second, each retrieved phrase can be traced back to its original document, enhancing the accountability of the output. … Beyond this point, the model’s performance peaks. Furthermore, even without phrases (i.e., the token rate is 1), our model can generate high-quality text, suggesting that our method also enhances token prediction learning. However, it’s important to note that while our model achieves a high MAUVE score based solely on token prediction, the factuality of the generated text is lower than when phrase retrieval is integrated. This highlights the need for more innovative metrics to precisely measure the quality of generated text. The impact of self-reinforcement on knowledge-intensive tasks.

#23
papers.nips.cc Long-form factuality in large language models

Large language models (LLMs) often generate content that contains factual errors when responding to fact-seeking prompts on open-ended topics. … In particular, they often produce factual errors in which a claim contradicts established ground-truth knowledge (Huang et al., 2023; Zhang et al., 2023, inter-alia). 2 For example, models may respond with incorrect information about established facts such as dates, statistics, or even a celebrity’s occupation (Li et al., 2023; Min et al., 2023; Muhlgay et al., 2023).Abstract Large language models (LLMs) often generate content that contains factual errors when responding to fact-seeking prompts on open-ended topics. To benchmark a model’s long-form factuality in open domains, we first use GPT-4 to generate LongFact, a prompt set comprising thousands of questions spanning 38 topics. We then propose that LLM agents can be used as automated evaluators for long form factuality through a method which we call Search-Augmented Factuality Evaluator (SAFE). … ##### 1 Introduction Large Language Models (LLMs) have significantly improved in recent years (Brown et al., 2020; Chowdhery et al., 2022; Google, 2023; OpenAI, 2023; Gemini Team, 2023, inter alia) but still lack reliability in responding to in-depth factuality questions. In particular, they often produce factual errors in which a claim contradicts established ground-truth knowledge (Huang et al., 2023; Zhang et al., 2023, inter-alia). 2 For example, models may respond with incorrect information about established facts such as dates, statistics, or even a celebrity’s occupation (Li et al., 2023; Min et al., 2023; Muhlgay et al., 2023). These factual errors weaken a language model’s factuality, making the model unreliable in many real-world settings where a factually-accurate response is expected. ∗Lead contributors. Author contributions listed at the end of the paper. 2We focus on factuality and factual errors, not hallucination, as our proposed evaluation method focuses on determining whether a response is factual with respect to external established knowledge (factuality) rather than …

#24

Our analysis applies to general density estimation and not only “next-word predictors” even though many language models are trained using self-supervised learning to predict each word based on the previous words. … A number of studies have shown how language models augmented with search or Retrieval-Augmented Generation (RAG) reduce hallucinations (Lewis et al., 2020; Shuster et al., 2021; Nakano et al., 2021; Zhang and Zhang, 2025).Not merely autocomplete. Our analysis applies to general density estimation and not only “next-word predictors” even though many language models are trained using self-supervised learning to predict each word based on the previous words. It is tempting to attribute hallucinations to poorly chosen prefixes (e.g., ‘‘Adam Kalai was born on’’) for which the language model cannot provide valid completions. However, from a purely statistical perspective, ignoring computation, the autocomplete view 5 5 5 Mathematically, any distribution $pp$ induces a distribution of completions $p⁡(wi​wi+1​…∣w1​w2​…​wi−1)p(w_{i}w_{i+1}\ldots\mid w_{1}w_{2}\ldots w_{i-1})$ for every prefix of words $w1​…​wi−1w_{1}\ldots w_{i-1}$ in its support. … ##### Search (and reasoning) are not panaceas. A number of studies have shown how language models augmented with search or Retrieval-Augmented Generation (RAG) reduce hallucinations (Lewis et al., 2020; Shuster et al., 2021; Nakano et al., 2021; Zhang and Zhang, 2025). However, Observation 1 holds for arbitrary language models, including those with RAG. In particular, the binary grading system itself still rewards guessing whenever search fails to yield a confident answer. Moreover, search may not help with miscalculations such as in the letter-counting example, or other intrinsic hallucinations. Latent context.

#25
arxiv.org Retrieval-Augmented Generation for AI-Generated Content: A Survey

Information retrieval is another pivotal application within the field of computer science. Different from generation, retrieval aims to locate relevant existing objects from a vast pool of resources. … Despite significant advancements in generative models, AIGC still grapples with challenges like outdated knowledge, lack of long-tail knowledge [27], and risks of leaking private training data [28]. Retrieval-Augmented Generation (RAG) aims to mitigate these issues with its flexible data repository [29]. … In RAG systems, the quality of retrieved content determines the information fed into the generators. Lower content quality increases the risk of model hallucinations or other degradation.… These advancements are further bolstered by the availability of rich, high-quality datasets [1, 18], which provide ample training samples to fully optimize model parameters. Information retrieval is another pivotal application within the field of computer science. Different from generation, retrieval aims to locate relevant existing objects from a vast pool of resources. The most prevalent application of retrieval lies in web search engines, which primarily focus on the task of document retrieval [19, 20]. In the present era, efficient information retrieval systems can handle document collections on the order of billions [21, 22]. Besides documents, retrieval has also been applied for many other modalities [23, 24, 25, 26]. Despite significant advancements in generative models, AIGC still grapples with challenges like outdated knowledge, lack of long-tail knowledge [27], and risks of leaking private training data [28]. Retrieval-Augmented Generation (RAG) aims to mitigate these issues with its flexible data repository [29]. The retrievable knowledge acts as non-parametric memory, which is easily updatable, accommodates extensive long-tail knowledge, and can encode confidential data. Moreover, retrieval can lower generation costs. RAG can reduce the size of large models [30], support long contexts [31], and eliminate certain generation steps [32]. … #### III-B2 Retriever Enhancement In RAG systems, the quality of retrieved content determines the information fed into the generators. Lower content quality increases the risk of model hallucinations or other degradation. In this section, we introduce efficient ways to enhance retrieval effectiveness. Recursive Retrieval: Recursive retrieval is to perform multiple searches to retrieve richer and higher-quality contents.

#26
arxiv.org 2025-03 | From Matching to Generation: A Survey on Generative Information Retrieval

However, the responses generated by language models may not always be reliable. They have the potential to generate irrelevant answers [85], contradict factual information [90, 104], provide outdated data [291], or generate toxic content [93, 263].… As illustrated in Figure 1, the generative approach marks a significant shift from traditional IR systems, which return a ranked list of documents (as shown in Figure 1(a,b)). Instead, response generation methods (depicted in Figure 1(c)) offer a more dynamic form of information access by directly generating detailed, user-centric responses, thereby providing a richer and more immediate understanding of the information need behind the users’ queries. However, the responses generated by language models may not always be reliable. They have the potential to generate irrelevant answers [85], contradict factual information [90, 104], provide outdated data [291], or generate toxic content [93, 263]. Consequently, these limitations render them unsuitable for many scenarios that require accurate and up-to-date information. To address these challenges, the academic community has developed strategies across four key aspects: enhancing internal knowledge [16, 37, 56, 119, 132, 193, 243, 267, 285]; augmenting external knowledge [5, 113, 139, 151, 204, 245, 333]; generating responses with citation [129, 142, 156, 204, 314]; and improving personal information assistance [149, 172, 295, 327]. …

#27
ar5iv.labs.arxiv.org 2022-09-21 | [2209.10063] Generate rather than Retrieve: Large Langu-age Models are Strong Context Generators

Besides, generated documents might suffer from hallucination error, resulting in incorrect predictions. … Table 15: Case studies of hallucination errors in InstructGPT generated documents. The documents contain contents that contradict to the facts and world knowledge, resulting in wrong predictions.… A major feature of retrieve-then-read is the ability to swap in new documents when new information is learned, such as temporally more recent documents, or adding in documents from a new domain to quickly adapt to a new downstream task. Our approach relies on a large language model to contain all this knowledge and adding new knowledge would likely require some retraining. Future work will explore how to efficiently incorporate new knowledge into our generate-then-read method. Besides, generated documents might suffer from hallucination error, resulting in incorrect predictions. We demonstrated case study in Table 15. Consideration in combination with recent approaches (Creswell & Shanahan, 2022) to boost generative faithfulness is a also direction worthy of future research. Ethics Statement … | and Darrel, as well as Johnny. | GPT generated documents: The Outsiders is a novel by S.E. Hinton. It is about a gang of greasers in Oklahoma in the 1960s. The greasers are from the poor side of town and constantly in trouble with the law. Q: Where are unipolar neurons found in spinal cord? · the granule region · dorsal root ganglia Explanation: The labled answer is “the distal dorsal root”, but the output “dorsal root ganglia” is the same. Table 15: Case studies of hallucination errors in InstructGPT generated documents. The documents contain contents that contradict to the facts and world knowledge, resulting in wrong predictions. Question: Who wrote the first declaration of human rights? Answer: Cyrus Cylinder Generated document: The first declaration of human rights was the Virginia Declaration of Rights, which was written by George Mason in 1776. | Retrieved document: John Peters Humphrey, OC (April 30, 1905 - 2013 March 14, 1995) was a Canadian legal scholar, jurist, and human rights advocate. He is most famous as the author of the first draft of the Universal Declaration of Human Rights. … … …

#28
arxiv.org Language Models Are Better Than Humans at Next-token ...

However, causal language models are not trained to perform well at these tasks; they are trained to accurately predict the next token given previous tokens in tokenized text. … LMs are not explicitly trained to perform well at natural language tasks. Their loss function is simply next-token prediction: accurately predicting the next token given previous tokens in tokenized text.Abstract Current language models are considered to have sub-human capabilities at natural language tasks like question-answering or writing code. However, causal language models are not trained to perform well at these tasks; they are trained to accurately predict the next token given previous tokens in tokenized text. It is not clear whether language models are better or worse than humans at next-token prediction. To try to answer this question, we performed two distinct experiments to directly compare humans and language models on this front: one measuring top-1 accuracy and the other measuring perplexity on OpenWebText. … ## 1 Introduction Recent language models (LMs) have demonstrated impressive capabilities in natural language tasks, like writing convincing human-like text, coding, or answering general knowledge questions. However, LMs are not considered to have yet surpassed human performance at these tasks. But performance at such tasks is not a fair way of comparing LMs and humans. LMs are not explicitly trained to perform well at natural language tasks. Their loss function is simply next-token prediction: accurately predicting the next token given previous tokens in tokenized text. How good are modern LMs compared to humans at next-token prediction? While one can construct tasks in which humans make next-token predictions better than any language model, there have been no “apples-to-apples” comparisons on non-handcrafted datasets. To answer this question, we performed two experiments that directly compare humans to language models on next-token prediction, using the OpenWebText dataset (Gokaslan & Cohen, 2019).

#29
arxiv.org Reducing hallucination in structured outputs via Retrieval ...

Retrieval-Augmented Generation is a common approach to limit generation of false or outdated information in classical NLP tasks such as question answering and summarization (Lewis et al., 2020; Izacard and Grave, 2021; Shuster et al., 2021). In the GenAI era, it refers to a process where relevant information from specific data sources is retrieved prior to generating text; the generation is then based on this retrieved information Gao et al. (2024). … Instead of retrieving facts, we retrieve JSON objects that could be part of the JSON output document.2 Related Work Retrieval-Augmented Generation is a common approach to limit generation of false or outdated information in classical NLP tasks such as question answering and summarization (Lewis et al., 2020; Izacard and Grave, 2021; Shuster et al., 2021). In the GenAI era, it refers to a process where relevant information from specific data sources is retrieved prior to generating text; the generation is then based on this retrieved information Gao et al. (2024). Our work differs from standard RAG as we apply it to a structured output task. Instead of retrieving facts, we retrieve JSON objects that could be part of the JSON output document. Providing plausible JSON objects to the LLM before generation increases the likelihood that the output JSON properties exist and that the generated JSON can be executed. A crucial ingredient of RAG is the retriever since its output will be part of the LLM input. …

#30
ar5iv.labs.arxiv.org 2025 | [2402.10612] Retrieve Only When It Needs: Adaptive Retrieval Augmentation for Hallucination Mitigation in Large Language Models

Despite their successes, it has been widely observed that even state-of-the-art LLMs often generate factually incorrect or nonsensical outputs, referred to as hallucinations (Ji et al., 2023a; Zhang et al., 2023b, d). These unreliable outputs pose significant risks in practical deployments of LLMs. … Internal hallucinations refer to the generation of wrong responses using only the parameterized knowledge of LLMs, while external hallucinations refer to the generation of incorrect responses using noisy documents introduced after retrieval.1. Introduction In recent years, large language models (LLMs) have demonstrated impressive abilities in natural language understanding (Hendrycks et al., 2021; Huang et al., 2023a), generation (Touvron et al., 2023; Taori et al., 2023), and reasoning (Zhang et al., 2023e; Wang et al., 2023a; Chu et al., 2023). Despite their successes, it has been widely observed that even state-of-the-art LLMs often generate factually incorrect or nonsensical outputs, referred to as hallucinations (Ji et al., 2023a; Zhang et al., 2023b, d). These unreliable outputs pose significant risks in practical deployments of LLMs. Figure 1. The limited knowledge of LLMs poses a challenge for generating accurate answers, referred to as Internal Hallucination, when faced with the latest or domain-specific questions. Additionally, retrieval-augmented generation occasionally faces the risk of error accumulation, where irrelevant evidence may infiltrate the generation phase and lead to nonfactual responses, known as External Hallucination. … ### 5.5. Quantitative Analysis (RQ5) To verify whether Rowen effectively reduces internal and external hallucinations, we present a comparative analysis of the prevalence of both types of hallucinations across three methods—Factool, Detect-and-Mitigate, and Rowen-Hybrid. Internal hallucinations refer to the generation of wrong responses using only the parameterized knowledge of LLMs, while external hallucinations refer to the generation of incorrect responses using noisy documents introduced after retrieval. Figure 5 shows that Factool and Detect-and-Mitigate are significantly prone to both types of hallucinations. Both baseline methods have high levels of internal and external hallucinations, struggling with internal coherence and external fact alignment. In contrast, Rowen-Hybrid effectively reduces external hallucinations by timely integrating external knowledge, thereby avoiding unnecessary information retrieval and mitigating potential errors. Table 9. …

#31
deepmind.google 2021-12-08 | Improving language models by retrieving from trillions of tokens — Google DeepMind

We explore an alternate path for improving language models: we augment transformers with retrieval over a database of text passages including web pages, books, news and code. … The RETRO architecture interleaves regular self-attention at a document level and cross-attention with retrieved neighbors at a finer passage level. This results in both more accurate and more factual continuations. … Below, we show two samples from our 7B baseline model and from our 7.5B RETRO model model that highlight how RETRO’s samples are more factual and stay more on topic than the baseline sample.… This has led to a tremendous increase in training energy cost and resulted in a generation of dense “Large Language Models” (LLMs) with 100+ billion parameters. Simultaneously, large datasets containing trillions of words have been collected to facilitate the training of these LLMs. We explore an alternate path for improving language models: we augment transformers with retrieval over a database of text passages including web pages, books, news and code. We call our method RETRO, for “Retrieval Enhanced TRansfOrmers”. Figure 1: A high-level overview of Retrieval Enhanced TransfOrmers (RETRO). … For each text passage (approximately a paragraph of a document), a nearest-neighbor search is performed which returns similar sequences found in the training database, and their continuation. These sequences help predict the continuation of the input text. The RETRO architecture interleaves regular self-attention at a document level and cross-attention with retrieved neighbors at a finer passage level. This results in both more accurate and more factual continuations. Furthermore, RETRO increases the interpretability of model predictions, and provides a route for direct interventions through the retrieval database to improve the safety of text continuation. In our experiments on the Pile, a standard language modeling benchmark, a 7.5 billion parameter RETRO model outperforms the 175 billion parameter Jurassic-1 on 10 out of 16 datasets and outperforms the 280B Gopher on 9 out of 16 datasets. Below, we show two samples from our 7B baseline model and from our 7.5B RETRO model model that highlight how RETRO’s samples are more factual and stay more on topic than the baseline sample. Figure 3: The baseline only generates 2 correct digits. With RETRO, the correct digits are generated after being retrieved by the database.

#32
research.facebook.com Retrieval Augmentation Reduces Hallucination in Conversation

Despite showing increasingly human-like conversational abilities, state-of-the-art dialogue models often suffer from factual incorrectness and hallucination of knowledge.Abstract Despite showing increasingly human-like conversational abilities, state-of-the-art dialogue models often suffer from factual incorrectness and hallucination of knowledge. In this work we explore the use of neural-retrieval-in-the-loop architectures – recently shown to be effective in open-domain QA – for knowledge-grounded dialogue, a task that is arguably more challenging as it requires querying based on complex multi-turn dialogue context and generating conversationally coherent responses. …

#33
aclanthology.org Faithfulness-Aware Uncertainty Quantification for Fact-Checking the Output of Retrieval-Augmented Generation

However, RAG remains prone to hallucinations: factually incorrect outputs may arise from inaccuracies in the model’s internal knowledge and the retrieved context.Abstract Large Language Models (LLMs) enhanced with knowledge retrieval, an approach known as Retrieval-Augmented Generation (RAG), have achieved strong performance in open-domain question answering. However, RAG remains prone to hallucinations: factually incorrect outputs may arise from inaccuracies in the model’s internal knowledge and the retrieved context. Existing approaches to mitigating hallucinations often conflate factuality with faithfulness to the retrieved evidence, incorrectly labeling factually correct statements as hallucinations if they are not explicitly supported by the retrieval. In this paper, we introduce FRANQ (Faithfulness-aware Retrieval-Augmented UNcertainty Quantification), a new method for hallucination detection in RAG outputs. …

#34
arxiv.org 2024-01-22 | [2401.11817] Hallucination is Inevitable: An Innate Limitation of Large Language Models

Specifically, we define a formal world where hallucination is defined as inconsistencies between a computable LLM and a computable ground truth function. By employing results from learning theory, we show that LLMs cannot learn all the computable functions and will therefore inevitably hallucinate if used as general problem solvers.Authors: Ziwei Xu, Sanjay Jain, Mohan Kankanhalli > Abstract: Hallucination has been widely recognized to be a significant drawback for large language models (LLMs). There have been many works that attempt to reduce the extent of hallucination. These efforts have mostly been empirical so far, which cannot answer the fundamental question whether it can be completely eliminated. In this paper, we formalize the problem and show that it is impossible to eliminate hallucination in LLMs. Specifically, we define a formal world where hallucination is defined as inconsistencies between a computable LLM and a computable ground truth function. By employing results from learning theory, we show that LLMs cannot learn all the computable functions and will therefore inevitably hallucinate if used as general problem solvers. Since the formal world is a part of the real world which is much more complicated, hallucinations are also inevitable for real world LLMs. Furthermore, for real world LLMs constrained by provable time complexity, we describe the hallucination-prone tasks and empirically validate our claims. Finally, using the formal world framework, we discuss the possible mechanisms and efficacies of existing hallucination mitigators as well as the practical implications on the safe deployment of LLMs.

Debate

Two AI advocates debated this claim using the research gathered.

Argument for

P
Proponent Argues TRUE

The claim is directly affirmed by NIST's authoritative Generative AI Profile (Sources 1 and 9), which states that confabulations are "a natural result of the way generative models are designed" because LLMs "predict the next token or word" by approximating the statistical distribution of training data, which can yield factually inaccurate outputs. Peer-reviewed work corroborates both halves of the claim: Nature (Source 2) shows next-word pretraining creates statistical pressure toward hallucination even with error-free data, NeurIPS (Source 3) notes LMs are "solely trained to model the statistical correlations between subword tokens" rather than a fact-related objective, and ICLR (Source 5) stresses that LLMs memorize facts implicitly in parameters rather than storing them as an explicit knowledge base, which is precisely why retrieval-augmented generation (Sources 4 and 11) was developed to anchor outputs in retrieved facts.

O
Opponent Rebuttal

The Proponent commits a scope fallacy by treating evidence about standalone, next-token-trained LLMs as evidence about all generative models, while retrieval-augmented systems explicitly retrieve external documents at inference and generate conditioned on that evidence (Source 4, "Retrieval-augmented generation for natural language processing: a survey"; Source 17, "Disentangling Faithfulness Hallucinations in Retrieval-Augmented Generation"). Moreover, NIST states that statistical prediction can produce factually accurate outputs as well as inaccurate ones (Source 1, NIST AI RMF Generative AI Profile), so its discussion of confabulation cannot establish the claim's categorical assertion that generative models predict sequences rather than retrieve facts.

Argument against

O
Opponent Argues FALSE

The claim is overbroad: retrieval-augmented generative models explicitly retrieve relevant documents or knowledge at inference time and generate conditioned on that evidence, rather than merely predicting word sequences (Source 4, “Retrieval-augmented generation for natural language processing: a survey”; Source 17, “Disentangling Faithfulness Hallucinations in Retrieval-Augmented Generation”). It also falsely implies an absence of factual knowledge, since LLMs can implicitly encode facts in their parameters and use them in question answering, while next-token prediction can produce factually accurate outputs; inaccuracies are a contingent limitation, not a necessary consequence of choosing generation over retrieval (Source 5, “Towards Understanding Factual Knowledge of Large Language Models”; Source 1, NIST AI RMF Generative AI Profile).

P
Proponent Rebuttal

The Opponent attacks a straw man: the claim says generative models "can lead to" inaccuracies, a modal statement that Source 1 (NIST) explicitly endorses when it notes statistical prediction "can produce factually accurate" outputs yet also inaccurate ones, so acknowledging occasional accuracy refutes nothing. The Opponent's RAG point also backfires, since Sources 4 and 17 present retrieval as an add-on remedy for the very hallucination problem arising from the next-token objective, and Source 24 shows that even with retrieval, models still hallucinate, confirming that the underlying generative mechanism, not retrieval, drives the inaccuracy risk.

Panel Review

3 reviewers assessed the evidence and the arguments.

Reviewer A · Claude

True
9/10

NIST's Generative AI Profile (Sources 1, 9) says confabulations are a natural result of generative model design: LLMs predict the next token and approximate the statistical distribution of training data, and this can produce factually inaccurate outputs. Nature (Source 2) shows next-word pretraining creates statistical pressure toward hallucination, and NeurIPS (Source 3) says models are trained on next-subword prediction rather than a fact-related objective. Source 5 adds that facts are stored implicitly in parameters rather than in an explicit knowledge base. The Opponent's RAG objection is weak: Sources 4 and 17 present retrieval as an add-on remedy for exactly this problem, and the claim's wording is hedged ('can lead to'). The claim is a standard, well-supported characterization of generative models, with only the minor caveat that it generalizes across all generative models and RAG hybrids exist.

Source issues

  • Sources 1 and 9 are the same NIST document, so they count as one source.

Evidence gaps

  • No source addresses non-language generative models such as image generators.

Precision issues

  • The phrase 'rather than retrieve facts' is slightly categorical and does not cover retrieval-augmented hybrid systems.

Reviewer B · GPT

Mostly True
8/10

NIST directly states that generative AI models approximate the statistical distribution of training data, with LLMs predicting the next token or word, and that this process can produce factually inaccurate or inconsistent outputs [1]. Nature likewise finds that next-word pretraining creates a statistical tendency toward hallucination, while the NeurIPS paper describes generative LMs as trained on statistical correlations between tokens rather than a fact-related objective [2, 3]. These sources directly support the claim for ordinary text-generating language models and its modal wording that inaccuracies can result. The wording is somewhat broad because retrieval-augmented generative systems can retrieve external evidence, but that augmentation does not negate the supported characterization of the underlying language-model generation mechanism; the claim is Mostly True.

Source issues

  • Sources 1 and 9 are duplicate publications of the same NIST profile and do not constitute independent corroboration.
  • Source 4 is a survey article and summarizes prior work rather than independently establishing every underlying mechanism.

Evidence gaps

  • The evidence chiefly concerns LLMs and text generation, whereas the phrase "generative models" could also encompass non-language generative systems.
  • No source establishes that every generative system lacks retrieval capabilities; retrieval-augmented generation is an important qualified case.

Precision issues

  • The phrase "Generative models" is broader than the LLM-focused evidence, although the reference to word sequences substantially signals the intended language-model scope.

Reviewer C · Gemini

True
10/10

The core assertion of the claim is directly supported by official government documentation and peer-reviewed scientific literature. Sources 1 and 9 (NIST AI RMF) specify that generative models approximate statistical distributions via next-token prediction rather than explicit factual retrieval, which inherently can cause confabulations and factual inaccuracies. Additional peer-reviewed studies (Sources 2, 3, and 5) confirm that LLMs optimize for token prediction rather than factual lookup, creating structural tendencies toward hallucination. The wording of the claim is accurately calibrated with the modal phrase 'can lead to,' which perfectly matches the evidence.

Source issues

  • Source 9 is a duplicate publication of the NIST framework found in Source 1.

Panel summary

NIST guidance and peer-reviewed research directly support the claim's central point: language models predict tokens, and that process can produce factual inaccuracies. The phrase “can lead to” accurately states a risk rather than an inevitable outcome. The main difference in assessment concerns scope: the evidence focuses on text-generating language models, while retrieval-augmented systems can also consult external information. Those qualifications do not overturn the claim's central description. The two NIST links are versions of the same document, not independent corroboration.

See the full panel summary

Create a free account to read the complete analysis.

Sign up free
The claim is
True
Score: 9/10
Confidence: 8/10 Spread: 2 pts

Only you will see this note.

Embed this verification

Every embed carries schema.org ClaimReview microdata — recognized by Google and AI crawlers.

True · Lenz Score 9/10 Lenz
“Generative models predict likely word sequences rather than retrieve facts, which can lead to factual inaccuracies.”
34 sources · 3-panel audit · Verified Oct 2026
See full report on Lenz →