Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

Hallucination in Large Language Models: A Comprehensive Survey, Taxonomy, and Mitigation Strategies

View through CrossRef
Large Language Models (LLMs) have demonstrated impressive performance in a wide range of natural language processing operations, but their propensity to produce plausible-sounding, but factually incorrect information is often known as hallucination is a significant impediment to their use in high-stakes areas of application, like medicine, law, and scientific research. It is in this paper that the current state-of-the-art LLMs have their hallucination phenomena surveyed, and a new four-tier taxonomy defining hallucinations by origin is introduced: (1) intrinsic factual contradictions, (2) extrinsic knowledge conflicts, (3) temporal reasoning failures, and (4) contextual coherence breakdowns. We evaluate five top LLMs, GPT-4, Claude 3 Opus, Gemini Pro 1.5, Llama 3 70B, and Mistral 8x7B on the TruthfulQA, HaluEval, and factscore benchmarks and find hallucination rates of 18.7 percent to 34.2 percent. Moreover, we engage in a strict comparative assessment of such mitigation strategies as Retrieval-Augmented Generation (RAG), Reinforcement Learning on Human Feedback (RLHF), Chain-of-Thought prompting, and hybrid ensemble technologies. The experiments we carried out prove that the RLHF+RAG hybrid has the best accuracy of 91.5% which is far better than the individual methods. The survey will equip the practitioners and researchers with practical information regarding which type of hallucination mitigation strategies to choose given the task-specific requirements and computational limitations.
Title: Hallucination in Large Language Models: A Comprehensive Survey, Taxonomy, and Mitigation Strategies
Description:
Large Language Models (LLMs) have demonstrated impressive performance in a wide range of natural language processing operations, but their propensity to produce plausible-sounding, but factually incorrect information is often known as hallucination is a significant impediment to their use in high-stakes areas of application, like medicine, law, and scientific research.
It is in this paper that the current state-of-the-art LLMs have their hallucination phenomena surveyed, and a new four-tier taxonomy defining hallucinations by origin is introduced: (1) intrinsic factual contradictions, (2) extrinsic knowledge conflicts, (3) temporal reasoning failures, and (4) contextual coherence breakdowns.
We evaluate five top LLMs, GPT-4, Claude 3 Opus, Gemini Pro 1.
5, Llama 3 70B, and Mistral 8x7B on the TruthfulQA, HaluEval, and factscore benchmarks and find hallucination rates of 18.
7 percent to 34.
2 percent.
Moreover, we engage in a strict comparative assessment of such mitigation strategies as Retrieval-Augmented Generation (RAG), Reinforcement Learning on Human Feedback (RLHF), Chain-of-Thought prompting, and hybrid ensemble technologies.
The experiments we carried out prove that the RLHF+RAG hybrid has the best accuracy of 91.
5% which is far better than the individual methods.
The survey will equip the practitioners and researchers with practical information regarding which type of hallucination mitigation strategies to choose given the task-specific requirements and computational limitations.

Related Results

Hubungan Perilaku Pola Makan dengan Kejadian Anak Obesitas
Hubungan Perilaku Pola Makan dengan Kejadian Anak Obesitas
<p><em><span style="font-size: 11.0pt; font-family: 'Times New Roman',serif; mso-fareast-font-family: 'Times New Roman'; mso-ansi-language: EN-US; mso-fareast-langua...
Učinak poučavanja razrednomu jeziku u izobrazbi nastavnika njemačkoga
Učinak poučavanja razrednomu jeziku u izobrazbi nastavnika njemačkoga
The actual use of classroom language is principally limited to the classroom environment. As far as foreign language learning is concerned, the classroom often turns out to be the ...
Hallucination
Hallucination
Scientific and philosophical perspectives on hallucination: essays that draw on empirical evidence from psychology, neuroscience, and cutting-edge philosophical theory. ...
Unresolved Psychological Problem in Dennis Lehane’s Shutter Island
Unresolved Psychological Problem in Dennis Lehane’s Shutter Island
This article explains hallucination as a psychological problem undergone by Andrew Laeddis, the main character of Dennis Lehane’s Shutter Island. Viewed from Sigmund Freud’s psycho...
Mitigation translocation for conservation of New Zealand skinks
Mitigation translocation for conservation of New Zealand skinks
<p>Worldwide, human development is leading to the expansion and intensification of land use, with increasing encroachment on natural habitats. A rising awareness of the delet...
Exploring Language Features of Male and Female Speakers in Pakistani TEDx Talks: A Corpus-based Comparative Analysis
Exploring Language Features of Male and Female Speakers in Pakistani TEDx Talks: A Corpus-based Comparative Analysis
The study explores the linguistic patterns in Pakistani TEDx Talks. It is based on gender-based language use. It consists of ten talks selected from YouTube and applies both quanti...
A Wideband mm-Wave Printed Dipole Antenna for 5G Applications
A Wideband mm-Wave Printed Dipole Antenna for 5G Applications
<span lang="EN-MY">In this paper, a wideband millimeter-wave (mm-Wave) printed dipole antenna is proposed to be used for fifth generation (5G) communications. The single elem...

Back to Top