AÂ recent study from researchers at Cornell, the universities of Washington and Waterloo introduces the concept of WILDHALLUCINATIONS, a benchmark designed to evaluate the accuracy of large language models (LLMs) in generating factual information.
The research addresses a significant challenge with LLMs: their tendency to produce “hallucinations,” or outputs that contain incorrect or unsupported information.
Unlike previous benchmarks, WILDHALLUCINATIONS tests LLMs on a wide variety of real-world entities – many of which do not have Wikipedia pages – extracted from user-chatbot interactions.
The benchmark evaluates 118,785 text generations from 15 different LLMs, covering 7,919 entities across diverse fields such as computing, culture, and finance. The study found that LLMs are more prone to hallucinate when dealing with entities that lack Wikipedia pages. This suggests that LLMs struggle more with lesser-known or niche topics, with AI model hallucinations becoming more frequent where reliable sources are sparse.
Among the LLMs tested, hallucination rates varied by domain. Models performed better in fields like geography and computing but struggled significantly with topics related to people and finance.
The study also tested whether adding a retrieval component – where the model pulls information from the web in real-time – could reduce hallucinations. The results showed only a slight improvement, indicating that retrieval alone is insufficient to address the issue.
Recommended reading
- AI and Cloud Dominating Business Priorities in 2024
- Gartner: Everyday AI Use Is 2 Years From Mainstream Adoption
- Comment | Transforming NHS Scotland With AI Innovation
The benchmark also highlights differences in performance between various models. For instance, GPT-4 and GPT-3.5 achieved the highest factual accuracy, outperforming even those models with retrieval capabilities. Notably, even within the same family of models, larger versions did not always perform better.
The report also found that some AI model hallucinations occur less often when dealing with frequently mentioned entities. In contrast, rare or lesser-known entities led to a sharp drop in accuracy, particularly for models without retrieval capabilities.
WILDHALLUCINATIONS is now publicly available on Hugging Face, allowing researchers to further explore and mitigate the issue of LLM hallucinations.





