Site navigation

Even the Best AI Models Hallucinate, New Study Finds

Graham Turner

,

AI model hallucinations
A new report from Cornell University uncovers the frequent inaccuracies of AI language models, especially on less common topics where reliable sources are sparse.

A recent study from researchers at Cornell, the universities of Washington and Waterloo introduces the concept of WILDHALLUCINATIONS, a benchmark designed to evaluate the accuracy of large language models (LLMs) in generating factual information.

The research addresses a significant challenge with LLMs: their tendency to produce “hallucinations,” or outputs that contain incorrect or unsupported information.

Unlike previous benchmarks, WILDHALLUCINATIONS tests LLMs on a wide variety of real-world entities – many of which do not have Wikipedia pages – extracted from user-chatbot interactions.

The benchmark evaluates 118,785 text generations from 15 different LLMs, covering 7,919 entities across diverse fields such as computing, culture, and finance. The study found that LLMs are more prone to hallucinate when dealing with entities that lack Wikipedia pages. This suggests that LLMs struggle more with lesser-known or niche topics, with AI model hallucinations becoming more frequent where reliable sources are sparse.

Among the LLMs tested, hallucination rates varied by domain. Models performed better in fields like geography and computing but struggled significantly with topics related to people and finance.

The study also tested whether adding a retrieval component – where the model pulls information from the web in real-time – could reduce hallucinations. The results showed only a slight improvement, indicating that retrieval alone is insufficient to address the issue.


Recommended reading


The benchmark also highlights differences in performance between various models. For instance, GPT-4 and GPT-3.5 achieved the highest factual accuracy, outperforming even those models with retrieval capabilities. Notably, even within the same family of models, larger versions did not always perform better.

The report also found that some AI model hallucinations occur less often when dealing with frequently mentioned entities. In contrast, rare or lesser-known entities led to a sharp drop in accuracy, particularly for models without retrieval capabilities.

WILDHALLUCINATIONS is now publicly available on Hugging Face, allowing researchers to further explore and mitigate the issue of LLM hallucinations.

Graham Turner

Sub Editor

Latest News

AI

Nvidia Launches Open Secure AI Alliance for AI Safety and Security

AI Business Recruitment

Nearly a Quarter of Orgs Reducing Entry-level Hiring Due to AI Automation

Business

Scottish Businesses Turn to Self-funding as Growth Confidence Dips in H2

Data Finance

Payment Leaders are Struggling to Get Real-time Data