AI chatbots are misrepresenting the news to users on a mass scale, new research from the BBC found.
The news corporation investigated four publicly available AI chatbots – OpenAI’s ChatGPT, Google’s Gemini, Microsoft’s Copilot, and Perplexity.
To conduct its research, the BBC asked each chatbot 100 queries having to do with the news, prompting the AI systems to use BBC News sources where possible, with full access to the BBC’s news and all other news sources.
Alarmingly, the research found that just over half (51%) of all AI responses were flagged as having ‘significant issues’ by expert journalists recruited to review the responses for inaccuracies and false claims.
Almost all (91%) of responses contained at least ‘some issues’ according to the reviewers.
Gemini saw the largest percentage of significant issues (46%), followed by Copilot, Perplexity, and ChatGPT.
About one in five answers had factual errors, such as on numbers, dates, or statements, and 13% of quotes sources directly to the BBC contained alternations or did not exist in the cited BBC article.
Other inaccuracies included that Nicola Sturgeon is still the first minister of Scotland, and that Rishi Sunak still serves as the UK prime minister, which were both produced by ChatGPT.
Gemini was found to have misrepresented NHS guidance on smoking and vaping; Perplexity misquoted a statement from One Direction singer Liam Paynes family after his death
Copilot provided false information on how Gisèle Pelicot – who took 50 men to court for raping her while unconscious last year – realised she had been the victim of rape. Microsoft’s chatbot was also found to have referenced out of date articles for news about Scottish independence.
Recommended reading
- Misinformation and Disinformation Top WEF Short Term Threat List, Again
- ICO Launches Consultation on Generative AI Accuracy
- ChatGPT to the Left, Meta’s LLaMA to the Right: Stuck in the Middle of AI Bias
The report from the BBC also found that chatbots often failed to discern between fact and opinion, and often added their own editorialisation to responses.
Reviewers found editorialisation in over 10% of Copilot and Gemini responses, 7% from Perplexity and 3% from ChatGPT.
According to the review, missing context was one of the most common issues identified with AI responses, especially when a response requires multiple perspectives or facets.
The review showed a startling lack of factual accuracy across chatbots, even on the most basic of principles, with one in five responses which used a BBC article as a source having factual inaccuracies.
These errors have real-world consequences, especially if they concern natural disasters, weather, current global conflicts, health advice, or politics.
The World Economic Forum has said that misinformation and disinformation is the top short-term threat facing global stability, and with the oft-unchecked proliferation of AI, it seems this is an appropriate concern.





