As more and more UK consumers turn to AI chatbots for help with product purchases, financial quandaries, medical tips, and general information, consumer researcher Which? investigated the accuracy of these AI tools.
Which?’s research found that more than half of UK consumers use AI to search for information across the internet, with around a third thinking that AI search is already more important than traditional web searches.
Around half of those using AI told Which? that they trust the information AI provides them to a ‘great’ or reasonable’ extent, but just how ‘reasonable’ is this level of trust?
Which? tested six of the most popular AI tools on the market: ChatGPT, Google Gemini on its own and as it appears in Google searchers, Microsoft Copilot, Meta AI, and Perplexity.
The consumer research advocacy asked the chatbots a series of questions ranging from medical queries to financial advice, product reviews, legal advice, and travel issues.
Which? found that while AI tools can be useful for general research, they often provide factual inaccuracies or incomplete advice.
All the AI chatbots made continual factual errors – Which? gave the example of getting broadband compensation prices wrong, as well as the ISA allowance figure.
The chatbots were found to often give incomplete advice which could be misleading, such as not considering different legal rules for different UK regions.
However, Which? also found that chatbots to often be overconfident in their presentation of advice without proper ethical considerations, often neglecting to recommend users turn to professionals when discussing legal and financial matters.
Further, Which? found that AI chatbots often relied on weak sources, which could even include online forum threads. They often had sources that were vague or even non-existent, making it difficult for consumers to track the AI’s source of information.
Finally, some of the tools also suggested dodgy services with high costs over free and safe tools, meaning that users could be duped into overpaying for services or employing a less reliable tool.
Recommended reading
- Google Launches AI Search in the UK | Will it Kill Web Traffic?
- Over Half of Consumers Don’t Trust AI Search Results
- AI Adoption Stumbles as Report Highlights Employee-leadership Disconnect
Overall, Which? found that Perplexity scored the highest overall across these different metrics, with a total score of 71%. It had the highest score for accuracy (72%), relevance (73%), clarity (73%) and usefulness (70%), but was surpassed by Google’s Gemini models for ethical responsibility.
Meta’s AI came in last place, with a 55% overall score, coming in last across the board.
Interestingly, ChatGPT, the most popularly used model, came in second to last, with a 64% overall score, trailing behind Microsoft Copilot (68%), Gemini (69%), and Google AI Overview (70%).
Interesting, when comparing Google AI Overview with Gemini, Which? also found discrepancies in accuracy across different questions.
Google AI Overview handled questions on legal and health better than Gemini, which scored higher on financial and consumer rights queries.
As AI tools become an increasingly popular choice for people’s first and last stop to find information, understanding their faults in accuracy and ethics is evermore vital.





