Site navigation

Security Experts Uncover Major Flaws in Hundreds of AI Benchmarks

Tom Quinn

,

AI benchmark
AI researchers have called for urgent reform of model testing, warning that current benchmarks give developers and users a false sense of safety.

Hundreds of the benchmark tests used to vet today’s most powerful AI models are riddled with flaws, a new research report has warned, raising serious questions about how developers measure the safety of widely used LLMs.

Researchers from the UK’s AI Security Institute, together with experts from Stanford, Berkeley, and Oxford, examined over 440 AI safety benchmarks, only to find that most rely on vague definitions and shaky analytical foundations.

The study found that around half of the benchmarks aimed to measure abstract ideas such as reasoning or harmlessness, without clearly defining what those terms mean, making it difficult to ensure that benchmarks are testing what they intend to.

Meanwhile, just 16% of the benchmarks used statistical methods when comparing model performance – meaning that differences between systems or claims of superiority could be due to chance rather than genuine improvement. 

The researchers warned that weak evaluation methods could fuel misleading, even reckless claims. For instance, a model that performs well on medical exams might be touted as having doctor-level expertise, despite lacking any real-world judgment, experience, or ethical reasoning.

“Benchmarks underpin nearly all claims about advances in AI,” said Andrew Bean, a researcher at the Oxford Internet Institute and lead author of the study. 

“But without shared definitions and sound measurement, it becomes hard to know whether models are genuinely improving or just appearing to.”

The lack of consistent, verifiable benchmarks (especially given the lack of national guardrails) presents a huge challenge in identifying safety issues among prominent AI models, which is particularly concerning given the rate at which high-profile cases of AI harm are stacking up.

In the latest headline-grabbing case of AI gone wrong, Google was forced to pull its Gemma models from the firm’s AI Studio after the bot spewed out unfounded and defamatory allegations of sexual misconduct against Republican Senator Marsha Blackburn.

While Google stressed that it ‘never intended’ for this particular model to be used by consumers for fact-finding purposes, the incident highlights the importance of tech firms using reliable, publicly available safety benchmarks in their race to produce bleeding-edge tech.


Recommended reading


The good news is, these problems are fixable. Among their recommendations for better benchmarking, the researchers said that AI developers should ensure the tests represent real-world conditions, cover the full scope of the target behaviour, and use precise, operational definitions for the concepts being measured.

They also said that detailed error analysis is needed to understand why a model fails, and that devs should use statistical methods to report uncertainty and allow for comparisons. 

To help this along, the team has created a Construct Validity Checklist, a practical tool for researchers, developers, and regulators to assess whether an AI benchmark follows sound design principles before relying on its results. 

“This work reflects the kind of large-scale collaboration the field needs,” said Dr. Adam Mahdi, one of the project’s researchers. “By bringing together leading AI labs, we’re starting to tackle one of the most fundamental gaps in current AI evaluation.”

Tom Quinn

Staff Writer, DIGIT

Latest News

AI

Nvidia Launches Open Secure AI Alliance for AI Safety and Security

AI Business Recruitment

Nearly a Quarter of Orgs Reducing Entry-level Hiring Due to AI Automation

Business

Scottish Businesses Turn to Self-funding as Growth Confidence Dips in H2

Data Finance

Payment Leaders are Struggling to Get Real-time Data