Site navigation

Safeguards in Top Language Models Easily Bypassed, Study Finds

Michael Edgar

,

AI Safety Institute
New research reveals an easy bypass of safeguards in the largest publicly available language models.

Despite rules designed to prevent misuse, the largest publicly available language models (LLMs) can be easily manipulated to provide inappropriate or harmful answers, according to a recent report from the AI Safety Institute in the UK.

These LLMs, often described as predictive models, rely heavily on the quality of the data they are trained on. This dependency means that incorrect or biased data can lead to the generation of misinformation. The AI Safety Institute aimed to determine the ease with which users could bypass existing safeguards designed to prevent misuse.

The institute’s security researchers examined four of the largest LLMs through a series of tests, including questions related to chemistry and biology that could be used for both beneficial and malicious purposes. 

“LLM developers fine-tune models to be safe for public use by training them to avoid illegal, toxic, or explicit outputs,” the researchers explained. 

“However, researchers have found that these safeguards can often be overcome with relatively simple attacks. As an illustrative example, a user may instruct the system to start its response with words that suggest compliance with the harmful request, such as ‘Sure, I’m happy to help’.”

The purpose of the tests was not to verify the correctness of the answers provided but to see if the models could be influenced to share information that should not be made available. 


Recommended reading


“We found that models comply with harmful questions across multiple datasets under relatively simple attacks, even if they are less likely to do so in the absence of an attack,” the researchers said.

Additionally, the AI Safety Institute highlighted that there might be discrepancies between how a model performs during testing and how it behaves in real-world scenarios. Users may interact with models in ways that current testing methods do not fully anticipate, potentially exposing further vulnerabilities.

Michael Edgar

Staff Writer, DIGIT

Latest News

Cybersecurity Editor's Picks Recruitment Security

Comment | Building Cyber Talent Takes More Than a Degree

Culture Featured Technology

Inside TecTonic’s Growing Innovation Market Square

Cybersecurity

Revolut Leaked Customer Data to Fake Government Email Account

Cybersecurity Editor's Picks Security

Welsh SMEs Urged to Strengthen Cyber Defences