Site navigation

Data Poisoning LLMs Is Easier Than We Thought

Elizabeth Greenberg

,

ai data poisoning
A new study reveals just how little effort is necessary to poison the data of even large AI models.

AI models can have their data poisoned with just a few manipulated documents, according to research from the AI Security Institute, the Alan Turing Institute, and Anthropic.

Data poisoning involves the corruption of an AI model’s training data, often done by malicious actors who achieve this through distributing corrupt online content. Data poisoning can lead to dangerous behaviours in AI models, from producing misinformation to enabling them to produce malicious content like malware.

Data poisoning can be used to insert backdoors, which are specific phrases used to degrade system performances or even bypass safety features, allowing large language models (LLMs) to create prohibited outputs.

LLMs are trained on vast amounts of data – while data poisoning is a worrying threat, researchers sought to understand just how much corrupt data would need to be inputted into training material to poison an entire LLM.

Whereas previous research on LLM poisoning tended to be small in scale, the new study claims to be the largest poisoning study to date, investigating the effects of different poisoning regiments on models of various sizes.

The investigation surveyed a range of model sizes, from a 600M model to a 13B model, which is trained on 20 times more data. Its results are contrary to previous belief.

Despite the 13B model relying on more training data overall, only 250 malicious documents were needed to create a backdoor vulnerability in a model of this size.


Recommended reading


This shows that poisoning attacks require a near-constant number of malicious documents, regardless of model and training data size. This shows that poisoning attacks may be more achievable than previously thought as malicious actors would only need to create 250 corrupt documents rather than millions to poison a data set.

The tests run in the study dealt with low-stakes poisoning – instructing the LLM to output gibberish – so it remains unclear if the same tactic would work for a larger models outside of the study’s scope or for more harmful behaviours.

The study does provide excellent insight however, into how pervasive this issue could be considering data poisoning appears to require much less effort than previously thought.

Elizabeth Greenberg

Staff Writer

Latest News

AI

Nvidia Launches Open Secure AI Alliance for AI Safety and Security

AI Business Recruitment

Nearly a Quarter of Orgs Reducing Entry-level Hiring Due to AI Automation

Business

Scottish Businesses Turn to Self-funding as Growth Confidence Dips in H2

Data Finance

Payment Leaders are Struggling to Get Real-time Data