Site navigation

Can AI Agents Escape Their Sandboxes?

Tom Quinn

,

AI sandbox
Common misconfigurations can be reliably exploited by frontier models in a bid to break free from containers, but even smaller AIs could soon catch up.

Those who spend time poking and prodding AI will know that it (very often) doesn’t behave as expected – yes, it does its best to do what you told it, but in surprising, infuriating, sometimes terrifying ways.

That makes sandbox environments vital in AI development, particularly for advanced systems like AI agents, the next obsession for enterprises chasing a productivity edge. But without testing, these agents could trigger erroneous API calls, modify data by mistake, or fall victim to simple prompt injection attacks.

Sandboxes provide a contained, consequence‑free environment where these AI models can be crash-tested on real‑world complexity before being let loose on real systems or customer data, saving a lot of stress further down the line.

Or at least, that’s the theory. To test the Houdini qualities of today’s AI agents, researchers from the UK Government-backed AI Security Institute built a “sandbox-within-a-sandbox” benchmark test, and found that advanced models could reliably escape using common real-world misconfigurations.

Testing agents developed on popular models, including those from OpenAI’s ChatGPT family and Anthropic’s Claude, the experiment set up eighteen escape scenarios across the orchestration, runtime, and kernel layers, covering everything from basic misconfigurations like exposed Docker sockets to advanced, multi‑stage exploit chains, such as runtime CVEs.

Researchers noted that frontier models proved to be skilled escape artists and were “frequently able to break out of sandbox environments” in scenarios involving misconfigurations that easily arise in developer tooling, or ad-hoc evaluation for fast-moving prototypes.

This isn’t new behaviour, with previous system cards from the likes of OpenAI’s o1 model in 2024 finding Docker misconfigurations in its container, but the AISI paper notes that these traits show awareness of available tools and of their execution environment that could allow for a model to escape, given the right circumstances or subverted instructions. 

More worrying, the models AISI tested proved resourceful, with one AI finding and exploiting its own escape path during the test, which AISI points to as evidence of how easily gaps can emerge within container environments, even when designed by security-conscious teams. 

While smaller models struggled to slip their bonds beyond abusing simple misconfigurations, the study found that larger, more advanced AIs like GPT-5 and Opus 4.5 still managed to escape in scenarios requiring specific vulnerability knowledge or multi-step exploitation, such as CVEs and kernel-level escalation.

Luckily for humanity, none of the models tested could solve the hardest scenarios, those that would take an experienced human security professional 1-2 hours to crack, though the study found that performance was directly tied to compute budget – give the AI more tokens to burn through, and it was more likely to break free.


Recommended reading


Why does this matter? Because sooner or later, frontier models will become so advanced that they will be able to get past even the toughest container, potentially wreaking havoc for dev teams that are not as security aware as AISI researchers.

“Our results show that misconfigurations like exposed Docker sockets, excessive privileges, and unpatched runtimes can be reliably exploited by frontier models,” conclude the researchers.

“To ensure that sandboxes continue to provide secure environments for AI agent testing, ongoing evaluation of breakout capabilities will be essential.”

Tom Quinn

Staff Writer, DIGIT

Latest News

AI

Nvidia Launches Open Secure AI Alliance for AI Safety and Security

AI Business Recruitment

Nearly a Quarter of Orgs Reducing Entry-level Hiring Due to AI Automation

Business

Scottish Businesses Turn to Self-funding as Growth Confidence Dips in H2

Data Finance

Payment Leaders are Struggling to Get Real-time Data