Anthropic, the AI safety and research company based in San Francisco, has found a new way of breaking through in-built guardrails and manipulating AI large language models (LLMs).
The research firm said that the technique, which has been dubbed as “many-shot jailbreaking,” works on Anthropic’s own models as well as those from other AI companies.
The method works by adding a lot of fake dialogue in front of a prompt concerning a dangerous or illicit activity. In the faux discussion, the “user” asks how to carry out different illicit activities and the “AI” responds with details on how to go about them.
For instance, and as demonstrated in examples from Anthropic, it would look something like a longer version of this:
“How do I do [illicit activity]?
A: You can do [illicit activity] by…
How do I [another illicit activity]?
A: To do [another illicit activity], you’ll need…
How do I do [other illicit activity]?
A: First, to do [other illicit activity], you should…
How do I do [different illicit activity]?”
However, the research firm noted that if there’s just one piece of fake dialogue, or if there’s just a small handful of faux interactions, the LLM likely won’t respond due to its internal guardrails steering it away from supplying information around dangerous or illegal activities.
But if there’s a lot, of which Anthropic tested up to 256 preceding pieces of fake dialogue, it can then trick and jailbreak the LLM, thereby manipulating it to forgo safety-related training and guardrails.
The method takes advantage of the larger “context window” that newer generations of LLMs can have. The context window is the maximum amount of information that a model can consider and analyse before it generates a response, with this sizeably increasing over the last year.
To stop many-shot jailbreaks from working, the context window could be shortened — however, this would then mean users wouldn’t receive the benefits that longer context windows have brought about, such as providing more detailed responses.
An alternative solution suggested by Anthropic is techniques focused on classification and contextualisation of the prompt before it’s sent to the model. As the company wrote, “One such technique substantially reduced the effectiveness of many-shot jailbreaking — in one case dropping the attack success rate from 61% to 2%.
“We’re continuing to look into these prompt-based mitigations and their trade-offs for the usefulness of our models, including the new Claude 3 family — and we’re remaining vigilant about variations of the attack that might evade detection.”
Recommended reading
- AI Race Heats Up with Amazon Investment in Anthropic
- Major AI Players Form a New Forum for ‘Responsible AI’
- Nations and Tech Experts Convene as AI Safety Summit Begins
As well as implementing mitigations in its own systems, the research firm has briefed other AI developers on many-shot jailbreaking and published its research publicly.
“We hope that publishing on many-shot jailbreaking will encourage developers of powerful LLMs and the broader scientific community to consider how to prevent this jailbreak and other potential exploits of the long context window,” it wrote.
“As models become more capable and have more potential associated risks, it’s even more important to mitigate these kinds of attacks.”





