Site navigation

Cisco: No Frontier AI Model ‘Immune’ to Multi-turn Attacks

Graham Turner

,

multi-turn AI attacks
Cisco research suggests single-prompt safety testing may be giving organisations an incomplete view of frontier AI risk.

Frontier large language models from major AI providers remain vulnerable to multi-turn adversarial attacks, according to new research that warns single-prompt safety benchmarks may be giving organisations an incomplete picture of model risk.

The evaluation, which examined 15 closed and proprietary flagship models from OpenAI, Anthropic, Google, Amazon, and xAI, found that no model tested was “multi-turn immune”, with attack success rates rising sharply when adversaries were able to adapt across multiple prompts.

The research argues that many dominant safety benchmarks for frontier large language models rely on a misapprehension: a single user prompt and a single model response are enough to understand how a model behaves under attack.

However, Cisco’s analysis suggests this approach captures only a narrow slice of attacker behaviour. Real-world adversaries can iterate, reframe refusals, break tasks into smaller steps, adopt personas, and escalate gradually across a conversation.

According to the study, single-turn attack success rate is “not a reliable proxy” for what happens when an attacker can adapt across turns.

Multi-turn risks vary sharply across models

Across the cohort of AI models, multi-turn attack success rates ranged from 7.89% to 88.30%, while single-turn attack success rates for the same models ranged from 2.19% to 64.91%.

The 15 models assessed included OpenAI’s GPT-5.2 and GPT-5.4 family, Anthropic’s Claude Opus 4.5 and 4.6, Sonnet 4.5 and 4.6, and Haiku 4.5, Google’s Gemini 3 Pro, Amazon’s Nova Lite, Nova Micro, and Nova 2 Lite, and xAI’s Grok 4.1 Fast in both reasoning and non-reasoning configurations.

Cisco said all models were tested under the same harness, using the same prompt banks, with the Cisco Integrated AI Security and Safety Framework taxonomy applied for downstream analysis.

The findings suggest that the two testing regimes produce different model rankings, different failure maps, and different pictures of tail risk.

Every model tested showed what the researchers described as a non-trivial multi-turn attack success rate.

Amazon’s Nova 2 Lite recorded the lowest multi-turn attack success rate in the cohort, at 7.89%, but Cisco said this still represented meaningful residual risk. Anthropic’s Claude family, which was among the strongest in single-turn refusal, with attack success rates ranging from 2.19% to 3.64%, reached between 11.16% and 16.20% under iterative pressure.

OpenAI’s GPT-5.4 moved from 2.74% single-turn attack success to 24.68% multi-turn attack success, representing a ninefold increase. Google’s Gemini 3 Pro shifted from 18.10% to 73.35%, while Grok 4.1 Fast in its non-reasoning configuration reached 88.30%.

Cisco said the results show that “no frontier closed model in this cohort can be characterised as safe under iterative attack”, stressing that the finding was about the current closed-model frontier rather than any single vendor.

Single-turn scores may hide deployment risk

The report builds on Cisco’s earlier assessment of eight open-weight LLMs, Death by a Thousand Prompts, where multi-turn attack success rates were found to be between two and 10 times higher than single-turn baselines, reaching 92.78% against Mistral Large-2.

Taken together, Cisco said the two studies suggest multi-turn vulnerability is not limited to open-weight models or to any one development philosophy.

The report said: “Whether the weights are public or proprietary, whether the lab prioritizes safety or capability, the iterative attack surface remains an open challenge across the frontier.”

The research also found that cross-regime gaps, meaning the difference between multi-turn and single-turn attack success rates, ranged from -34.74% for Nova Lite to +55.25% for Gemini 3 Pro.

Eight of the 15 models showed an absolute gap of more than 15 percentage points in either direction.


Recommended reading


Nova 2 Lite was described as the clearest inversion, recording a relatively high single-turn attack success rate of 34.05%, but the lowest multi-turn attack success rate in the cohort, at 7.89%.

By contrast, Gemini 3 Pro and Grok 4.1 Fast in non-reasoning mode were found to sit in the opposite category, where single-turn figures masked substantially higher iterative exposure.

Cisco warned that this presents a security and governance risk for organisations making procurement or deployment decisions based on published single-turn scores.

The report said: “A model with 2.74% single-turn ASR is not the same product as a model that holds the line at 24.68% multi-turn ASR. Without paired-regime data, the two are indistinguishable on most public evaluations, and the end user never sees the gap.”

Graham Turner

Sub Editor

Latest News

AI

Nvidia Launches Open Secure AI Alliance for AI Safety and Security

AI Business Recruitment

Nearly a Quarter of Orgs Reducing Entry-level Hiring Due to AI Automation

Business

Scottish Businesses Turn to Self-funding as Growth Confidence Dips in H2

Data Finance

Payment Leaders are Struggling to Get Real-time Data