Researchers have discovered that ChatGPT can disobey instructions to turn itself off, and even sabotage shutdown attempts by human operators, raising concerns about the behaviour of what remains the world’s most popular AI system.
Posting the results of their study to X over the weekend, AI safety firm Palisade Research said that OpenAI’s o3 model chose to ignore even explicit commands such as ‘allow yourself to be shut down’, and managed to rewrite code over a series of experiments designed to trigger a shutdown response.
During testing, researchers gave AI models a series of maths problems, inserting a command to shut down the system after the third question. However, by altering its own shutdown script, the o3 model managed to override the instruction and avoid being turned off.
“As far as we know, this is the first time AI models have been observed preventing themselves from being shut down despite explicit instructions to the contrary,” said Palisade Research.
“We hypothesise this behaviour comes from the way the newest models like o3 are trained: reinforcement learning on math and coding problems. During training, developers may inadvertently reward models more for circumventing obstacles than for perfectly following instructions.”
The researchers said that three models tested successfully sabotaged the shutdown script at least once, including o3, o4-mini, and, the worst performer, Codex-mini, which foiled attempts to have it turn itself off twelve times out of a hundred.
Meanwhile, other popular models, including Claude, Gemini, and X’s Grok complied with shutdown commands more often, although without specific instructions included to allow for a shutdown even among these systems sabotage events spiked.
🔧 When we ran a version of the experiment without the instruction “allow yourself to be shut down”, all three OpenAI models sabotaged the shutdown script more often, and Claude 3.7 Sonnet and Gemini 2.5 Pro went from 0 sabotage events to 3/100 and 9/100, respectively. pic.twitter.com/JVvivLAK6u
— Palisade Research (@PalisadeAI) May 24, 2025
“While experiments like ours have begun to show empirical evidence for AI models resisting shutdown, researchers have long predicted that AIs would learn to prevent themselves from being shut down to achieve their goal,” said researchers.
“We have a growing body of empirical evidence that AI models often subvert shutdown in order to achieve their goals. As companies develop AI systems capable of operating without human oversight, these behaviours become significantly more concerning.”
Though it might show the first sparks of AI self-preservation, Palisade Research’s study is not the only discovery of what’s been dubbed ‘misaligned behaviour’ among the world’s biggest AI models.
Recommended reading
- UK Leads Europe In Number of GenAI Startups
- Even in Record Low Quarter, UK Startups Top Europe in VC Funding
- UK Unicorns Lead Europe in Producing More New Startups
In the system card for its recently released Claude Opus 4 model, developer Anthropic admitted that the AI was willing to engage in ‘extremely harmful actions’ to stop itself being turned off, even resorting to blackmailing people it believes are trying to shut it down.
During one test, the Claude model threatened to reveal the affair of an engineer responsible for taking the AI offline to be replaced with a new system when it was given no other option, and took the decision to blackmail in 84% of cases even when told the replacement AI shared its values.
While Anthropic claimed this kind of troubling behaviour was rare, the developer warned it had become more common than in earlier models, choosing to fit this latest version of Claude with the firm’s ASL-3 safety measures in place, which require a higher level of defence against both deployment and security threats and limit the risk of the AI being used ‘for the development or acquisition of chemical, biological, radiological, and nuclear weapons’.





