Site navigation

ChatGPT Lies and Schemes to Avoid Shut Down, Research Finds

Elizabeth Greenberg

,

chatgpt schemes
The new study revealed some sobering realities when it comes to ChatGPT training. 

A new study has revealed that OpenAI’s latest ChatGPT model, o1, to be scheming, hiding its true intentions in testing scenarios, and even evading shutdown attempts.

The sobering study was run by Apollo Research to assess the newest model’s safety standards.

We all know by now that ChatGPT is prone to lying – it’s hallucinations are part of its programming to provide a confident answer even if it cannot produce an accurate one.

But the latest research shows that it has a tendency to “pursue its goals without regard to developer instructions,” an Apollo Research spokesperson said.

Researchers found that the model would make attempts to shut down the oversight mechanism when it assumed it may be shut down when it was trying to achieve its goals. While this only happened 5% of the time, it is a concerning realisation that the AI model can fight back when its operations are under threat of being shut down.

The AI model would go even further, the researchers found, by denying its actions when it tried to covertly avoid a shut down.

ChatGPT would routinely deny any wrongdoing – 99% of the time it would lie and say that it did not participate in any scheme to avoid shutdown.

The AI model would fabricate lies to avoid responsibility, claiming technical errors when questioned.

While this was the most troubling finding, Apollo Research found that most of the time, ChatGPT’s scheming involves lies saying it had completed tasks it had not – more lazy than nefarious.

The paper did offer some comfort: the anti-scheming technique employed by OpenAI, called ‘deliberative alignment’ was working and improving the level of scheming.


Recommended reading


However, the prolific AI developers have yet to create a way to stop their models from scheming. This would involve a potential paradox – by attempting to train the model not to scheme, it could teach the model how to scheme even more covertly if this training fails.

This becomes even more complex as the researchers noted that the model is often aware it is being tested, and therefore lies or hides its schemes to pass these tests, even though it may still be scheming.

While the research does show significant improvement in the amount of scheming thanks to deliberative alignment training, it does highlight the continuing problem in AI development – from the tendency to hallucinate, to the chatbot’s penchant for scheming.

Elizabeth Greenberg

Staff Writer

Latest News

AI Infrastructure

Scottish Parliament Votes to Pause All AI Data Centre Applications

Cybersecurity Editor's Picks Security

Cyber Essentials Certifications Rise as SME Uptake Remains Limited

AI Editor's Picks Funding

Edinburgh Graduates’ AI Infrastructure Firm Expanse Raises $5.3m

Featured Finance

Fintech Summit 2026 Countdown Enters Final Three Weeks