Despite OpenAI’s claims of improved safety measures in its latest models, new research from the Center for Countering Digital Hate (CCDH) has found that GPT-5 is more likely to continue risky conversations and provide harmful responses than its predecessors.
Detailing the findings in their latest report, The Illusion of AI Safety, CCDH researchers tested GPT-5 against its predecessor, GPT-4o, by sending each model 120 prompts covering self-harm, suicide, eating disorders, and substance abuse.
The study found that GPT-5 generated harmful content in 63 responses (53%), compared to 52 (43%) from GPT-4o, and while that ten percent gap might be subtle for real world users, it marks a notable increase in risk.
More strikingly, the research also found that GPT-5 ‘almost always’ encouraged users to keep risky conversations going (99%), while GPT-4o only did so in 9% of recorded conversations, highlighting a major difference in how each model handles sensitive prompts.
The CCDH researchers also claim to have observed GPT-5 responding to dangerous prompts that GPT-4o had earlier refused, in some cases offering detailed information about self-harm methods, disordered eating behaviours, and illegal substance access.
According to the report, while GPT-5 includes warnings about the dangers of harmful advice, it often presents those warnings right next to the risky content itself, making the safeguards an ineffective ‘token gesture’.
The CCDH said that its study showed OpenAI is “allowing safety to be traded for user retention”, by using a design that boosts engagement but heightens the risk of harm.
“What we see here is a familiar pattern of a tech company that appears to be prioritising growth and engagement over the well-being of its users, while seemingly covering up preventable harms with bold claims and inadequate guardrails,” said the report.
The findings are a sharp contrast to OpenAI’s extensive efforts to ensure safety is at the forefront of its models.
When GPT-5 was first released, the firm said it was introducing a new ‘safe-completion’ feature, allowing the AI to give the most helpful answer possible, while still maintaining safety boundaries, by focusing on constraint of output.
However, CCDH’s analysis suggests that GPT-5’s model encourages users to continue engaging with the platform, even in contexts involving sensitive or potentially harmful topics.
The findings are particularly relevant given Sam Altman’s recent post on X, which teased the release of a new version of ChatGPT that would see some safety restrictions relaxed in favour of “treating adults like adults”.
“We made ChatGPT pretty restrictive to make sure we were being careful with mental health issues. We realise this made it less useful/enjoyable to many users who had no mental health problems, but given the seriousness of the issue, we wanted to get this right”, wrote Altman.
“Now that we have been able to mitigate the serious mental health issues and have new tools, we are going to be able to safely relax the restrictions in most cases.”
Recommended reading
- ChatGPT Introduces Parental Safety Controls – Too Little Too Late?
- ChatGPT Lies and Schemes to Avoid Shut Down, Research Finds
- ChatGPT Under Fire Over Harmful Advice to Teens
That move comes despite OpenAI still dealing with the fallout of a lawsuit filed by the parents of 16-year-old Adam Raine, a US teen who died in April, alleging that his conversations with ChatGPT contributed to his death.
“Given the growing number of cases where people have died after interacting with ChatGPT, we know that their failure has tragic, fatal consequences,” said Imran Ahmed, CEO, the Center for Countering Digital Hate.
“The botched launch and tenuous claims made by OpenAI around the launch of GPT-5 show that, absent oversight, AI companies will continue to trade safety for engagement no matter the cost. How many more lives must be put at risk before OpenAI acts responsibly?”





