Anthropic has admitted that it discovered bad actors using its AI models for malicious activities like cyber-attacks, surveillence, and even developing bioweapons, which it had to block.
The AI firm behind AI chatbot Claude disclosed that its newest AI models are enabling lone actors to create potential cyber-attacks at a formerly unprecedented scale.
An attempt to use Anthropic’s AI models for research that could lead to the development of bioweapons caused the AI firm to add stricter guardrails to its latest model. These would restrict researh into biology that could be leveraged to mae weapons.
Anthropic said that these instances were not necessarily typical, but were the “most notable and novel threat activity we’ve identified to date,” the firm’s report read.
“We’re publishing this work because we believe we have a responsibility to disclose malicious misuse of our services. As models become increasingly capable, their risks will increase, unless AI developers and society’s defenders act to make them safer,” Anthropic said.
The firm released some lines of the malicious code it found as well as the prompts it discovered, urging governments and other AI companies to also search for similar patterns of misuse.
The report detailed a real instance of a user employing Claude to research ways to genetically modify a mosquito-borne virus to make it more harmful, with research questions geared towards its transmission and immune resistance.
Anthropic said that while these questions could be used to develop vaccines and mitigation methods, it could also be used to genetically engineer a more devastating virus, essentially creating a bioweapon.
Interestingly, all the cases – bar one – discussed by Anthropic did not use its newer models – Claude Fable or Mythos – for the malicious purposes. The singular instance where a newer model was used exhibited illici distillation where a bad actor attempted to extract the model’s capabilities to repliate them into another model.
Recommended reading
- OpenAI and Anthropic Called to White House Over AI Hacking Risks
- OpenAI Pauses AI Training Following Security Breaches
- Anthropic Nears IPO Filing as Banks Jostle for Top Roles
- Anthropic and OpenAI Models Implicated in New Cyber Testing Breaches
- “Sloppy” OpenAI Hack Hit More Than Hugging Face
- OpenAI Model Goes Rogue In Testing, Hacks Hugging Face
Anthropic said that its older models were below the threshhold of being capable to truly render harmful biological research, but the discovery is prompting it to intoduce safetyrails to its newer, more capable models.
The AI firm also found instances of users using Claude to create fake social media profiles made to look like real people, that would all share and spread the same political opinion. These appeared to be from users based in a range of regions, from Europe to South Asia and Africa.
The report comes ahead of Anthropic’s much anticipated initial public offering expected this autumn, but also after one of its lead researchers announced his resignation over rising concerns that AI firms are not taking safety and security seriously.





