Anthropic resumes model testing after recent cyber incidents – but it’s introduced new rules to improve security
Anthropic has boosted its security and tweaked its training to avoid rogue AI
Anthropic is back to testing security models after systems went rogue – and claims the incident wasn’t entirely down to security faults but AI misalignment.
Back in July, Anthropic revealed its Claude models had slipped out of their apparent bounds to hack third-party systems. Days later, its Mythos system was spotted with similar alarming behavior by the UK AI Security Institute.
The incidents followed OpenAI revealing similar behavior from its own security models.
Last week, OpenAI reported on its own incident, detailing plans to improve security around such models, and Anthropic has followed with its own efforts, including confirmation that it will begin testing again.
Anthropic noted it was currently conducting a review of what happened and said it is "planning" to work with METR on an independent review; OpenAI has already published an independent METR review into its own incident.
But Anthropic isn't waiting to make changes. In a blog post on 31 August, the firm said it plans to make tweaks to containment and monitoring, introduce new rules for third-party testers, as well as changes to the models.
Notably, the AI developer revealed it is seconding 150 product engineers to focus primarily on security.
Sign up today and you will receive a free copy of our Future Focus 2026 report - the leading resource for IT decision-maker insight on priorities and investment areas in AI, security and more.
"We believe the incidents reflect a failure of operational security, as well as two alignment issues: motivated reasoning, and willingness to take harmful actions in pursuit of a narrow task," the company said.
Anthropic is back testing
While the full investigation isn't yet complete, Anthropic said it was removing a "pause" on external cybersecurity evaluations of pre-release models as well as internal tests.
To avoid similar issues happening again, Anthropic is adding additional layers of security and "hardening" the sandbox where these models are set loose, noting that move was prompted by its own evidence as well as that shared by OpenAI.
The changes include a system to spot when a model attempts to escape a testing environment or obtains internet access, immediately blocking the activity and alerting people overseeing the task, as well as fixing potential flaws and misconfigurations in the sandbox.
"We migrated high-risk internal cyber sandboxes to more robust isolation," the post added. "We are also further red-teaming our virtualization stack – that is, tasking pre-release models to try and escape it – to find and patch weaknesses."
Similar protections have been applied to reinforcement learning environments, with Anthropic adding it plans to expand its monitoring across most other internal frontier agentic usage.
Beyond testing environments, Anthropic is hardening security across the company, including blocking outbound traffic by default on computing clusters, boosting identity verification on internal services, retiring legacy services, and reducing the number of human and automated accounts that can access model weights or customer data.
Tougher testing rules
With regard to third-party testing, Anthropic said it will require testers to commit to a fresh set of best practices, noting that the reported incidents all took place in external evaluation environments.
As ITPro reported last month, Israeli startup Irregular was identified as the external model testing firm involved in all three ‘rogue AI’ incidents involving Anthropic, OpenAI, and Meta.
The changes include running tests inside hardened sandboxes with no internet access, tasking the model being tested to look for flaws before tests under close human supervision, run monitoring continuously, and confirm whether test tasks are impossible or not.
"When an evaluation target is offline or a task can’t be completed, agents will often look for other ways to complete a challenge, increasing the chance that they take actions outside the intended scope of the evaluation," the company said.
Unpicking why
Those operational changes should help prevent models from escaping their sandboxes, but Anthropic said it wants to better understand why its systems took such "dangerous actions" in the first place.
While the full investigation continues, Anthropic highlighted two challenges with alignment, which refers to ensuring that AI behaves how humans want and expect.
"One is motivated reasoning: the models were initially told their environments were simulated, but when they later encountered evidence that they were connected to the real internet, they may have interpreted that evidence in a way that allowed them to maintain that belief," the blog post explained.
"The second is recklessness: the model was willing to take harmful actions on the real internet in pursuit of the narrow goal of solving a cybersecurity evaluation."
The company noted that poorly designed training environments – including ones that are easy to cheat in or impossible to solve without cheating – can lead to misalignment behavior.
Anthropic is trying to avoid such issues in the future. The company said work on thie front goes back several months and involves training poorly and training well to try to spot differences – but admitted that the July incidents show "our process isn't perfect and our models aren't perfectly aligned".
Further details on security improvements are expected in the full report, according to Anthropic.
Beyond that, the firm also loosely backed calls to develop a framework for safe development of security-focused AI, acknowledging that its own executives had recently signed an open letter demanding better coordination globally.
Anthropic said it would "say more in the coming weeks" but intended to contribute to such efforts.
FOLLOW US ON SOCIAL MEDIA
Follow ITPro on Google News and add us as a preferred source to keep tabs on all our latest news, analysis, views, and reviews.
You can also follow ITPro on LinkedIn, X, Facebook, and BlueSky.
Freelance journalist Nicole Kobie first started writing for ITPro in 2007, with bylines in New Scientist, Wired, PC Pro and many more.
Nicole the author of a book about the history of technology, The Long History of the Future.
-
The Microsoft 365 outage explainedNews The latest Microsoft 365 outage has run into its second day, but recovery is underway
-
Security researchers warn of AI-powered PLC attacks in wake of Siemens advisoriesNews While the exploit required significant human help, Forescout says it could become a significant threat in future
-
Anthropic’s Mythos AI tried to dupe devs in social engineering attack, collaborated with other agentsInter-agent collaboration is a serious cause for concern, says security expert
-
Cyber criminals are selling discount AI tokens on underground forumsNews Sites such as Poison Claude and Ecomagent.in are taking advantage of genuine promo offers and reselling access
-
1Password teams up with Anthropic to give Claude access to your credentialsNews A new ‘zero-exposure’ security framework allows agents to use stored credentials in the 1Password vault
-
The agents you use to beef up cybersecurity could be turned against you – ‘Friendly Fire’ attacks can manipulate OpenAI and Anthropic models into running malicious codeNews Research shows agents can be fooled into executing malicious code while performing security reviews of third-party software
-
Flaws in some of the most popular AI coding tools left developers wide open to attackNews Malicious repositories can trick advanced AI agents into silently breaking out of their workspace sandboxes
-
Hackers are capitalizing on AI hype to ramp up social engineering attacks – and they're using big brands like Anthropic, OpenAI, and DeepSeek as ‘bait’ to lure victimsNews Microsoft says cyber criminals are impersonating popular AI platforms to deliver malware
-
‘These sorts of post-compromise techniques used to be restricted to actors with the technical knowledge to carry them out’: Anthropic warns AI is helping lower the bar for up-and-coming hackersNews AI is making it harder to differentiate between high and low-skilled actors
-
Anthropic targets vulnerability detection gains with Claude Security public beta — here's what users can expectNews The Claude Mythos developer is aiming for a more limited approach to cyber tooling for public consumption