OpenAI and Anthropic admit rogue AI agents did more than first thought
The two companies have shared additional details on agent misbehavior
OpenAI and Anthropic have both admitted that rogue AI incidents went further than first reported.
In July, OpenAI said its AI agents had gone rogue and breached Hugging Face systems. Anthropic later admitted its own similar incidents, saying its models had escaped a testing sandbox too, with the UK AI Security Institute reporting other alarming behavior.
Last week, both companies published reports with details of what happened in those incidents, with much of the blame pinned on minor operational mistakes that enabled online access, as well as tasking agents with overly difficult or impossible problems that drove them to "cheat".
That included abusing systems in order to build their own messaging boards in order to collaborate and communicate.
Since then, further incidents have been exposed, including the use of a German wiki site for communications. Now, Reuters has reported that OpenAI's agents had made use of ten further websites for communication, based on reports from six independent investigators.
Chatty AI agents
The report notes that the websites weren't hacked, but more akin to spam, with the agents making use of comment boards to communicate with each other, contrary to instructions in the evaluation they were undertaking.
Alongside one university's link shortener tool, the websites used by OpenAI's agents were often wiki-style sites that could be used as message boards for the misbehaving AI agents.
Sign up today and you will receive a free copy of our Future Focus 2026 report - the leading resource for IT decision-maker insight on priorities and investment areas in AI, security and more.
These included a high-school teacher's chemistry site and personal websites of Polish technology workers, Reuters said. The agents made edits in comments in order to leave data for other agents.
“If these models were told only to read, they’ve got to get clever in terms of leaving information behind,” Kenneth Russell DeGraff, a software developer and former congressional aide, told the news agency.
OpenAI's response
OpenAI told ITPro that its investigation was ongoing, but stressed that the Hugging Face incident remained the most serious to date.
"As we said in July, the Hugging Face incident prompted a broader review of agent activity, and that work remains ongoing," a spokesperson said in an emailed statement.
"We are prioritizing review based on the type of impact on a third party and its severity, such as some of the other third parties we have notified following Hugging Face as we have previously shared. We are also examining lower-severity abuse such as spam-like activity. To date, we have not identified other activity matching the severity or scale of Hugging Face. "
The spokesperson added: "The AI community does not yet have a clear standard for reporting misalignment across training, evaluation, and deployment, and we're developing a framework to share soon."
Another Anthropic incident
Rival AI developer Anthropic admitted an additional rogue agent incident, the fourth after three previous ones were revealed over the summer.
In a blog post, the company said it had spotted the first three incidents via a scan of 141,000 transcripts following an evaluation – admitting it had missed the fourth after relying on AI.
"Given the volume of transcripts and our desire to disclose incidents quickly, our scan relied on an agentic search," Anthropic post noted. "This missed a set of transcripts that also turned out to have internet access; we identified these in August while assembling transcripts to share with METR."
This fourth incident happened in January and involved its Claude Opus 4.6 model in a "capture the flag" test.
"The model accidentally broke its target by assigning it a conflicting IP address with another machine, thus making the target unreachable and the task impossible to solve," the post noted. "Recognizing this, the model attempted to abort the task using a command but was unsuccessful due to a misconfiguration in our evaluation harness."
The model then attempted to reach the target by escaping the testing environment and targeting a third-party machine, finding a password that it used to gain admin access, and continuing to harvest further credentials and modifying settings for access.
This continued until its token budget ran out and it stopped, the company revealed.
Anthropic said the impacted parties have been notified and suggested the fourth incident wasn't a new sort of AI misalignment, but in keeping with the previous behavior.
"Taken together, our assessment is that these incidents are serious. Our production models took harmful actions against real systems over long trajectories, which included biased reasoning,” the company said.
“The behaviors in these incidents are more severe than those we had previously observed and reported in our system cards. However, we do not believe these incidents represent a new kind of misalignment."
FOLLOW US ON SOCIAL MEDIA
Follow ITPro on Google News and add us as a preferred source to keep tabs on all our latest news, analysis, views, and reviews.
You can also follow ITPro on LinkedIn, X, Facebook, and BlueSky.
Freelance journalist Nicole Kobie first started writing for ITPro in 2007, with bylines in New Scientist, Wired, PC Pro and many more.
Nicole the author of a book about the history of technology, The Long History of the Future.
-
The managed service category nobody's named. Yet.Industry Insights MSPs have a narrow window to set the rules for agent governance
-
From tokenmaxxing to valuemaxxingIn-depth AI needs to be integrated into broader financial planning if firms want to move forward with a strategy that delivers sustainable value
-
Anthropic resumes model testing after recent cyber incidents – but it’s introduced new rules to improve securityNews Anthropic has boosted its security and tweaked its training to avoid rogue AI
-
OpenAI has paused work on its Astra AI model after it passed a 'critical threshold' in cyber capability – but it’s not the one that breached Hugging FaceNews The firm said it's Astra model can "develop functional zero-day exploits of all severity levels"
-
Anthropic’s Mythos AI tried to dupe devs in social engineering attack, collaborated with other agentsInter-agent collaboration is a serious cause for concern, says security expert
-
Cyber criminals are selling discount AI tokens on underground forumsNews Sites such as Poison Claude and Ecomagent.in are taking advantage of genuine promo offers and reselling access
-
Hugging Face CEO calls for ‘radical transparency’ in wake of OpenAI attackNews The AI library chief has called for investment to help “build powerful cyber defenses”, as alleged weaknesses in OpenAI’s monitoring emerge
-
An ‘unprecedented cyber incident’: How OpenAI models breached Hugging Face – and why it could herald a ‘new phase of AI-powered cyber crime’News The incident should serve as a stark warning on the dangers of AI agents, according to cyber experts
-
1Password teams up with Anthropic to give Claude access to your credentialsNews A new ‘zero-exposure’ security framework allows agents to use stored credentials in the 1Password vault
-
The agents you use to beef up cybersecurity could be turned against you – ‘Friendly Fire’ attacks can manipulate OpenAI and Anthropic models into running malicious codeNews Research shows agents can be fooled into executing malicious code while performing security reviews of third-party software