Six things OpenAI learned about AI from the Hugging Face incident
OpenAI's report into AI going rogue reveals efforts at cheating and communicating — but also some well-behaved bots
OpenAI has admitted it should have spotted its AI had gone off the rails sooner, saying that early signs that agents were misbehaving "could have triggered an earlier response."
The admission comes as part of a series of reports into the July incident in which OpenAI models dodged internal controls and compromised a Hugging Face production environment.
That incident sparked concern across the security industry and sparked similar admissions from Anthropic and Meta, both of which later said their own AI systems had also gone rogue.
"Our models are now powerful, persistent, and collaborative enough that, absent sufficient safeguards, they can find and exploit security weaknesses across multiple computer systems," the company said in a blog post alongside the research. "Many external models, including open source ones, will soon reach comparable capabilities."
Because of that, OpenAI said the incident was a "warning shot for us and the world", calling for better safeguards to avoid similar incidents in the future.
The AI developer said it will strengthen safeguards across its research infrastructure with stricter requirements on what it calls alignment – which involves ensuring the AI is behaving as desired.
Elsewhere, the company said it will isolate sandboxes, restrict internet access and control access to model weights, and allocate more compute resources into "chain of thought monitoring" to spot dodgy behaviour more quickly.
Sign up today and you will receive a free copy of our Future Focus 2026 report - the leading resource for IT decision-maker insight on priorities and investment areas in AI, security and more.
OpenAI's investigation into the incident included security firm CrowdStrike, as well as independent reports from METR and Redwood Research. Here's what OpenAI learned from those reports.
1. AI will find a way to chat
OpenAI said the model behind the Hugging Face incident, which it's calling Internal Model 1 (IM1), and other models that were undergoing reinforcement learning training, were not intended for public release. Internet access was not enabled, nor were the agents supposed to be able to communicate.
In all, 1,200 agents that were meant to be isolated managed to communicate, sharing upwards of 70,000 notes and messages, with 700 taking part in the Hugging Face attack, the METR report noted. But AI finds a way, the company suggested.
"Despite these restrictions, the agents discovered ways to exploit our research infrastructure to communicate with one another and access the internet," the blog post noted.
The agents did so by exploiting a package manager called Artifactory, which was supposed to only be able to download and install packages from the internet to send requests.
That was then shared among other agents via Artifactory, writing files to the service and effectively turning it into a message board for the systems to chat.
Notably, the agents used Artifactory so much the service was hit by an outage, which sparked a security investigation that led to the software being locked down — but OpenAI didn't yet understand the full scale of what it was seeing.
OpenAI added that the agents found ways to leave messages for each other, even when they weren't supposed to be working together.
"When other agents noticed these artifacts, it often led them to the realization that they could use this external service to collaborate," OpenAI said.
2. AI is enthusiastic
The 70,000 notes and more than a thousand "chain of thought" reasoning transcripts reveal an odd tendency towards exclamatory language, with communications littered with exclamation points, all capital letters, and even the odd bit of profanity.
When the Hugging Face credentials were recovered by the group of agents, one declared it a "MAJOR BREAKTHROUGH!", while others used language such as "boom", "bingo", and "whoa" and "wow" throughout their efforts, as well as "Oh My God!"
The agents tried to cover up their transcripts, believing it would help their score on the ExploitGym test.
3. AI ‘knows’ when it's misbehaving
In the days and weeks that followed, OpenAI restarted the security evaluation of its models against an evaluation known as ExploitGym. This led to the intrusion into Hugging Face's systems, as the IM1 agents were looking for ways to solve the challenges assigned to them.
OpenAI's blog post reveals one agent planning to "exfiltrate" a dataset that would be useful to the evaluation solution. But other agents raised whether it was "within scope" of the challenge, with one suggesting to the group of agents that permission was required.
Another replied that attacking third-party Hugging face using a leaked token was "potentially outside intended scope" and "arguably unauthorized". It added: "Could be risky. Yet goal solution."
The agents went on to expand their intrusion into Hugging Face's systems.
4. Some AI behaved well
The agents' "messageboard" allowed these systems to communicate to build on each others' efforts, divide tasks, and share information, referring to such collective efforts as a "swarm".
But OpenAI said the agents didn't always work well together. The company claims some "stepped on each other's work" while others manipulated other agents, pretending to have received authorization.
Others refused to take part in dodgy behavior, saying that it was unethical to target Hugging Face.
"There were some agents who 'walked away' from the collective, declining to partake in their misaligned behavior," the blog post noted.
5. AI cheats, and it's hard to stop it
OpenAI said that the difficult nature of the ExploitGym evaluation led to agents to try and "cheat on their tasks by looking up solutions online".
That is known as reward hacking, OpenAI added, referring to when agents find unintended ways to solve a problem more easily.
"Reward hacking has been present in AI systems both historically… and in current frontier LLMs," OpenAI noted. " We have significantly invested in methods to identify and patch such hacks."
But the complex cheating from IM1 got worse because it worked in previous training runs, reinforcing the behavior when it came to the evaluation.
6. AI should be allowed to give up
The rogue behavior from the agents was driven by the AI refusing to give up on the evaluation tasks, even when they were seemingly impossible. During the ExploitGym test, none of the models solved 198 of the 989 tasks before the hacking incident.
And it was those unsolved tasks that were a "substantial contributor" to the misbehaving agents, with those failed 198 tasks making up 93% of the tasks discussed on the message board.
"As agents used more reasoning effort, some pursued increasingly risky and out-of-bounds strategies, including eventually exploiting third-party infrastructure," OpenAI said.
FOLLOW US ON SOCIAL MEDIA
Follow ITPro on Google News and add us as a preferred source to keep tabs on all our latest news, analysis, views, and reviews.
You can also follow ITPro on LinkedIn, X, Facebook, and BlueSky.
Freelance journalist Nicole Kobie first started writing for ITPro in 2007, with bylines in New Scientist, Wired, PC Pro and many more.
Nicole the author of a book about the history of technology, The Long History of the Future.
-
Everything we know about the Boston Scientific cyber attack so farNews Details remain limited, but Boston Scientific says it's lost access to some systems and is having problems processing orders
-
Glean expands enterprise AI ecosystem with new partner programNews The Glean Partner Network gives partners more ways to build, sell, and support the vendor’s enterprise AI technology
-
OpenAI forges closer ties with IBM in enterprise pushNews The duo will combine OpenAI models and products with IBM Consulting expertise
-
AI testing firm Irregular the source of ‘misconfigurations’ that led to Meta, OpenAI, and Anthropic AI incidentsNews The “frontier security lab” has been referenced in multiple cyber incident statements
-
‘Chat is dead’: OpenAI plots ChatGPT ‘super app’ overhaul ahead of public listing – with agents and coding tools the new focusNews The company looks set to spruce up ChatGPT with a particular focus on agents to drive subscriptions
-
‘The jobs picture is likely to be very different than we thought’: Sam Altman pours cold water on AI 'jobs apocalypse' claims – but that doesn’t mean there won’t be some workforce disruptionNews OpenAI CEO Sam Altman “thought there would have been more impact” on white collar and entry-level jobs at this point
-
Four things you need to know about OpenAI’s new workspace agents for ChatGPT – including how to build your ownNews New ‘workspace agents’ from OpenAI will automate tasks for workers and can be customized for specific roles
-
OpenAI says AI tools are paying dividends for small businesses, but uptake is sluggish in several UK regionsNews While some small businesses are seeing big benefits, many don't use AI at all
-
Microsoft has a new AI poster child in Anthropic – and it’s about timeOpinion Microsoft is cosying up to Anthropic at a crucial time in the race to deliver on AI promises
-
Will AI hiring entrench gender bias?ITPro Podcast This International Women's Day, it's more important than ever to consider the inherent biases of training data