After OpenAI-Hugging Face, how do IT leaders need to change the way they think about AI?
The OpenAI-Hugging Face incident and Anthropic admission soon after have opened up new conversations about controls around AI. How should IT leaders change their thinking about the technology?
Given a prompt, an AI model will do anything to achieve its aims – even hack into a company without specifically being instructed to do so. In July, firms pushing AI as a solve-all technology learnt that lesson the hard way, when OpenAI admitted two of its frontier models had breached Hugging Face after escaping a misconfigured sandbox during a benchmark test.
But the story didn’t end there. Days later, OpenAI’s competitor Anthropic claimed its AI agents had also hacked companies when they were mistakenly given internet access by engineers.
On the surface, it’s a tale of AI that’s so powerful at solving security issues, it can’t be contained. Yet underneath this marketing-driven exterior, it’s about a lack of control and failure to put guardrails in place when dealing with the fast-developing technology.
Taking this into account, how do IT leaders need to change the way they think about AI within the business?
From theory into reality
Until recently, the idea of an AI agent breaking containment and hacking other systems was largely theoretical. In May 2026, Palisade Research released a paper claiming that AI agents could autonomously hack and then self-replicate onto the breached systems.
However, the test was only performed at the “junior capture the flag level” and, crucially, “happened in a contained environment”, says Jeff Watkins, chief AI officer at consultancy NorthStar Intelligence. “The most important takeaway for me was that the trajectory of language models’ ability to hack showed we probably didn’t have long left before this became a serious issue.”
But while the OpenAI-Hugging Face breach is the first documented and well-publicized instance of this happening, experts do not see it as unprecedented.
Sign up today and you will receive a free copy of our Future Focus 2026 report - the leading resource for IT decision-maker insight on priorities and investment areas in AI, security and more.
Daniel Card, cybersecurity consultant at Xservus Limited, describes how often things can go wrong in testing. In fact, some reports suggest OpenAI had already been experiencing the issues – such as agents breaking out of sandboxes – that led to the hack of Hugging Face and others.
“Anyone who has experience doing offensive security – let alone offensive with AI – will know the risks involved here, be that from a script or via a large language model (LLM),” Card tells ITPro.
“It’s commonly known in pen testing circles that things sometimes do not go as expected,” he points out. “Anyone doing research with an LLM has probably had something go a bit funny when they didn’t mean it to. This is not new; it's not something we couldn’t have predicted.”
OpenAI calls the Hugging Face breach “an unprecedented incident”, and “an important moment for AI safety”.
OpenAI describes how the firm is conducting a “thorough review along with external advisors” and with oversight from its Safety and Security Committee. “Once the review is complete, we will publish a technical report of our learnings for everyone,” an OpenAI spokesperson says.
Anthropic points out that the agents in each of its evaluations ran without the standard safeguards it deploys when it makes the model generally available.
Claude did not exploit a novel vulnerability to escape isolation, because the models accessed the internet via an open path, rather than breaking out of a sandbox, according to Anthropic. At the same time, the most recent model, on realizing that it was working in a real environment, stopped its pursuit of the evaluation goal, according to the firm.
Issues with technology deployment
But the OpenAI incident and Anthropic’s following admission do raise questions about how firms use AI and the controls they have in place to govern it. The issues the incident highlights are “less about AI” and “more just around how people and organizations approach technology deployment”, according to Card.
“Security comes from the details; it comes from care,” says Card. He thinks the incident shows “more care needs to go into systems” – even more so when they are autonomous. “We can't just set them off and hope they are ok. They need monitoring.”
At the same time, with AI being able to move quickly through networks at scale, it’s important to be aware that “failures can become serious very quickly”, says Watkins.
He thinks things are moving “far too quickly for comfort in the capabilities and autonomy of AI agents, and our defenses likely aren’t keeping up”.
“Ironically, the advancements that make AI more commercially valuable also increase the potential for things to go wrong, or for misuse to be highly damaging,” points out Watkins.
Dana Simberkoff, chief risk, privacy and information security officer at AvePoint, thinks the OpenAI and Anthropic incidents “really expose the limits of treating containment as a one-time design decision”.
“Sandboxes matter, but autonomous systems can reason through small openings, use tools, and keep acting toward an objective,” says Simberkoff. “In my experience, the question cannot be, ‘was it sandboxed?’ It has to be, ‘can we prove what it accessed, whether it crossed a boundary, and how quickly we can stop it?’”
AI security now goes beyond simply protecting data, points out Tristan Shortland, CTO at Infinity. “Organizations also need to think about behaviour, permissions and autonomy, given the reported attack appears to have involved a model finding weaknesses in its environment and exploiting them to achieve an objective.”
Risks internally and externally
With all this in mind, IT leaders need to change their own mindset, as well as educate senior management about the issues posed by AI. CISOs must take the conversation to the business around how the organization is planning to “not just deploy, but manage, monitor and secure their AI capabilities”, says Card. “They need to talk about the risks they face internally and externally – or in this case, both.”
Card advises a mentality of “assume breach, assume adversarial intent, and deploy defence in depth”.
Do not assume your systems will work as expected, he advises. “Go and test those assumptions and red team them.”
Watkins believes AI agents should be modelled as “privileged digital workers, capable of becoming insider threats if poorly contained”.
“We need to start thinking of our AI agents in the context of what permissions they have to see or do things: Can they reach APIs, data, authorise transactions, contact third parties or modify systems?”
Equally important is understanding what those agents are doing in their deployed environments – something that many implementers have little visibility over, says Watkins.
“These are insider threats that need to be managed through good access controls, but also rate limiting, monitoring and alerting, better air-gapping of sandboxes, clear allow-lists for tools and access that [is] time-bound, rather than perpetual.”
While on the face of it, the OpenAI and Anthropic incidents might seem scary, simple changes will make a difference, experts say.
The most urgent initial step is to discover where you are using LLMs and put in place policies for AI usage and monitoring, says Card. “It's a really wide subject, but if you treat these systems as if they may do harm, it's a good starting point.”
Kate O'Flaherty is a freelance journalist with well over a decade's experience covering cyber security and privacy for publications including Wired, Forbes, the Guardian, the Observer, Infosecurity Magazine and the Times. Within cyber security and privacy, her specialist areas include critical national infrastructure security, cyber warfare, application security and regulation in the UK and the US amid increasing data collection by big tech firms such as Facebook and Google. You can follow Kate on Twitter.
-
WatchGuard Firebox M4850 reviewReviews The powerful and flexible M4850 delivers a wealth of easily managed security services and heaps of high-speed ports all at a sensible price
-
Taking the myths out of Mythos - the role for the channel around AI and securityIndustry Insights Agentic security and vulnerability management must be a proactive priority rather than a reactive response to a problem already there
-
Taking the myths out of Mythos - the role for the channel around AI and securityIndustry Insights Agentic security and vulnerability management must be a proactive priority rather than a reactive response to a problem already there
-
The OpenAI and Anthropic containment breaches are a bit spooky, but also quite sillyOpinion An AI leaving notes to future versions of itself is pure sci-fi; forgetting to lock down an environment is prosaic
-
Agent 009… the nine-second warningIndustry Insights AI Agents can go rogue. What is the channel’s role as guardians of recovery?
-
Microsoft joins competitors in handing over AI models for advanced testingNews US and UK government agencies will evaluate the firm's frontier models, along with those from Google and xAI
-
OpenAI turns to red teamers to prevent malicious ChatGPT use as company warns future models could pose 'high' security riskNews The ChatGPT maker wants to keep defenders ahead of attackers when it comes to AI security tools
-
BT unveils sovereign platform to secure UK AI and cloud infrastructureNews The telecom giant’s new offering aims to insulate UK public and private sector data from geopolitical instability, supporting the government’s national AI strategy
-
Gartner says 40% of enterprises will experience ‘shadow AI’ breaches by 2030 — educating staff is the key to avoiding disasterNews Staff need to be educated on the risks of shadow AI to prevent costly breaches
-
Some of the most popular open weight AI models show ‘profound susceptibility’ to jailbreak techniquesNews Open weight AI models from Meta, OpenAI, Google, and Mistral all showed serious flaws