OpenAI has paused work on its Astra AI model after it passed a 'critical threshold' in cyber capability – but it’s not the one that breached Hugging Face
The firm said it's Astra model can "develop functional zero-day exploits of all severity levels"
OpenAI has paused work on one model over concerns it could launch its own attacks, while at the same time expanding its system for accessing such security models.
The firm said its own "internal evaluations" of next-gen model Astra showed "significant advancements" in security capabilities, leading it to suspect it had breached the "Critical" threshold, used to rate such systems.
That is the first time OpenAI has rated one of its own models as Critical, rather than High.
"Under our Preparedness Framework, a model reaches the Critical cybersecurity threshold if it can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention, or can devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high level desired goal," OpenAI said in a blog post.
The move by OpenAI follows a series of security incidents from AI developers in which their models went rogue, escaping guardrails to attack third parties.
Last month, OpenAI said one of its models had gone rogue and breached Hugging Face. Shortly after, Anthropic piped up with details of a similar incident; last week, Meta said it too had a model capable of such nefarious activity.
As ITPro reported last week, it's since become clear that misconfigurations at one testing firm, Irregular, could be how the models escaped.
Sign up today and you will receive a free copy of our Future Focus 2026 report - the leading resource for IT decision-maker insight on priorities and investment areas in AI, security and more.
Now, OpenAI is looking to tighten up its internal controls on one model and how external organizations access its models for security work. However, it stressed that Astra was not involved in the Hugging Face attack.
In response, OpenAI said it would roll out stricter security around the model and pause "internal activities" using Astra that couldn't meet the stronger controls.
Astra now has "universal monitoring for risky actions and misalignment," with monitors watching over its "chain of thought" in order to stop the model if it starts to move into high risk activity.
OpenAI said it was also working with governments and AI safety organizations to test the Astra model.
OpenAI splits up Daybreak program
OpenAI has also announced plans to expand its Daybreak program, through which companies and governments can access its security models.
The company revealed it will reconfigure access for organizations to better control how its technology is used in the security sector.
"As these capabilities spread, defenders have a narrowing window to prepare," the company said in a blog post. "Our answer is to put frontier intelligence in the hands of trusted defenders everywhere before attackers deploy offensive AI capabilities at scale."
This will now see the Daybreak scheme split into two distinct tiers.
Daybreak Blue offers access to a wider group of organizations but limits that to general frontier models for security tasks such as incident response, malware analysis, and to "support" vulnerability discovery.
Meanwhile, Daybreak Red offers access to security-specific models for flaw hunting and validation. The latter will include a new model, GPT-5.6-Cyber.
Alex Goller, Principal Solution Architect EMEA at Illumio, said the two-tier approach helped solve two issues.
"Daybreak Red restricts models from doing more harm than they should and limits people's exposure to knowing what the models do, while Daybreak Blue fixes the gap that left Hugging Face's responders unable to use frontier models during their incident.”
FOLLOW US ON SOCIAL MEDIA
Follow ITPro on Google News and add us as a preferred source to keep tabs on all our latest news, analysis, views, and reviews.
You can also follow ITPro on LinkedIn, X, Facebook, and BlueSky.
Freelance journalist Nicole Kobie first started writing for ITPro in 2007, with bylines in New Scientist, Wired, PC Pro and many more.
Nicole the author of a book about the history of technology, The Long History of the Future.
-
Logistics firm supply chain breach hits Valve and other customersNews After the hack of Ceva Logistics, customers are being warned to look out for scam emails
-
OVHcloud chief issues customer price warning amid surging hardware costsNews The chief executive warned of 87% price increases for some cloud products, pinning the blame on AI
-
Cyber criminals are selling discount AI tokens on underground forumsNews Sites such as Poison Claude and Ecomagent.in are taking advantage of genuine promo offers and reselling access
-
Hugging Face CEO calls for ‘radical transparency’ in wake of OpenAI attackNews The AI library chief has called for investment to help “build powerful cyber defenses”, as alleged weaknesses in OpenAI’s monitoring emerge
-
An ‘unprecedented cyber incident’: How OpenAI models breached Hugging Face – and why it could herald a ‘new phase of AI-powered cyber crime’News The incident should serve as a stark warning on the dangers of AI agents, according to cyber experts
-
The agents you use to beef up cybersecurity could be turned against you – ‘Friendly Fire’ attacks can manipulate OpenAI and Anthropic models into running malicious codeNews Research shows agents can be fooled into executing malicious code while performing security reviews of third-party software
-
OpenAI expands 'Daybreak' cyber program: New tools, partnerships, and a cyber-focused GPT-5.5 aim to help 'patch the world'News The company has added new tools, signed up partners, and released its GPT-5.5-Cyber model more widely
-
Hackers are capitalizing on AI hype to ramp up social engineering attacks – and they're using big brands like Anthropic, OpenAI, and DeepSeek as ‘bait’ to lure victimsNews Microsoft says cyber criminals are impersonating popular AI platforms to deliver malware
-
Everything you need to know about ChatGPT’s new Advanced Account Security featuresNews OpenAI has introduced new tools to tightening up access to ChatGPT, Codex, and its other AI tools
-
OpenAI is cracking down on AI misuse with a new bug bounty programNews Submissions don't have to be security vulnerabilities, OpenAI says, just the potential to cause material harm