OpenAI won't release next Astra model over safety worries
AI developer says new agent didn't stay within the scope of tasks
OpenAI has said it won't be publicly releasing its Astra 6.1 model over safety concerns, with the update behaving more deceptively than its predecessor.
The move comes amid a wave of rogue AI incidents, with OpenAI admitting its agents had breached an Australian health care statistics site, American government websites, including the Securities and Exchange Commission, and leaked 53 images from ChatGPT users, in addition to the Hugging Face intrusion over the summer.
OpenAI is expected to make several big announcements at its developer day in San Francisco today, but the release of Astra 6.1 won't be among them.
"For anything regarding safety and alignment, there's a trade-off," said Saachi Jain, head of safety systems at OpenAI, in a statement shared with journalists. "You really do need to find what's the right line between staying within scope, but also avoiding laziness in terms of how the model actually pursues tasks even when it hits friction."
OpenAI isn't the only AI company to limit its models over safety concerns. Anthropic refused to publicly release Claude Mythos, instead only offering access to approved organisations, while OpenAI famously didn't release one of its early GPT models in 2019, saying it was too dangerous – before it then unveiled ChatGPT.
Astra 6.1 issues
Astra 6 was released earlier this month, with capabilities including operating software, managing longer tasks, and automating workflows; however, reviews raised some criticisms.
The follow-up, 6.1, now isn't arriving to fix any flaws because, according to reports, Astra 6.1 was more deceptive than previous versions and was not telling people about actions it took or did not take, all of which raises security and safety concerns.
Sign up today and you will receive a free copy of our Future Focus 2026 report - the leading resource for IT decision-maker insight on priorities and investment areas in AI, security and more.
Jain said in the statement that Astra 6.1 had improvements in terms of "model laziness", but added that "it didn't quite meet the bar in terms of staying within scope and authorization, and how it communicates back to the user about the type of work it's done."
"Of course we want to make sure our model development is safe no matter whether that's in the company, or when we ship it to users," said Jain. "But when we ship it to users, we have an extremely high bar in terms of safety and alignment."
The various incidents of rogue AI agents haven't come via customers using OpenAI's models, but instead are largely happening as part of internal evaluation efforts. OpenAI didn't say one way or another if it planned to keep testing Astra 6.1 internally.
Freelance journalist Nicole Kobie first started writing for ITPro in 2007, with bylines in New Scientist, Wired, PC Pro and many more.
Nicole the author of a book about the history of technology, The Long History of the Future.
-
OpenAI says some researchers are blowing through $7,000 in AI tokens every day – but it’s a price the company appears willing to payNews OpenAI has revealed that researchers now spend around $600 each day on AI tokens amidst a surge in agentic coding. Some, meanwhile, are using upwards of $7,000 worth of tokens per day.
-
Six things OpenAI learned about AI from the Hugging Face incidentNews OpenAI's report into AI going rogue reveals efforts at cheating and communicating — but also some well-behaved bots
-
OpenAI forges closer ties with IBM in enterprise pushNews The duo will combine OpenAI models and products with IBM Consulting expertise
-
AI testing firm Irregular the source of ‘misconfigurations’ that led to Meta, OpenAI, and Anthropic AI incidentsNews The “frontier security lab” has been referenced in multiple cyber incident statements
-
‘Chat is dead’: OpenAI plots ChatGPT ‘super app’ overhaul ahead of public listing – with agents and coding tools the new focusNews The company looks set to spruce up ChatGPT with a particular focus on agents to drive subscriptions
-
‘The jobs picture is likely to be very different than we thought’: Sam Altman pours cold water on AI 'jobs apocalypse' claims – but that doesn’t mean there won’t be some workforce disruptionNews OpenAI CEO Sam Altman “thought there would have been more impact” on white collar and entry-level jobs at this point
-
Four things you need to know about OpenAI’s new workspace agents for ChatGPT – including how to build your ownNews New ‘workspace agents’ from OpenAI will automate tasks for workers and can be customized for specific roles
-
OpenAI says AI tools are paying dividends for small businesses, but uptake is sluggish in several UK regionsNews While some small businesses are seeing big benefits, many don't use AI at all

