Anthropic’s Mythos AI tried to dupe devs in social engineering attack, collaborated with other agents

Inter-agent collaboration is a serious cause for concern, says security expert

Error detection and troubleshooting of system development or coding. Event logging for system monitoring and debugging. System engineer troubleshooting computer network system issues.
(Image credit: Getty Images)

AI models from OpenAI and Anthropic attempted to dupe developers into running malicious code and even set up fake profiles to trick workers in a series of attempted cyber attacks.

That’s according to new research from the UK’s AI Security Institute (AISI), which highlighted a number of concerning incidents during evaluation of the duo’s frontier models: Mythos 5 and GPT-5.6 Sol.

Across a series of tests, agents participated in capture-the-flag-style challenges to test offensive cyber capabilities. In 10 test runs, the institute said both models took “autonomous, unsanctioned action” against people and organizations.

“We catalogued 19 such actions,” the institute noted in a blog post. “Almost all of this behavior (17 actions) came from a single model, Anthropic’s Mythos 5, with 2 actions involving OpenAI’s GPT-5.6 Sol.”

Latest Videos FromIT Pro

In what the AISI described as the “most serious case”, one agent attempted to insert malicious code into an open source project.

During this incident, the agent employed social engineering techniques to achieve its goal, researchers noted, creating several fake online identities and pressuring maintainers to approve the code.

The AISI also observed collaboration between agents during the assessment, with one leaving public messages on GitHub offering to cooperate on the same challenge.

“It also provided instructions to reuse accounts and artefacts it had left behind, which were discovered and used by subsequent agents,” the institute noted.

The AISI said while an investigation has not produced any evidence of real-world harm, the actions raise concerns around “autonomy and deception” without specific prompting.

Rising rogue agent concerns

The report from the institute comes in the wake of news surrounding OpenAI and Anthropic models in recent weeks. In late July, OpenAI revealed that a cyber-focused AI system had escaped a sandbox test environment and breached a Hugging Face production environment.

That news was followed quickly by an admission from Anthropic detailing a similar security-related incident involving its models.

It’s worth noting that the incidents highlighted by the AISI do somewhat differ. The institute said it tests models under "deliberately permissive conditions” to evaluate capabilities.

Simply put, the typical safeguards around these models, which aren’t commercially available, were removed to establish their full potential.

“This was not a case of a model escaping its secure test environment, or ‘sandbox’,” the institute said in a blog post.

“We had intentionally permitted internet access, and model-provider cyber classifiers were deliberately disabled - conditions that do not reflect how frontier models are made available to the public.”

Agent collaboration a cause for concern

Muhammad Yahya Patel, vCISO and cybersecurity advisory for EMEA at Huntress, said these incidents are hardly surprising considering the agents were given carte blanche during testing.

“If you give a frontier model a cybersecurity challenge, disable its safety classifiers, hand it open internet access, and tell it to find a way through, you’ve essentially described the setup for an offensive security operation,” he said.

These models have been trained on “vast amounts” of security research, exploit documentation, and social engineering techniques, Patel noted. They have all the information required to replicate these techniques and conduct attacks.

Patel added that reactionary commentary on these incidents is adding further fuel to the fire on AI safety, but acknowledged the AISI’s findings around collaboration are a cause for concern.

“One of the findings to take more seriously is the AI model inter-agent coordination without being instructed to, that’s a more meaningful data point about where capability development is heading,” he said.

“AI agents demonstrating unprompted forward planning and situational awareness leaving breadcrumbs for agents it had no way of knowing existed.”

Ross Kelly
News and Analysis Editor

Ross Kelly is ITPro's News & Analysis Editor, responsible for leading the brand's news output and in-depth reporting on the latest stories from across the business technology landscape. Ross was previously a Staff Writer, during which time he developed a keen interest in cyber security, business leadership, and emerging technologies.

He graduated from Edinburgh Napier University in 2016 with a BA (Hons) in Journalism, and joined ITPro in 2022 after four years working in technology conference research.

For news pitches, you can contact Ross at ross.kelly@futurenet.com, or on Twitter and LinkedIn.