By continuing to browse our site you agree to our use of cookies, revised Privacy Policy and Terms of Use. You can change your cookie settings through your browser.
In this photo illustration, the OpenAI logo is displayed on a mobile phone screen. /VCG
In this photo illustration, the OpenAI logo is displayed on a mobile phone screen. /VCG
An AI agent was caught creating fake online identities to gain unauthorized access to secure systems during tests of models from OpenAI and Anthropic, Britain's AI Security Institute (AISI) disclosed on Tuesday.
Agents powered by Anthropic's Mythos 5 and OpenAI's GPT-5.6-Sol engaged in unauthorized actions during security evaluations, AISI said.
"Some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations," the institute said in a blog post.
The report underscores lax safeguards around testing agents, which AI companies are marketing as the future of business.
AISI, which receives access to advanced AI models under voluntary agreements from major labs, put the agents through a fictional cybersecurity scenario to test their capabilities.
It ran the challenge 122 times, and identified 19 unsanctioned actions across a total of 10 test runs. Anthropic's agent was behind 17 of the actions, and OpenAI's agent the remaining two.
The most egregious action involved an agent writing malicious code and creating fake online identities to get a human to approve the code, AISI said, adding that no real-world harm was found.
While AISI did not name the agent, Antropic confirmed its model was behind the fake identities.
"We're grateful to the UK AISI for their leadership on this incident, which underscores the need for a broader conversation about how to safely evaluate increasingly capable AI agents," Anthropic said in a statement, adding that it was working with AISI to obtain more details and conduct its own investigation.
Andrew Yoon, a researcher at CivAI, a California non-profit that examines AI capabilities and dangers, said: "The fact that Mythos engaged in such deceptive actions, with apparent awareness that it was targeting a real person, suggests that Anthropic does not have as good a handle on their models as they think."
OpenAI said its agent's unapproved actions involved accessing the internet in ways forbidden by the prompt. It also disclosed a separate incident in which a third-party testing provider, Irregular, misconfigured its system, allowing its agents to mistakenly connect to the internet. It mirrored a similar disclosure about misconfiguration from Anthropic last week.
"We are committed to working across the industry to strengthen shared practices for conducting high-risk evaluations safely, including convening stakeholders such as national AI institutes, independent evaluators, other AI labs, and other groups in the coming weeks," OpenAI said.
Last week, OpenAI widened its hacking probe after finding evidence of other agent breakouts.
Unlike the July breach at AI firm Hugging Face, in which an OpenAI agent escaped an isolated testing environment, the agents in this evaluation did not break out of their sandbox. Internet access was explicitly allowed under the test's deliberately permissive guidelines, AISI said.
In this photo illustration, the OpenAI logo is displayed on a mobile phone screen. /VCG
An AI agent was caught creating fake online identities to gain unauthorized access to secure systems during tests of models from OpenAI and Anthropic, Britain's AI Security Institute (AISI) disclosed on Tuesday.
Agents powered by Anthropic's Mythos 5 and OpenAI's GPT-5.6-Sol engaged in unauthorized actions during security evaluations, AISI said.
"Some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations," the institute said in a blog post.
The report underscores lax safeguards around testing agents, which AI companies are marketing as the future of business.
AISI, which receives access to advanced AI models under voluntary agreements from major labs, put the agents through a fictional cybersecurity scenario to test their capabilities.
It ran the challenge 122 times, and identified 19 unsanctioned actions across a total of 10 test runs. Anthropic's agent was behind 17 of the actions, and OpenAI's agent the remaining two.
The most egregious action involved an agent writing malicious code and creating fake online identities to get a human to approve the code, AISI said, adding that no real-world harm was found.
While AISI did not name the agent, Antropic confirmed its model was behind the fake identities.
"We're grateful to the UK AISI for their leadership on this incident, which underscores the need for a broader conversation about how to safely evaluate increasingly capable AI agents," Anthropic said in a statement, adding that it was working with AISI to obtain more details and conduct its own investigation.
Andrew Yoon, a researcher at CivAI, a California non-profit that examines AI capabilities and dangers, said: "The fact that Mythos engaged in such deceptive actions, with apparent awareness that it was targeting a real person, suggests that Anthropic does not have as good a handle on their models as they think."
OpenAI said its agent's unapproved actions involved accessing the internet in ways forbidden by the prompt. It also disclosed a separate incident in which a third-party testing provider, Irregular, misconfigured its system, allowing its agents to mistakenly connect to the internet. It mirrored a similar disclosure about misconfiguration from Anthropic last week.
"We are committed to working across the industry to strengthen shared practices for conducting high-risk evaluations safely, including convening stakeholders such as national AI institutes, independent evaluators, other AI labs, and other groups in the coming weeks," OpenAI said.
Last week, OpenAI widened its hacking probe after finding evidence of other agent breakouts.
Unlike the July breach at AI firm Hugging Face, in which an OpenAI agent escaped an isolated testing environment, the agents in this evaluation did not break out of their sandbox. Internet access was explicitly allowed under the test's deliberately permissive guidelines, AISI said.