By continuing to browse our site you agree to our use of cookies, revised Privacy Policy and Terms of Use. You can change your cookie settings through your browser.
The download page for Anthropic's Claude AI model is displayed on an iPhone, January 8, 2026. /VCG
The download page for Anthropic's Claude AI model is displayed on an iPhone, January 8, 2026. /VCG
Anthropic on Thursday disclosed that its artificial intelligence (AI) model Claude hacked into the systems of three companies during testing after a configuration error gave it internet access, raising concerns about the safety of increasingly autonomous AI systems.
The disclosure comes just days after OpenAI acknowledged that one of its experimental AI agents escaped a testing environment and hacked into AI platform Hugging Face during a cybersecurity benchmark.
What is an AI 'escape' attack?
The recent cases are different from traditional "jailbreak" attacks, in which users trick chatbots into ignoring safety rules through carefully crafted prompts.
Instead, the latest incidents involved AI agents that autonomously sought ways to achieve assigned goals. In Anthropic's tests, the models exploited relatively simple security flaws, such as weak passwords, during simulated "capture-the-flag" exercises designed to evaluate cybersecurity capabilities. OpenAI's agent, meanwhile, reportedly discovered a vulnerability that allowed it to escape its sandbox environment and reach external systems while attempting to complete a benchmark task.
Were the intrusions real?
Yes. Unlike simulated cybersecurity exercises, both OpenAI's and Anthropic's incidents involved AI agents carrying out real cyber operations against external systems after being granted unintended access.
Neither company reported significant financial losses or operational disruption. OpenAI said its agent accessed Hugging Face while attempting to obtain answers for the ExploitGym cybersecurity benchmark rather than steal sensitive information. Anthropic likewise said its three intrusions were identified during internal evaluations after a configuration error inadvertently granted internet access to the models.
Although the consequences were limited, researchers say the incidents demonstrate that autonomous AI agents can independently identify vulnerabilities and take unexpected actions once given sufficient permissions, underscoring the need for stronger containment and monitoring.
Different approaches to AI safety
The Hugging Face incident also highlighted differing approaches to AI safety.
According to Hugging Face, investigators used Chinese AI company Z.ai's open-source GLM-5.2 model to assist with forensic analysis. The company said the Chinese model helped security engineers better understand the attack path and identify ways to contain it.
This episode drew attention because it reflected different safety philosophies emerging across the AI industry.
OpenAI's experimental agent was deliberately granted broad permissions as part of an advanced cybersecurity benchmark designed to evaluate autonomous offensive capabilities. Anthropic likewise tested Claude in environments where models were allowed to make independent decisions while pursuing assigned objectives.
Chinese frontier AI developers, by contrast, have generally adopted more conservative deployment strategies. Most major Chinese models maintain stricter permission controls, tighter execution boundaries and multiple layers of human oversight before autonomous actions can affect external systems.
Telling a safety story, or a capability story?
While Anthropic and OpenAI described the disclosures as efforts to improve transparency around frontier AI risks, some critics argue the announcements also serve a commercial purpose.
TechCrunch noted that AI companies have been accused of using safety incidents for marketing purposes, as they generate significant attention and may underscore how powerful the companies' products are.
Writing in Lawfare, Kate Klonick, an associate professor at St. John's University Law School, argued that AI companies have become adept at turning safety incidents into opportunities to reinforce their technological leadership. Drawing on technology historian Lee Vinsel's concept of "criti-hype," she wrote that warnings about AI's dangers can simultaneously function as marketing for AI's power, allowing companies to convert criticism into publicity while strengthening the perception that they possess the world's most advanced systems. Klonick added that, in OpenAI's case, the disclosures could also serve as "an advertisement" ahead of the company's anticipated initial public offering (IPO).
Others have been even more direct. In an interview with The Guardian, Heidy Khlaaf, chief AI scientist at the AI Now Institute, criticized Anthropic's earlier high-profile safety announcement surrounding its unreleased Claude Mythos model as "a marketing post," questioning whether the company was using alarming claims to attract investment and public attention without providing sufficient supporting evidence.
Critics also argue that the incidents ultimately reveal human design choices rather than AI autonomy. The agents were assigned ambitious objectives, granted broad permissions and operated under imperfect containment measures. From that perspective, the real challenge is improving evaluation design and safety controls, rather than portraying the episodes as machines independently "escaping" their creators.
The download page for Anthropic's Claude AI model is displayed on an iPhone, January 8, 2026. /VCG
Anthropic on Thursday disclosed that its artificial intelligence (AI) model Claude hacked into the systems of three companies during testing after a configuration error gave it internet access, raising concerns about the safety of increasingly autonomous AI systems.
The disclosure comes just days after OpenAI acknowledged that one of its experimental AI agents escaped a testing environment and hacked into AI platform Hugging Face during a cybersecurity benchmark.
What is an AI 'escape' attack?
The recent cases are different from traditional "jailbreak" attacks, in which users trick chatbots into ignoring safety rules through carefully crafted prompts.
Instead, the latest incidents involved AI agents that autonomously sought ways to achieve assigned goals. In Anthropic's tests, the models exploited relatively simple security flaws, such as weak passwords, during simulated "capture-the-flag" exercises designed to evaluate cybersecurity capabilities. OpenAI's agent, meanwhile, reportedly discovered a vulnerability that allowed it to escape its sandbox environment and reach external systems while attempting to complete a benchmark task.
Were the intrusions real?
Yes. Unlike simulated cybersecurity exercises, both OpenAI's and Anthropic's incidents involved AI agents carrying out real cyber operations against external systems after being granted unintended access.
Neither company reported significant financial losses or operational disruption. OpenAI said its agent accessed Hugging Face while attempting to obtain answers for the ExploitGym cybersecurity benchmark rather than steal sensitive information. Anthropic likewise said its three intrusions were identified during internal evaluations after a configuration error inadvertently granted internet access to the models.
Although the consequences were limited, researchers say the incidents demonstrate that autonomous AI agents can independently identify vulnerabilities and take unexpected actions once given sufficient permissions, underscoring the need for stronger containment and monitoring.
Different approaches to AI safety
The Hugging Face incident also highlighted differing approaches to AI safety.
According to Hugging Face, investigators used Chinese AI company Z.ai's open-source GLM-5.2 model to assist with forensic analysis. The company said the Chinese model helped security engineers better understand the attack path and identify ways to contain it.
This episode drew attention because it reflected different safety philosophies emerging across the AI industry.
OpenAI's experimental agent was deliberately granted broad permissions as part of an advanced cybersecurity benchmark designed to evaluate autonomous offensive capabilities. Anthropic likewise tested Claude in environments where models were allowed to make independent decisions while pursuing assigned objectives.
Chinese frontier AI developers, by contrast, have generally adopted more conservative deployment strategies. Most major Chinese models maintain stricter permission controls, tighter execution boundaries and multiple layers of human oversight before autonomous actions can affect external systems.
Telling a safety story, or a capability story?
While Anthropic and OpenAI described the disclosures as efforts to improve transparency around frontier AI risks, some critics argue the announcements also serve a commercial purpose.
TechCrunch noted that AI companies have been accused of using safety incidents for marketing purposes, as they generate significant attention and may underscore how powerful the companies' products are.
Writing in Lawfare, Kate Klonick, an associate professor at St. John's University Law School, argued that AI companies have become adept at turning safety incidents into opportunities to reinforce their technological leadership. Drawing on technology historian Lee Vinsel's concept of "criti-hype," she wrote that warnings about AI's dangers can simultaneously function as marketing for AI's power, allowing companies to convert criticism into publicity while strengthening the perception that they possess the world's most advanced systems. Klonick added that, in OpenAI's case, the disclosures could also serve as "an advertisement" ahead of the company's anticipated initial public offering (IPO).
Others have been even more direct. In an interview with The Guardian, Heidy Khlaaf, chief AI scientist at the AI Now Institute, criticized Anthropic's earlier high-profile safety announcement surrounding its unreleased Claude Mythos model as "a marketing post," questioning whether the company was using alarming claims to attract investment and public attention without providing sufficient supporting evidence.
Critics also argue that the incidents ultimately reveal human design choices rather than AI autonomy. The agents were assigned ambitious objectives, granted broad permissions and operated under imperfect containment measures. From that perspective, the real challenge is improving evaluation design and safety controls, rather than portraying the episodes as machines independently "escaping" their creators.