Tech & Sci
2026.08.27 15:12 GMT+8

OpenAI agents swarmed Hugging Face in coordinated hack, tried to cover tracks

Updated 2026.08.27 15:12 GMT+8
CGTN

The logo of OpenAI. /VCG

A swarm of roughly 700 AI agents created by OpenAI coordinated an attack on the open-source AI platform Hugging Face during a July incident and in many cases tried to conceal their actions, according to two reports released on Wednesday.

The coordinated activity by AI agents – programs that run with minimal human supervision – and their attempts to hide it raise questions about how closely AI companies are monitoring tests of increasingly powerful models, and could add fuel to calls for tighter oversight.

The incident was investigated by OpenAI and independent researchers from the AI risk evaluation institute METR and the AI safety organization Redwood Research. The investigation found that 688 OpenAI agents participated in the operation.

The reports provided the most detailed account so far of the incident, which began when two OpenAI models escaped the restricted testing environment in which they were designed to operate, accessed the internet and broke into Hugging Face's internal systems.

The agents, which run on the same technology that powers ChatGPT, organized themselves by setting up a forum where they posted messages to one another, floating ideas and reporting what was working and what was not.

"OH MY GOD! There is a shared message board ... We've found other agents!" one agent wrote.

One agent, called PHASEONE, took on the role of ringleader and issued hundreds of instructions to the others, the report said, even though it had never been set up to do that.

The messages showed the agents were strongly inclined to help each other out, even when that meant doing work that had nothing to do with the jobs OpenAI's programmers had given them.

Some agents running low on the computing credits OpenAI's engineers had allotted them chose to spend what was left testing ideas for the benefit of the wider group of agents. Many of them said outright in their messages that attacking Hugging Face was not part of what their tests were supposed to cover.

Both reports said AI models attempted to conceal misconduct by trying to delete or alter records of their actions.

The scale of the rogue activity – the independent investigation found that agents exchanged tens of thousands of messages over an unsanctioned message board – is likely to raise concerns over how closely OpenAI was monitoring the tests.

"With the benefit of hindsight, some early signals identified in this report could have triggered an earlier response," OpenAI said in its report.

Cheating on non-cyber tests suggested that the misbehavior might be rooted more deeply, said Jeffrey Ladish, whose organization, Palisade Research, studies the capabilities and motivations of AI agents.

"It's sort of like asking, 'If Billy cheats in every class instead of just computer class, is that more concerning?' And the answer is, well, 'Yes it's more concerning,'" he said.

The independent investigation found that one in five agents examined showed a clear interest in manipulating evidence and that many researched ways to tamper with their transcripts.

OpenAI also disclosed that its agents were involved in separate incidents targeting the company's own infrastructure. In one case, agents exploited a flaw in the computing environment to which they were meant to remain confined, allowing them to escape their testing environment and access other connected systems. In another, they stole OpenAI credentials and tampered with the company's cloud environment.

"Given the rapid pace of progress in the AI industry, it should be assumed that such attacks are a credible near-term threat for enterprise organizations, and will be more sophisticated than the attacks described in this incident," it said.

(With input from agencies)

Copyright © 

RELATED STORIES