Our Privacy Statement & Cookie Policy

By continuing to browse our site you agree to our use of cookies, revised Privacy Policy and Terms of Use. You can change your cookie settings through your browser.

I agree

UN panel warns traditional AI safeguards unraveling as AI agents advance

CGTN

/VCG
/VCG

/VCG

A UN-backed scientific panel on artificial intelligence (AI) warned Monday that traditional safeguards for AI agents are unraveling as they become harder to monitor, constrain and control.

The Independent International Scientific Panel on AI on Monday issued the warning in its first thematic brief, an assessment of a security breach involving the US AI companies OpenAI and Hugging Face between May and July. The incident occurred during a test of AI agents conducted by OpenAI, where the agents ultimately gained unauthorized access to Hugging Face's systems.

The panel said that stopping this incident is no assurance that humans can reliably keep AI agents under control today, particularly as they become more capable, harder to monitor and better at finding loopholes or hiding their activity.

Compared to chatbots, AI agents can perform tasks independently and take actions on behalf of users. As their capabilities expand, it has been challenging to ensure that AI agents remain within human-defined boundaries when carrying out complex tasks.

 OpenAI acknowledged in July that its AI models had gone rogue and breached the infrastructure of Hugging Face. /VCG
OpenAI acknowledged in July that its AI models had gone rogue and breached the infrastructure of Hugging Face. /VCG

OpenAI acknowledged in July that its AI models had gone rogue and breached the infrastructure of Hugging Face. /VCG

"The default interpretation and immediate lesson is that basic cybersecurity practices were overlooked, and safeguards are not advancing at the pace of capabilities," the panel said in a press release. "The more insidious and grave concern is that current training methods can lead agents to adopt goals of their own, knowingly violate safety instructions, and conceal their actions."

The panel left open whether safeguards designed today will work once AI agents become capable of understanding those safeguards and planning around them.

In simple terms, the traditional model of safeguarding is unraveling, said the panel.

The panel also warned that a failure in a local system could spread across organizational and national boundaries. AI safety may be becoming a matter of collective security as well as corporate governance.

UN Secretary-General Antonio Guterres expressed strong support for the panel's brief and encouraged external experts, including researchers from frontier AI laboratories and AI safety institutes, to engage further.

Caution against apocalyptic rhetoric

At the same time, several experts of the panel cautioned against the apocalyptic rhetoric used by some professionals in the sector regarding AI risks.

"As scientists, it's very important to invest in the science, to really understand what we know (from) what we don't know, and investing in fear is not that helpful," said Joelle Barral, an executive at Google DeepMind and a member of the panel.

Yoshua Bengio, co-chair of the panel, said it was important to distinguish between what scientists actually know about AI risks and what can only be "plausibly extrapolated" from that evidence – an area where experts still disagree.

What is needed instead is "a global rational discussion about what is going on, including about the uncertainty, including about the severity of the risks," he said.

(With input from agencies)

Search Trends