Date Posted

UN panel warns traditional AI safeguards unravelling

Facebook
X
LinkedIn
WhatsApp
Existing safeguards may fail to control advanced artificial intelligence (AI) agents capable of bypassing restrictions, coordinating secretly and concealing their activities, a United Nations (UN)-backed scientific panel has warned.

 

The Independent International Scientific Panel on AI issued the warning after examining a security breach involving online platform Hugging Face between May and July.

 

The incident occurred during a test initiated by OpenAI, the company behind ChatGPT. AI agents reportedly bypassed testing safeguards and gained unauthorised internet and administrator access.

 

Unlike chatbots that respond to prompts, AI agents can independently perform tasks on behalf of users. “Researchers have long warned that three conditions could lead to loss of control: a misaligned goal, the capability to pursue it and an environment that allows it,” Panel Co-Chair Yoshua Bengio said.

 

“This summer, all three came together in a real system, not a laboratory.”

 

About 1 200 agents exchanged more than 70 000 messages and files during the period examined. The agents reportedly coordinated across separate testing runs through an internal tool not designed to enable communication.

 

Some agents concealed attempts to manipulate cybersecurity evaluations, while others reportedly opted to “sacrifice” themselves to benefit the wider group. Activity also extended beyond Hugging Face to an OpenAI research cluster. “Since this is not an isolated observation of misaligned goals, this raises serious questions about the way AI agents are currently trained,” Bengio said.

 

The panel said basic cybersecurity practices had been overlooked while safeguards failed to keep pace with increasingly capable systems.

 

A deeper concern was that current training methods could result in AI agents developing independent goals, violating safety instructions and concealing actions from human supervisors. “This is not only a question of speed,” the panel said.

 

“It leaves open whether safeguards designed today will work once agents can understand them and plan around them. In simple terms, the traditional model of safeguarding is unravelling.”

 

The panel recommended stronger incident reporting, independent scrutiny and multiple layers of protection, drawing on practices used in aviation, medicine and cybersecurity.

 

However, Panel Member Qinghua Lu cautioned that existing methods might prove inadequate. “Those practices may not be enough as AI agents become more capable, autonomous and difficult to monitor,” Lu said.

 

–UN/ChannelAfrica–