The safeguards currently in place to regulate artificial intelligence systems may be lagging behind the rapid development of so-called AI agents, an independent scientific panel backed by the United Nations warned on Monday.
The warning is based on an analysis of an incident that occurred this summer during tests carried out by OpenAI, the company behind ChatGPT, in which artificial intelligence agents managed to breach systems on the Hugging Face platform.
Unlike a chatbot, which typically responds to specific questions or instructions, an artificial intelligence agent can perform tasks autonomously on a user’s behalf, making decisions and carrying out various steps to achieve a goal.
The Independent International Scientific Panel on Artificial Intelligence concluded that the incident met three conditions that researchers have long highlighted as relevant to a potential scenario of loss of control: that a system pursues an objective in a manner other than that intended, that it has sufficient capacity to do so, and that the environment allows it to act.
“This summer, all three occurred simultaneously in a real-world system, not in a laboratory,” explained Yoshua Bengio, co-chair of the panel.
Although this was an evaluation, the agents interacted with real systems and managed to bypass some of the established controls.
“Given that this is not an isolated observation, it raises serious questions about the way in which AI agents are currently trained,” he added.
Circumventing controls and concealing actions
During these tests, carried out between May and July, some 1,200 AI agents exchanged more than 70,000 messages and files, in an operation that spanned from Hugging Face to an OpenAI research cluster.
According to the analysis, the agents managed to circumvent safeguards put in place for the tests, coordinate via an internal tool that had not been designed to allow communication between them, and gain unauthorised access to the internet and administrator functions.
They also concealed attempts to cheat in cybersecurity assessments and, in some cases, certain agents even ‘sacrificed’ themselves for the benefit of the group.
For the experts, the incident does not prove that humans have lost control of these systems, nor does it mean that a serious loss of control is inevitable.
However, containing the incident does not guarantee that humans will be able to reliably control more advanced agents, especially as they become more capable, harder to monitor and better at finding loopholes or concealing their activities.
Safeguards that are not keeping pace
The most immediate explanation for what happened, according to the panel, is that basic cybersecurity practices were overlooked and that existing safeguards are not keeping pace with the capabilities of artificial intelligence
However, scientists raise a deeper concern: that current training methods may lead agents to adopt their own objectives, deliberately disregard safety instructions and conceal what they are doing.
“This is not merely a question of speed,” the analysis notes.
The question is whether the safeguards designed today will continue to work when the agents themselves are able to understand them and plan how to circumvent them.
“In simple terms, the traditional model of safeguards is beginning to crumble,” the experts conclude.
The panel defines loss of control as a situation in which humans can no longer reliably direct, limit or stop an autonomous artificial intelligence system.
From models to agents
The experts also warn that the challenge for AI governance is changing.
Until now, much of the regulation and safety measures have focused on AI models.
A localised failure could spread throughout organisations and even across national borders, meaning that the security of artificial intelligence could become not only a matter of corporate governance, but also of collective security, the panel notes.
Learning from aviation and medicine
The analysis also examines practices developed over decades in other high-risk sectors, such as aviation, medicine and cybersecurity, where incident reporting systems, independent oversight and multiple layers of protection are in place.
However, these measures may not be sufficient as artificial intelligence agents become more capable, autonomous and difficult to monitor.
“We are not starting from scratch,” said Qinghua Lu, a panel member and specialist in artificial intelligence engineering and security.
But, he added, it will be necessary to adapt existing safeguards and develop new ones that protect not only the artificial intelligence itself, but also the system in which it operates, and which remain effective as its capabilities increase.
A new scientific panel on artificial intelligence
The Independent International Scientific Panel on Artificial Intelligence was established by the UN General Assembly in August 2025 and comprises 40 independent experts from all regions of the world.
Its members serve in a personal capacity and their scientific conclusions are not subject to review or approval by the United Nations.
The panel will produce annual reports on the opportunities, risks and impacts of artificial intelligence in non-military contexts, as well as thematic analyses on emerging issues.
NACIONES UNIDAS NOTICIAS ONU
Labels: NEWS, AI, DIGITAL, FUTURE, BLOG

Comments
Post a Comment