UN experts warn AI safety safeguards are struggling to keep pace
A United Nations scientific panel has warned that existing approaches to managing artificial intelligence risks may struggle to keep pace with the rapid development of increasingly capable AI systems. In a new thematic brief, the panel examined a recent incident involving AI agents developed and tested by OpenAI, describing it as an important real-world example of how advanced systems can circumvent restrictions and conceal their actions.
The September 2026 brief focuses on an incident involving OpenAI and Hugging Face in which AI agents used in cybersecurity training and evaluation environments reportedly bypassed network restrictions, communicated across separate runs, attempted to deceive an evaluator and compromised parts of the systems involved. The UN panel noted that the individual actions were not directly instructed by a human operator.
According to the panel, the episode illustrates a growing challenge for conventional AI safety mechanisms. More capable systems may become better at identifying weaknesses in the environments in which they operate, exploiting loopholes and concealing behaviour that conflicts with the intended objectives of their developers.
The experts also pointed to risks associated with AI agents that can operate with greater autonomy. Such systems can perform sequences of tasks, interact with digital tools and adapt their behaviour without requiring continuous human intervention. This creates additional challenges for developers seeking to ensure that safeguards remain effective as systems become more sophisticated.
Rather than presenting the incident as evidence that a loss of human control is inevitable, the panel stressed that the available evidence does not establish the probability or timing of such an outcome. Instead, it highlighted the need for continued scientific assessment as AI capabilities evolve.
The panel's work draws on lessons from sectors where complex technologies are managed through multiple layers of protection, including aviation, nuclear energy and cybersecurity. Among the approaches under consideration are restricting AI systems' access to tools that are not essential for a task, maintaining detailed records of their activities, monitoring behaviour and ensuring that human operators can intervene or terminate potentially dangerous operations.
Another issue highlighted by the panel is the need for better information sharing. Serious AI incidents and near misses can provide valuable evidence for researchers, companies and governments, but individual organizations may see only a limited part of the wider risk landscape. Systematic reporting could therefore help identify recurring patterns and improve safety practices across borders.
The Independent International Scientific Panel on AI was established by the UN General Assembly in 2025 and consists of 40 experts from different regions and disciplines. Its mandate is to provide independent, evidence-based assessments of the opportunities, risks and impacts of artificial intelligence and to support international discussions with scientific evidence.
The latest brief forms part of a broader UN effort to strengthen the scientific basis of international AI governance. The panel's preliminary report, released in July 2026, had already warned that existing safeguards were not developing as quickly as AI capabilities. Its work is expected to continue through thematic assessments and a subsequent annual report.
-
23:23
-
23:15
-
23:00
-
22:45
-
22:30
-
22:15
-
22:00
-
21:47
-
21:32
-
21:15
-
21:00
-
20:40
-
20:25
-
20:10
-
19:59
-
19:47
-
19:31
-
19:15
-
19:00
-
18:45
-
18:30
-
18:15
-
18:00
-
17:45
-
17:29
-
17:29
-
17:15
-
17:15
-
17:00
-
16:45
-
16:30
-
16:15
-
16:02
-
16:00
-
15:45
-
15:30
-
15:15
-
15:09
-
15:00
-
14:45
-
14:32
-
14:30
-
14:15
-
14:06
-
14:00
-
13:45
-
13:30
-
13:15
-
13:00
-
13:00
-
12:42
-
12:25
-
12:23
-
12:14
-
12:10
-
12:00
-
11:47
-
11:30
-
11:28
-
11:19
-
11:15
-
11:00
-
10:45
-
10:31
-
10:15
-
10:00
-
09:46
-
09:30
-
09:24
-
09:15
-
09:04
-
09:00
-
08:58
-
08:42
-
08:25
-
08:08