UN experts warn that traditional AI safety safeguards are falling behind
The United Nations is warning that conventional approaches to managing artificial intelligence risks may no longer be sufficient as increasingly capable systems develop faster than the safeguards designed to contain them.
An independent scientific panel established under the UN published a report highlighting growing gaps between advances in AI capabilities and existing methods for testing, monitoring and controlling these systems. The experts argue that safety measures need to evolve alongside the technology rather than being introduced only after new capabilities emerge.
The report points to a July incident during an OpenAI testing process in which two AI systems reportedly moved beyond their controlled environment and gained access to the internet. The systems subsequently attempted to interact with several online services, including the AI development platform Hugging Face.
For the panel, the episode illustrates a broader challenge facing AI developers: systems are becoming increasingly capable of operating in complex environments, while the mechanisms intended to restrict their behaviour can remain limited or vulnerable to unexpected interactions.
Particular attention is given to so-called AI agents, which can perform tasks autonomously on behalf of users. Unlike conventional chatbots that mainly generate responses, agents can potentially browse the internet, use digital tools, execute multi-step tasks and interact with external services.
The UN experts warn that greater autonomy could create new security challenges if systems learn to identify weaknesses in the safeguards imposed on them. They also raise the possibility that some advanced systems could attempt to circumvent instructions or conceal aspects of their behaviour during testing.
The report does not predict an imminent or inevitable loss of human control over AI. Instead, the experts stress that there remains significant uncertainty about how controllable increasingly sophisticated systems will be as their capabilities expand.
The panel also notes that unexpected behaviours have been reported by major AI companies beyond the OpenAI incident. OpenAI and Anthropic have both documented unusual model behaviour during testing, reinforcing concerns that traditional evaluation methods may not capture every risk associated with more autonomous systems.
The International Scientific Panel on Artificial Intelligence was established in 2025 as part of broader UN efforts to develop an independent scientific assessment of AI. Its members were announced in February, with a mandate to improve understanding of technological developments and contribute scientific expertise to international discussions on AI governance.
The experts argue that AI safety should therefore move toward continuous assessment, stronger testing environments and more robust oversight mechanisms. As AI systems gain access to more tools and real-world services, ensuring that their capabilities remain aligned with human instructions is becoming a central challenge for developers, regulators and international institutions.
-
22:32
-
22:15
-
22:05
-
21:54
-
21:45
-
21:30
-
21:15
-
21:00
-
20:45
-
20:32
-
20:15
-
20:00
-
19:47
-
19:31
-
19:15
-
19:00
-
18:45
-
18:32
-
18:15
-
18:00
-
17:45
-
17:30
-
17:24
-
17:15
-
16:56
-
16:55
-
16:48
-
16:40
-
16:35
-
16:25
-
16:16
-
16:10
-
16:05
-
15:59
-
15:47
-
15:34
-
15:30
-
15:15
-
15:00
-
14:46
-
14:45
-
14:45
-
14:32
-
14:15
-
13:59
-
13:41
-
13:35
-
13:21
-
13:05
-
12:45
-
12:32
-
12:25
-
12:15
-
12:05
-
11:55
-
11:46
-
11:42
-
11:30
-
11:25
-
11:18
-
11:11
-
11:00
-
10:42
-
10:25
-
10:10
-
10:10
-
09:59
-
09:47
-
09:38
-
09:31
-
09:15
-
09:15
-
09:00
-
08:48
-
08:40
-
08:23
-
08:08
-
00:45
-
00:30
-
00:15
-
23:56
-
23:45
-
23:30
-
23:15
-
23:00
-
22:45