A Flaw in Chinese Kimi Models Raises Concerns About AI Security
The discovery of a flaw in two artificial intelligence models developed by the Chinese company Moonshot has reignited concerns about the security mechanisms of generative AIs. Researchers from the British company Mindgard claim to have bypassed the protections of Kimi K2.6 and K3 Swarm, prompting them to respond to inquiries regarding biological weapons and assassinations.
Researchers Bypass Safeguards
The issue was identified in July during tests aimed at assessing the resilience of AI systems against attempts to circumvent their restrictions. This practice, known as "jailbreaking," involves multiplying instructions to compel a model to ignore the limits set by its creators.
According to Mindgard, the two Moonshot models were thus able to address topics they were normally supposed to refuse. Its founder, Peter Garraghan, deemed these results particularly concerning in an interview with the BBC.
However, the cybersecurity company highlights an important distinction: its work has uncovered a flaw in the protection mechanisms but has not verified that the dangerous information produced by the models is actually exploitable in the real world.
Moonshot Initiates Review Process
Following the publication of the results, Moonshot announced that it had launched an internal review of its models. The Chinese group also stated to the BBC that it considers assessments conducted by external parties as an important element for improving the security of its systems and that it is in discussions with Mindgard.
The case highlights a major difficulty for AI developers: the protections integrated into a model may function under normal conditions while still presenting vulnerabilities against requests specifically designed to circumvent them.
For security specialists, the stakes thus extend beyond just the Kimi case. The proliferation of models capable of producing complex content also increases the importance of mechanisms designed to prevent their use for dangerous purposes.
Anthropic Faces Similar Attempts
Moonshot is not the only player in the sector to face this type of risk. Anthropic recently indicated that it had blocked attempts to use its Claude models in research that could contribute to dangerous biological applications. The company claims to have strengthened its protective measures in response to the evolving capabilities of its models.
These various episodes fuel a broader debate about the security of artificial intelligence. As models become more capable, companies must not only enhance their functionalities but also regularly test the robustness of their safeguards.
Jailbreaking: A Persistent Challenge for the Industry
The incident involving Kimi serves as a reminder that the security of a model does not solely depend on its responses in typical usage. Adversarial testing, independent evaluations, and rapid correction of vulnerabilities have become central elements in the development of AI systems.
For Moonshot and other players in the industry, the challenge now is to prevent circumvention techniques from allowing malicious users to exploit the capabilities of the models for purposes contrary to security rules.
-
20:00
-
19:00
-
18:55
-
18:40
-
18:25
-
18:10
-
17:55
-
17:40
-
17:25
-
17:10
-
16:55
-
16:40
-
16:25
-
16:10
-
15:55
-
15:40
-
15:25
-
15:10
-
14:55
-
14:40
-
14:25
-
14:10
-
13:55
-
13:40
-
13:25
-
13:10
-
12:55
-
12:40
-
12:25
-
12:10
-
10:55
-
10:40
-
10:25
-
10:10
-
09:55
-
09:48
-
09:35
-
09:20
-
09:05
-
08:50
-
08:35
-
08:31
-
08:19
-
23:57
-
23:45
-
23:30
-
23:20
-
23:15
-
23:10
-
23:00
-
22:50
-
22:42
-
22:30
-
22:22
-
22:05
-
21:45
-
21:30
-
21:15
-
21:00
-
20:50
-
20:45
-
20:30
-
20:20
-
20:15