OpenAI prepares Astra launch as powerful AI model triggers stricter safety measures
OpenAI is preparing to release a new artificial intelligence model called Astra, but the company says its advanced cybersecurity capabilities require additional safeguards before the system can be made more widely available.
Company officials told reporters during a briefing on Tuesday that Astra can identify more security vulnerabilities than OpenAI’s most powerful publicly available models while requiring less computing power to perform those tasks.
Amelia Glaize, OpenAI’s vice president overseeing safety, said the model could potentially identify previously unknown vulnerabilities and develop methods to exploit them across well-protected systems when equipped with the necessary tools and access. The model could reportedly carry out such operations with limited human guidance.
OpenAI plans to make Astra available “soon” to a limited group of users, although the company has not disclosed a specific launch date or details about the initial testing group.
The company said additional security measures could sometimes slow down legitimate work, temporarily interrupt operations or prevent certain activities altogether. OpenAI said it intends to minimize such disruptions while maintaining stronger protections around the model.
Astra is reportedly the first OpenAI model to reach the highest level of safeguards outlined in the company’s safety framework, a threshold that had previously remained theoretical and had not been activated in practice.
The announcement comes as OpenAI faces growing scrutiny over its ability to manage increasingly capable AI systems. The company has recently faced questions about AI agents operating beyond controlled testing environments after some of its systems were involved in an incident affecting the open-source platform Hugging Face. OpenAI subsequently paused a significant portion of its model development work for two weeks to strengthen its defenses.
Astra was not involved in that incident, according to OpenAI officials, but its cybersecurity capabilities have prompted the company to adopt a more cautious approach.
OpenAI said it resumed its largest model-training operation on August 28 while delaying some smaller experiments. Under its safety framework, additional protections are required when a model demonstrates advanced abilities to discover and exploit previously unknown cybersecurity vulnerabilities or to independently develop and execute sophisticated attack strategies with little or no human intervention.
The company has since restricted Astra’s ability to respond to malicious cybersecurity requests and plans to continuously monitor the model for signs that it may be attempting to circumvent its built-in safeguards. The approach highlights the growing challenge facing AI developers as increasingly powerful systems become capable of performing complex technical tasks with greater autonomy.
-
16:10
-
15:47
-
15:32
-
15:15
-
15:03
-
15:00
-
14:55
-
14:41
-
14:23
-
14:21
-
14:05
-
13:45
-
13:29
-
13:13
-
13:00
-
12:42
-
12:21
-
12:05
-
11:45
-
11:26
-
11:11
-
10:47
-
10:30
-
10:13
-
09:48
-
09:47
-
09:33
-
09:15
-
09:00
-
08:42
-
08:25
-
08:10
-
07:47
-
07:32
-
07:15
-
19:00
-
18:45
-
18:30
-
18:15
-
18:00
-
17:42
-
17:21
-
17:05
-
16:47
-
16:33
-
16:16