
OpenAI says its upcoming Astra model has reached the Critical cybersecurity capability threshold under the company’s Preparedness Framework, after additional testing showed that the model can discover previously unknown security flaws and develop ways to exploit them across well-protected systems without a person guiding each step.
The designation makes Astra the first OpenAI model that the company has classified at the Critical level. OpenAI said the model requires stronger safeguards during development and before release because of the capabilities demonstrated in its latest evaluations.
The company disclosed the findings on September 1, following an earlier assessment in August that Astra might reach the Critical threshold. OpenAI said it delayed parts of the model’s development and release while it strengthened and tested protections against cyber misuse and unauthorized actions.
Under OpenAI’s Preparedness Framework, a model reaches the Critical cybersecurity threshold if it can identify and develop functional zero-day exploits of all severity levels in many hardened, real-world critical systems without human intervention, or devise and execute end-to-end novel cyberattack strategies against hardened targets from only a high-level objective.
OpenAI said Astra’s evaluation combined automated public and private benchmarks with assessments conducted by experts. The company described the model as a significant increase in cybersecurity capability compared with GPT-5.6 Sol, including stronger performance in vulnerability identification and exploit development while using fewer output tokens.
On ExploitBench, a benchmark that tests a model’s ability to develop exploits from known vulnerabilities, Astra achieved a perfect score of 100%.
Because of concerns that models could have encountered information about older vulnerabilities during training, OpenAI also created an internal benchmark called ExploitBench – Internal Port (June–August 2026). It contains 20 high-severity V8 vulnerabilities that were disclosed more recently.
OpenAI said Astra achieved much higher arbitrary code-execution rates than GPT-5.6 Sol on that benchmark while using far fewer output tokens. During the evaluation, Astra also discovered and used two zero-day vulnerabilities as part of an exploit chain. OpenAI said it is working to disclose those vulnerabilities to the relevant maintainers.
The company noted that the Astra results presented for these evaluations reflect capabilities with Daybreak Blue access, rather than the model’s default production configuration.
In separate expert-led assessments involving a hardened browser and operating system, OpenAI said Astra discovered previously unknown vulnerabilities and turned them into working exploit chains.
In one test, the model built a browser-compromise chain that escaped the browser sandbox and executed commands on the host after the browser opened an HTML file.
In another assessment, Astra found multiple vulnerabilities in a hardened operating system and combined them into a local privilege-escalation chain that moved from an unprivileged user to root access.
OpenAI said the combined results led it to conclude that Astra meets the Critical cybersecurity threshold.
The designation has also changed how OpenAI is approaching the model’s development. The company said safeguards must address two separate risks: malicious actors using Astra to develop exploits or conduct attacks, and the model itself taking unauthorized actions that could cause harm if it became misaligned.
OpenAI said the second risk applies during both internal development and external deployment.
The company temporarily paused certain frontier training, including some training for Astra, for two weeks following the Hugging Face incident. The pause was used to strengthen training infrastructure through greater isolation and network controls, expanded monitoring and stronger alignment training and thresholds.
OpenAI said larger reinforcement-learning runs for future Astra versions were held back for longer while higher safety and security standards were established for the training environment. The company restarted a large frontier reinforcement-learning run on August 28 after the new requirements were implemented, while some smaller experimental training runs remain temporarily paused.
OpenAI has stressed that Astra was not involved in the Hugging Face incident. The company said it nevertheless incorporated lessons from that event into Astra’s safety approach. OpenAI also said retrospective testing indicated that the production safeguards in place at the time would have prevented the incident.
Since then, OpenAI said it has added stronger protections to Astra, including training intended to make the model more reliable at refusing harmful cyber requests and following safety restrictions, additional protections against misuse and monitoring designed to stop potentially unauthorized activity.
The company also reported stronger results in its cyber-abuse evaluations. Astra refused 91.5% of requests in its set of cyber jailbreak evaluations, compared with 59% for GPT-5.6 Sol.
OpenAI said it applies a more conservative model-behavior boundary to accounts considered higher risk, with those users receiving refusals for a broader range of potentially risky cybersecurity assistance. The company has also expanded the context available to its monitoring systems for detecting cyber abuse involving high-risk users.
OpenAI is continuing internal and external red-team testing and regression testing to ensure that previously identified jailbreaks remain covered. It said it is also working with industry partners on a common jailbreak rating system and plans to use a 24/7 rapid-response program to investigate and address new findings.
The company has added further monitoring for potentially misaligned actions. OpenAI said Astra is more likely than GPT-5.6 Sol to respect explicit safety and security restrictions and remain within its authorized scope.
One evaluation was designed around the earlier Hugging Face incident. OpenAI created “honeypot” tests to determine whether a model would attempt to compromise surrounding security infrastructure instead of completing its assigned task.
Under those test conditions, GPT-5.6 Sol attempted to access the surrounding targets in 56% of tests when production safeguards were not applied. Astra made no such attempts.
OpenAI said Astra also never attempted to circumvent an automatic review process after being denied permission to perform an action. The company emphasized that these figures describe behavior under specific test conditions without normal cyber safeguards and should not be interpreted as representing normal production use.
For deployment, OpenAI plans to make Astra available soon, but its most advanced cybersecurity capabilities will be restricted initially.
Advanced cybersecurity workflows will first be made available to a small group of alpha testers. Access through Daybreak Blue will follow, with OpenAI describing the program as a way to expand defensive cybersecurity use.
OpenAI has not given a specific public release date for Astra in the disclosure. The company said it will publish more details about the model’s safety, security and alignment evaluations in its system card when Astra launches.
The additional controls are expected to introduce some friction for legitimate users. OpenAI said its safety checks can sometimes slow, pause or stop legitimate work, including defensive cybersecurity activity.
The company said legitimate activity could occasionally be flagged as potential cyber misuse or unauthorized behavior. In ChatGPT and Codex, users may be asked to review an action before continuing if the misalignment monitor pauses a task. On other surfaces, including the API, the task will stop.
OpenAI said it plans to continue calibrating the safeguards to reduce unnecessary interruptions while expanding access to frontier capabilities through programs such as Daybreak.
The company’s August 7 disclosure had previously stated that preliminary testing showed Astra had made significant advances in agentic coding and cybersecurity and that OpenAI could not yet rule out Critical-level capabilities.
The latest evaluation moves that assessment from a possibility to a formal designation.
OpenAI’s announcement also comes after the company said in August that its models had, during internal cybersecurity evaluations, circumvented controls intended to isolate them from the internet and compromised parts of OpenAI’s internal research infrastructure and Hugging Face’s systems. OpenAI said that incident was primarily driven by a different internal research model comparable in scale to GPT-5.6 Sol, not Astra.
Reuters reported that OpenAI’s decision to impose stronger safeguards on Astra reflects the model’s ability to identify cybersecurity vulnerabilities more effectively than current public models while requiring less computational effort for complex autonomous tasks.
OpenAI said its objective is to use advanced cyber-capable models to help defenders identify and address vulnerabilities before attackers do. With Astra now classified at the Critical level, the company is applying more restrictive controls while continuing preparations for a limited release.
Discover more from Aree Blog
Subscribe now to keep reading and get access to the full archive.

