OpenAI pauses some Astra work as tests flag possible critical cyber capabilities
The company is again tightening safeguards under its Preparedness Framework, following an earlier capability transition involving biology risks in June 2025
OpenAI has paused internal Astra work that does not meet strengthened security requirements after preliminary cybersecurity evaluations.
OpenAI is pausing some internal work involving Astra after preliminary tests suggested its upcoming frontier model could reach the Critical cybersecurity threshold in the company’s Preparedness Framework.
The company has not concluded that Astra possesses Critical cyber capabilities. It says, however, that recent internal evaluations and expert assessments found enough progress in agentic coding and cybersecurity that this classification can no longer be ruled out.
This is not the first capability shift to change how OpenAI handles a developing model.
When its models approached the High capability threshold for biology in June 2025, the company strengthened safeguards, expanded testing, worked with external experts and introduced additional security controls. OpenAI says it is applying the same principle to Astra.
The difference is the potential risk level. Previous models, including GPT-5.6-Sol, were assessed at the High rather than Critical cybersecurity threshold.
Katrina Mulligan, Head of National Security Partnerships at OpenAI for Government, described the latest response on LinkedIn as “measuring twice, cutting once before we release Astra.”
She added: “My team is supporting OpenAI's work with relevant government agencies and AI safety organizations to test the capabilities for this model.”
What OpenAI’s Critical threshold means
Under the Preparedness Framework, a model reaches the Critical cybersecurity threshold if it can independently identify and develop working zero-day exploits across many hardened, real-world critical systems.
The threshold can also apply if a model can devise and execute a new end-to-end cyberattack against hardened targets after receiving only a high-level objective.
OpenAI says Astra’s preliminary results were strong enough to warrant additional precautions while benchmarking and assessment continue. It has not published the underlying scores or detailed evaluation results, meaning the reported increase in capability cannot yet be independently assessed.
Development moves into tighter security conditions
OpenAI is introducing stricter controls for higher-capability models and related work. These include isolated testing environments, restricted network and tool access, stronger model weight protection and encryption, sandboxed execution, and additional monitoring and detection systems.
Internal Astra activities that do not meet the strengthened requirements have been paused.
The company has also introduced universal monitoring for risky actions and signs of misalignment across Astra’s agentic applications, including during training and evaluation. OpenAI says the monitors assess the model’s chain of thought and can trigger a security response to review and interrupt high-risk activity.
Recommended security controls will be provided to third-party partners conducting higher-risk evaluations and workloads. OpenAI also plans to test Astra with relevant government agencies and selected AI safety organizations.
The company argues that advanced cyber-capable models could help defenders identify and address vulnerabilities before attackers exploit them. However, it has not provided evidence of defensive outcomes or named the external organizations involved in the current testing.
OpenAI has not announced a release date for Astra. Benchmarking, safeguard testing and external capability assessments will continue before the model is deployed.