OpenAI Says Upcoming Astra Model May Reach Critical Cyber Capability Threshold

OpenAI said preliminary evaluations of its upcoming Astra model show significant advances in agentic coding and cybersecurity, prompting the company to strengthen security controls because it cannot rule out that the model may reach the Critical cybersecurity capability threshold under its Preparedness Framework.

The assessment represents a potentially significant capability transition for OpenAI. Previous models, including GPT-5.6 Sol, were evaluated for frontier cybersecurity capabilities but were assessed at the High rather than Critical threshold.

OpenAI said its conclusion about Astra followed internal evaluations conducted over several days along with assessments from cybersecurity experts. The company emphasized that testing remains underway and that the preliminary results do not yet establish definitively that Astra has reached the Critical level.

Under OpenAI’s Preparedness Framework, a model can reach the Critical cybersecurity threshold if it is capable of independently identifying and developing functional zero-day exploits across many hardened, real-world critical systems, including vulnerabilities spanning different severity levels.

The threshold can also be reached if a model is capable of devising and carrying out novel, end-to-end cyberattack strategies against hardened targets after receiving only a high-level objective.

OpenAI said Astra’s preliminary performance is strong enough that Critical-level capabilities currently cannot be excluded, leading the company to begin preparing its safeguards as though deployment of such capabilities may become necessary.

The company is introducing stricter controls around higher-capability models and related research activities. Measures include isolated testing environments, tighter restrictions on network and tool access, stronger protections and encryption for model weights, additional monitoring and detection systems, and sandboxed execution environments.

OpenAI is also pausing internal Astra-related activities that do not yet satisfy the strengthened security requirements.

The company said it has implemented universal monitoring for risky actions and potential misalignment across Astra’s agentic applications, including during training and evaluation.

Those monitoring systems can evaluate the model’s reasoning activity and trigger a security response that allows high-risk behavior to be reviewed and interrupted.

OpenAI also plans to work with relevant government agencies and selected AI safety organizations to evaluate Astra’s capabilities.

Third-party testing organizations conducting higher-risk evaluations and workloads will receive recommended security controls intended to reduce the risks associated with testing increasingly capable AI systems.

The company said the approach follows the structure of its Preparedness Framework, which OpenAI first published in December 2023 to establish processes for identifying and responding to emerging frontier AI capabilities.

The framework covers areas including cybersecurity, biological and chemical risks, and AI self-improvement.

OpenAI previously applied similar measures as its models approached the High capability threshold for biological applications in 2025, increasing safeguards, expanding testing and working with external experts.

At the same time, OpenAI argues that increasingly capable cybersecurity models could provide significant defensive benefits by helping organizations find and remediate vulnerabilities before attackers exploit them.

The company plans to work with governments, safety institutes and civil society organizations on the responsible deployment of advanced cyber-capable models.

OpenAI also clarified that Astra was not involved in the exploitation of Hugging Face, distinguishing the upcoming model’s cybersecurity evaluation from that separate security incident.

The disclosure highlights the growing cybersecurity implications of frontier AI models as agentic systems become more capable of independently writing code, using tools, analyzing software environments and completing longer sequences of technical tasks.

For OpenAI, the Astra evaluations are now triggering safeguards designed for a capability level beyond that assigned to GPT-5.6 Sol and previous models.

KEY QUOTE:

“We are sharing this because we believe it’s important to be transparent with the public and the safety and security communities.”

OpenAI statement