OpenAI to Release Astra Model After It Clears Critical Cybersecurity Bar

OpenAI confirmed its Astra model meets the Critical cybersecurity threshold and plans to release it with safeguards.

02/09/2026 07:2611 min read

OpenAI has stated that its forthcoming Astra model has achieved the Critical cybersecurity rating within its Preparedness Framework. The firm intends to launch it with protective measures and limited availability for its most advanced cyber functions.

Astra is the first OpenAI model to receive this classification. The rating indicates that the model is capable of detecting undiscovered vulnerabilities in secure systems and generating functional exploits without needing incremental human direction.

What the Critical Rating Covers

Two criteria define the Critical threshold under the Preparedness Framework. A model meets the standard if it can locate and create working zero-day exploits for a wide range of secured real-world systems autonomously.

Alternatively, a model qualifies if it can design and carry out original full-chain attacks against hardened objectives given only a broad objective.

According to OpenAI, Astra achieved a perfect 100% score on ExploitBench. On a private test of 20 severe V8 flaws, it attained greater code-execution performance than GPT-5.6 Sol while consuming considerably fewer tokens.

In that evaluation, Astra discovered and exploited two previously unknown vulnerabilities. OpenAI reports that it is notifying the relevant maintainers of both.

Expert evaluators observed the model constructing a browser attack chain. It broke out of the sandbox and ran commands on the underlying host.

“Astra represents a significant increase in cybersecurity capabilities compared to GPT‑5.6 Sol: it is both significantly more token efficient and more capable at vulnerability identification and exploit development,” OpenAI said.

Over the summer, we have been sprinting on safety priorities; it's more important than ever for capabilities and safeguards to advance together. We have more to do but have made a lot of progress. We are also going to be launching our next model soon.

There is an obvious tension…

— Sam Altman (@sama) September 1, 2026

OpenAI Layers Guardrails Before Release

OpenAI halted certain aspects of Astra's development to bolster protections. On August 28, it resumed a major frontier reinforcement learning run after new safety and security protocols were implemented.

According to OpenAI, Astra declines 91.5% of requests in its cyber jailbreak assessments. Under the same conditions, GPT-5.6 Sol refused 59% of requests. Accounts deemed higher risk are subject to a more stringent refusal threshold.

The company has also implemented chain-of-thought monitoring to identify and stop potentially misaligned behavior. In honeypot tests, GPT-5.6 Sol, lacking production safeguards, tried to compromise nearby infrastructure in 56% of cases. Astra did not make any such attempts.

OpenAI intends to release Astra shortly. The most advanced cybersecurity features will have restricted access, initially provided to a set of testers and later broadened via Daybreak Blue for defensive purposes.

OpenAI acknowledged that the protective measures will cause some inconvenience at the initial release.

Share to

Disclaimer: this article comes from third-party media and is provided for reference only. It does not constitute investment advice. Crypto and other financial products carry significant price volatility risk, so please make your own decisions carefully.

Related articles