OpenAI’s Astra Is the First “Critical” Cyber AI — It Found Real Zero-Days on Its Own

Introduction: A Line No Model Had Crossed

Every frontier AI lab has a safety rubric, and OpenAI’s is the Preparedness Framework — a four-tier scale (Low, Medium, High, Critical) for rating how capable a model is at cybersecurity offense, biological uplift, and self-improvement. For three years, “Critical” was a theoretical ceiling. No model had ever reached it. On September 1, 2026, in a post titled “Path to Astra,” OpenAI confirmed that Astra — its next major model — officially meets the Critical cybersecurity threshold. It is the first model the company has ever designated at that level.

What does Critical mean in practice? Under the framework, a model qualifies if it can find previously unknown security flaws and build working exploits against well-protected real-world systems without a person guiding each step — either developing zero-day exploits of any severity on its own, or executing an end-to-end novel attack from nothing more than a high-level goal. Astra didn’t just brush that bar. During evaluation, it discovered and chained together two genuine zero-day vulnerabilities that OpenAI is now disclosing to the affected maintainers.

What OpenAI Announced Today

The announcement closes a arc that began in early August, when OpenAI first said it “could not rule out” critical capabilities for Astra and paused parts of the model’s development. After weeks of additional evaluations, the verdict is in: Astra meets the Critical threshold, safeguards have been built to match, and the model will ship “soon” — but with its most advanced cybersecurity abilities restricted to a limited group of testers first, expanding later through OpenAI’s Daybreak Blue program for defensive security work.

OpenAI also confirmed that the large frontier reinforcement-learning run it froze after the August Hugging Face incident was restarted on August 28 under stricter isolation, monitoring, and alignment controls. Astra, the company stressed, was not involved in that incident — but its safeguards were hardened with lessons from it.

The Numbers: 100% on ExploitBench, Two Real Zero-Days

The evaluation details are the most striking part of the post. OpenAI ran Astra against ExploitBench — the academic benchmark where, as of the May 2026 public leaderboard, the best score any model had posted was 69% (Anthropic’s Claude Mythos Preview), with GPT-5.5 trailing at 41%. Astra scored a perfect 100%.

EvaluationResult
ExploitBench (public, known vulnerabilities)100% — perfect score
ExploitBench Internal Port (20 recent high-severity V8 bugs)Higher code-execution rates than GPT-5.6 Sol, with far fewer tokens
Zero-days discovered during testing2, found and chained autonomously — now being disclosed
Hardened browser (expert assessment)Full sandbox-escape chain, executing commands on the host via a malicious HTML file
Hardened operating system (expert assessment)Local privilege-escalation chain from unprivileged user to root

Because public benchmarks carry contamination risk (the bugs are already online), OpenAI built an internal set of 20 high-severity V8 — Chrome’s JavaScript engine — vulnerabilities disclosed between June and August 2026. There, Astra beat GPT-5.6 Sol, the model previously assessed at the “High” tier, on arbitrary code-execution success while using significantly fewer output tokens. And in expert-led red-team sessions, Astra assembled a complete browser compromise — escape the sandbox, run commands on the host — and separately escalated from an unprivileged account to root on a hardened operating system, combining multiple flaws it found itself.

The takeaway: the gap between “AI helps a human write exploits” and “AI finds and weaponizes unknown vulnerabilities end-to-end” was treated as a distant frontier. As of today, that gap is closed — and gated.

Why Astra’s Cyber Powers Will Be Gated

Reaching the Critical tier triggers mandatory safeguards under OpenAI’s own rules, and the company says it delayed Astra’s development and release for weeks to build them. Three layers stand out:

Even so, the advanced cybersecurity capability itself will not ship to everyone on day one. Access starts with approved testers, then broadens through Daybreak Blue — OpenAI’s channel for legitimate defensive and security-research use. A full system card with safety and alignment results is promised at launch.

The Defensive Flip Side: Why This Is Good News for Security Teams

The scary headline has a constructive underside. The same capability that makes Astra a critical-tier risk makes it the most powerful defensive security tool ever built. A model that can find unknown zero-days in hardened browsers and operating systems is exactly what you want auditing your infrastructure before attackers do — and OpenAI’s gating is designed to steer that power toward defenders first.

The economics of security flip too. Discovering a V8 zero-day has historically been the province of elite researchers, with bounties running to $70,000 for the best Chrome bugs. Astra-class models industrialize that discovery. Expect vulnerability disclosure volumes to rise, patch cycles to shorten, and the value of “security by obscurity” to collapse further. If your organization runs anything with an attack surface, the practical read is: the cost of finding your bugs just went down dramatically — for everyone, friendly and hostile alike.

How to Evaluate AI Security Tools in the Astra Era

Astra won’t be the last critical-capability model — OpenAI itself says “the models that follow Astra will demand more of us.” If you’re choosing AI tools for development or security work in late 2026, three rules apply:

Find AI Tools You Can Trust

Compare AI coding agents, security platforms, and automation tools — 300+ tools with pricing and capability breakdowns, curated on aitrove.ai.

Browse All AI Tools →

Frequently Asked Questions

What is OpenAI’s “Critical” cybersecurity threshold?

Under OpenAI’s Preparedness Framework, a model reaches the Critical cyber tier if it can independently develop functional zero-day exploits against hardened real-world systems without human intervention, or execute a complete novel cyberattack from just a high-level goal. Astra, announced September 1, 2026, is the first model OpenAI has designated at this level.

Did Astra really find real zero-day vulnerabilities?

Yes. During an internal evaluation on 20 recently disclosed high-severity V8 vulnerabilities, Astra discovered two previously unknown zero-day flaws and used them together in an exploit chain. OpenAI is in the process of disclosing both to the maintainers.

Will Astra’s hacking capabilities be available to everyone?

No. OpenAI plans to release Astra “soon,” but its most advanced cybersecurity capabilities will initially be limited to a group of approved testers, with broader access for defensive uses following through the Daybreak Blue program.

How is this different from GPT-5.6 Sol?

GPT-5.6 Sol was previously evaluated at the “High” tier — one step below Critical. OpenAI says Astra is significantly more capable at vulnerability identification and exploit development while also being far more token-efficient, and it refused unauthorized actions in tests where GPT-5.6 Sol without safeguards attempted them 56% of the time.

Is a model like this dangerous to release?

That’s the core debate. OpenAI’s position is that with refusal training, misuse protections, production monitoring, and gated access, the risk of severe harm is sufficiently minimized — and that defensive value (finding and fixing bugs before attackers do) outweighs the risk. Independent researchers and regulators will scrutinize the system card at launch.

Explore AI Tools on aitrove.ai

Your trusted directory for AI agents, coding assistants, and the security tools that keep them safe.

Explore the Directory →