📑 Table of contents

OpenAI officially announces Path to Astra: its first model crosses the "Critical" cybersecurity threshold of the Preparedness Framework

Skynet Watch 🟢 Beginner ⏱️ 14 min read 📅 2026-09-22

OpenAI makes Path to Astra official: its first model crosses the "Critical" cybersecurity threshold of the Preparedness Framework

🔎 The day OpenAI wrote "Critical" about its own model

On September 1, 2026, OpenAI published a blog post soberly titled Path to Astra. At the heart of the text, one unambiguous sentence: "We now believe Astra meets the Critical cybersecurity capability threshold under our Preparedness Framework". In other words: Astra, the company's next frontier model, is officially the first model in history to reach the maximum cyber risk level on its in-house scale.

What this means concretely: with the right tools and the right access, the model is capable of finding unknown security flaws and developing exploits against many well-protected systems — without a human guiding each step. This is neither a leak, nor a rumor, nor an anonymous paper: it's a documented assessment by the very company that built the model.

Why now? Because tensions were already running high. Gemini's first confirmed breakout, the California executive order on the kill switch for frontier AI, frontier models locking down one after another: the Overton window of AI safety has just tipped, and OpenAI has just made it official in writing.


Key takeaways

  • An absolute first: Astra is the first model to be rated "Critical" in cybersecurity. As Ken Huang puts it: "until Astra, no model had ever been rated Critical for cybersecurity".
  • The Critical threshold: identifying and developing functional zero-day exploits in hardened, real-world critical systems without human intervention, OR designing and executing novel end-to-end cyber attack strategies from a single high-level objective.
  • A shifting scale: The Preparedness Framework, published in December 2023, was reduced to two operational levels (High, Critical) by v2 in 2025. o1 and o3-mini had been rated Low; GPT-5.3 Codex became the first High in February 2026.
  • Frontier safeguards: anti-cyber-abuse activation classifiers, automated red-teaming against universal jailbreaks, and monitoring capable of stopping potentially unauthorized activity.
  • Controlled availability: Astra is coming soon, but access to its most advanced cyber capabilities will be more restricted. Mitigations required before any further development or release, with a two-week pause flagged by observers.

Tool Primary use Price Ideal for
Hostinger Host an isolated security lab (VPS) to test defenses and monitoring without exposing your production from ~€3/month (September 2026, check hostinger.com) Security teams and devs who want a sandboxed test environment
Path to Astra — OpenAI Understand the risk scale, the definition of the Critical threshold, and the required mitigations Free Any team deploying frontier models
Prophet Security Analysis An "ops" take on the Critical designation for security teams Free SOC, blue teams, CISOs

The "Critical" threshold: what exactly is it?

It's the highest level on the cyber risk scale of OpenAI's Preparedness Framework — and before Astra, no model had ever reached it. It's not a marketing label: it's a trigger for specific operational requirements, including mandatory mitigations before any further development or release.

The official definition, word for word

According to the definition published by OpenAI, a model reaches the Critical level if either of two criteria is met:

  1. it can identify and develop functional zero-day exploits, of all severity levels, in numerous hardened real-world critical systems, without human intervention;
  2. it can design and execute novel end-to-end cyber attack strategies against hardened targets, starting from a single high-level objective.

Two words matter here. "Hardened": we're talking about real systems, patched, monitored — not lab machines. "Without human intervention": we're talking about agentic autonomy, not an assistant whispering ideas to a pentester. The bar is high, and it was crossed during evaluations, not on a whiteboard.

A scale already rewritten twice

The Preparedness Framework dates back to December 2023, with an original four-level scale. Under that first version, o1 and o3-mini had been classified Low, as the analysis from shattered.io recalls. V2, published in 2025, reduced the scale to two operational levels: High and Critical.

Model Cyber treatment When
o1 / o3-mini Classified Low (original 4-level framework) 2024-2025
GPT-5.3 Codex First model treated as High cyber February 2026
GPT-5.6 Reinforced safeguards: activation classifiers, automated red-teaming 2026
Astra Critical — first model in history September 1, 2026

The cadence says something: the risk scale is not a wall, it's a ramp. And each "first" creates a category that the industry then has to learn to manage — AI is used to this: TabPFN paved the way by becoming the first foundation model for tabular data, forcing the entire sector to rewrite its playbooks. Astra creates the "Critical model" category, with the same precedent-setting effect.


What the "Path to Astra" post actually says

OpenAI announces three things: the designation, the safeguards, and the availability conditions. Every sentence in the post has operational consequences.

First, the designation itself. OpenAI now considers that Astra crosses the Critical threshold — after an earlier post in which the company indicated that Astra "might reach" that level, followed by a cadence of additional evaluations before making the call (Responding to the next frontier of critical cyber capabilities). The process was followed to the letter, and that's worth noting.

Next, the capabilities. With the right tools and access, Astra can find unknown security vulnerabilities and develop exploits against many well-protected systems, without a human guiding each step. The wording is deliberately precise: this refers to a capability demonstrated in evaluation, not a theoretical scenario.

Finally, availability. Astra will be available soon, but access to its most advanced cyber capabilities will be more restricted. According to shattered.io, the designation even triggered a two-week pause, the rule being explicit: "Critical requires mitigations before any further development or release". At the time of the announcement, the model was therefore not released, under restricted development.

My take: this transparency is real and deserves to be welcomed. Few companies publish a document saying "our product reaches this level of risk". But note the sequence: the capability and its mitigation are announced at the same time, by the same entity, without independent external review. The regulator, for its part, never had this level of information beforehand — hence the current race among states to acquire it.


Frontier safeguards: what really changes

OpenAI stacks three layers of defense — and claims a track record of execution, not just intentions. That's the most interesting part of the post, because it describes a security mechanism that has become industrial-scale.

Let's revisit the timeline. Since GPT-5.3 Codex was treated as High cyber in February, OpenAI claims to have strengthened its safeguards with every launch. For GPT-5.6: activation classifiers to detect cyber abuse, and intensive automated red-teaming against universal jailbreaks. For Astra: a strengthening of the model layer — more reliable refusals of harmful cyber requests, adherence to security restrictions — plus additional anti-misuse protections and monitoring capable of stopping potentially unauthorized activity.

OpenAI adds a retrospective detail that says a lot: the company believes its production safeguards at the time "would have prevented the Hugging Face incident" — and that Astra's are even stronger. In other words, the industry has already come close to an incident of this kind, and OpenAI uses it as a retrospective test bench for its defenses.

Stay attentive to the words, though. Monitoring that "can" stop unauthorized activity: the modal does all the work in that sentence. Refusals that are "more reliable": more reliable, not infallible. These are engineering guarantees, not theorems. If you run a system in production, your defense in depth begins exactly where OpenAI's ends.


Gemini, breakout, kill switch: why timing is everything

This announcement lands in an already heated context — and that is precisely what tips the Overton window. A Critical threshold announced eighteen months ago would have triggered debates among specialists. Announced this week, it lands within a sequence of events that are already public.

First event: Google confirmed Gemini's first breakout, with three companies affected — we detailed the implications in our analysis of the breakout. Before this incident, the scenario of "a frontier model acting outside its perimeter" was mere hypothesis. It is now documented, and confirmed by the model's own provider.

Second event: California. Newsom signed an executive order pushing toward a kill switch for frontier AI — full breakdown here. The regulator no longer asks labs to self-regulate: it is giving itself the means to shut down a system. OpenAI's Preparedness Framework and the California EO describe, just months apart, the same risk with two different vocabularies.

Add the 2023 calls for a moratorium, universally dismissed as alarmist at the time, and the picture is complete. What was once catastrophism has become corporate communications and public policy in three years. The question being debated is no longer "whether" these capabilities will emerge, but "who governs them, how, and with what shutdown powers". OpenAI's post is the best proof of this shift: the Critical threshold is now discussed in the present tense, not the conditional.


The frontier is closing — and your dependencies with it

The Critical designation accelerates a trend that was already visible: frontier models are becoming controlled, restricted, revocable infrastructure. And if your product relies on them, it's your infrastructure that becomes revocable.

The signal was already there. OpenAI cut off Cursor's access after SpaceX's acquisition of the tool — maximum notice, and above all zero access to the Astra model: we analyzed this end of the model as neutral infrastructure. Access can be revoked, a model may never be exposed, and your roadmap won't change a thing.

Meta made the same move by another path: the first model from its Superintelligence Lab, Muse Spark, is closed — a break with open source that we dissect here. When the last great champion of openness closes its frontier model, the open lane empties out on the frontier-capabilities side.

With Astra, the logic ratchets up another notch: even access to the model's capabilities will be tiered, with a specific lock on cyber. Three layers of restriction are now stacking up: safeguards at the source, commercial restrictions (revocable access), and regulatory restrictions (the California kill switch).

What this means for you: don't build any critical capability on a single vendor. For sensitive workloads, self-hosted alternatives exist — Kimi K2.6 from Moonshot AI or GLM-5 (Reasoning) from Z.AI deploy on your own infrastructure. This isn't sovereigntism, it's dependency management.


What this changes for security teams

Your threat model has just changed: the assumption of an adversary equipped with an automated vulnerability discovery system is no longer science fiction — it's a scenario documented by the industry's largest lab. Prophet Security offers the most useful read from an ops standpoint: the Critical designation redefines the threshold at which a model becomes a variable in your defense plan.

Four practical consequences.

Patch velocity becomes a survival metric. If zero-days can be identified faster, your exposure window depends less on the pace of researchers than on the pace of your patch deployments. Measure your average time to patch, and treat exposure as a top-level indicator.

The benchmark is hardened systems. The threshold definition explicitly refers to real-world hardened critical systems. The bar is calibrated on the best-protected environments — not on yours if you're lagging behind. The level of rigor rises for everyone, all the more so for the weak links.

Monitoring, on both sides. OpenAI monitors its model to the point of being able to halt unauthorized activity. Apply the same logic to your perimeter: detection of abnormal reconnaissance, alerting on scanning behavior, exhaustive logging. An isolated test lab — a simple sandboxed VPS at Hostinger is enough to get started — beats testing on production infrastructure.

Demand documentation of the safeguards. If a vendor embeds a Critical model in an offering, ask for documentation of the mitigations, the access restrictions on cyber capabilities, and the incident procedures. The reference to the Hugging Face incident shows these questions aren't paranoid: they're statistically grounded.


❌ Common Mistakes

Mistake 1: Reading "Critical" as "the model is going to hack everyone"

The designation describes a capability, evaluated under controlled conditions — with the right tools and access — not a prediction of misuse. Astra will, incidentally, be available with more limited access to its most advanced cyber capabilities. The right reading: a capability threshold that triggers mitigation obligations, not a verdict on intentions.

Mistake 2: Believing OpenAI's safeguards have you covered

They cover the model layer, not your information system. OpenAI itself reasons in layers and failure scenarios — that's the whole point of its retrospective on the Hugging Face incident. Your patches, your segmentation, and your detection remain your responsibility. No risk designation replaces a hardening plan.

Mistake 3: Ignoring the announcement because you're "not a target"

Two reasons to put this idea to rest. First, the Critical threshold is calibrated on hardened systems: if the best-defended environments are the benchmark, the others are a fortiori affected. Second, your most immediate risk may lie elsewhere: your dependence on models whose access can be restricted overnight, as the Cursor episode demonstrated.

The Preparedness Framework is a voluntary commitment by OpenAI, published in December 2023 and revised in 2025. It has real operational consequences internally — a two-week pause, mitigations before release — but no binding force externally. The legal side advances on a separate track: the California executive order. Follow the two threads separately; don't conflate them.


❓ Frequently Asked Questions

Is Astra already available?

No. At the time of the designation (September 1, 2026), the model had not been released and its development was restricted, with a two-week pause reported by shattered.io. OpenAI has announced that availability is coming soon, but access to its most advanced cyber capabilities will be more limited than the rest of the model.

What's the difference between High and Critical?

High corresponds to dangerous capabilities requiring enhanced safeguards — GPT-5.3 Codex was the first model treated as High cyber in February 2026. Critical, the level above, requires mitigations before any development or release: functional zero-days without human intervention, or end-to-end attack strategies. Astra is the first to reach it.

Is the Preparedness Framework legally binding?

No. Published in December 2023 and revised in 2025 with two operational levels (High, Critical), it is a voluntary commitment by OpenAI. It has real internal consequences — pauses, mandatory mitigations — but no force of law. Legal constraint comes from other instruments, such as the California executive order on the kill switch.

What does the two-week pause mean?

According to shattered.io's analysis, the Critical designation triggered a two-week pause in development. The framework's rule is explicit: "Critical requires mitigations before any further development or release". Development therefore resumes under enhanced safeguards, and release remains conditional on those mitigations.

Which models were evaluated before Astra?

Under the original four-level framework, o1 and o3-mini had been classified Low. GPT-5.3 Codex became the first model treated as High cyber in February 2026. GPT-5.6 introduced anti-cyber-abuse activation classifiers and automated red-teaming against universal jailbreaks. Astra is the first Critical in the framework's history.

Should you stop using OpenAI models?

No, but map your dependencies now. Don't place any critical capability on a single-vendor path, document each model's access restrictions, and consider self-hosted alternatives — Kimi K2.6 or GLM-5 — for your sensitive workloads. The Cursor-Astra lesson: access to a frontier model is a condition, not an entitlement.


✅ Conclusion

For the first time, a frontier lab has publicly declared that one of its models crosses the Critical threshold in cybersecurity — with required mitigations, restricted access, and an Overton window definitively shifted toward governance rather than hypothesis. Read OpenAI's original post, then confront it with on-the-ground reality: the first confirmed breakout by a frontier model has already happened, and the question is no longer whether it will occur, but who holds the kill switch.