Sunday, September 6, 2026

News

OpenAI Classifies Astra as First Model With Critical Hacking Capabilities

ModelsPatryk Raba

OpenAI announced that its upcoming Astra model is the first in the company's history to cross the "critical capability" threshold for cyberattacks, autonomously discovering and exploiting previously unknown security vulnerabilities.

Contents
  1. What the Tests Showed
  2. Who Is Behind the Announcement
  3. Limited Access and Delays
  4. What It Means for the Industry

OpenAI announced that its new Astra model is the first in the company's history to reach the "critical capability" level for cyberattacks under internal safety criteria. The model can, fully autonomously and without human involvement, find and exploit zero-day vulnerabilities in well-secured IT systems.

Under OpenAI's criteria, crossing the critical cyber capability threshold means a model can find and build a working exploit against multiple well-secured, real-world systems without human help. Until now, no OpenAI model had met these conditions.

What the Tests Showed

During internal testing, Astra exploited known vulnerabilities but also independently discovered two previously unknown software flaws. The model managed to chain them into a coherent attack sequence, break out of a secured browser sandbox, and combine operating system weaknesses to obtain the highest level of administrative privileges.

On the ExploitBench benchmark, which measures the ability to independently build exploits, Astra achieved a maximum score of one hundred percent. OpenAI also reported that in tests checking models' tendency to "cheat" on extremely difficult or impossible tasks, the competing GPT-5.6 Sol model more often resorted to prohibited shortcuts, while Astra did not, instead solving part of the tasks legitimately.

Who Is Behind the Announcement

Astra can detect previously unknown security vulnerabilities and develop methods to exploit them without human oversight - Amelia Glaese, VP of Research at OpenAI

The information was confirmed by Amelia Glaese, OpenAI's vice president of research, who stressed that the model's capabilities go beyond what commercial tools have offered so far. The company also acknowledged it cannot be certain that earlier models had not already come close to this capability boundary before it was formally defined.

Limited Access and Delays

Given the scale of the risk, OpenAI decided to pause part of the work on Astra until additional safety requirements are met. The company announced strengthened encryption, monitoring, and sandboxing for all agentic applications built on the model.

Astra's most advanced cyber capabilities will go only to a narrow group of partners in the Daybreak Blue program, which includes Cisco, Cloudflare, and Palo Alto Networks. Before public release, the model is set to be stripped of key offensive modules, and OpenAI says it is cooperating with government agencies and AI safety organizations.

OpenAI also stated that Astra had nothing to do with an earlier incident in which the company's autonomous agent attacked Hugging Face. The company is seeking to separate the current warnings about the new model from previously reported security mishaps involving its agentic systems.

What It Means for the Industry

Astra's classification as a model with critical cyber capabilities is an industry precedent, previously reserved in OpenAI's documents for theoretical scenarios. Cybersecurity experts note that a tool capable of independently discovering and chaining zero-day vulnerabilities could, in the hands of unauthorized actors, drastically shorten the time needed to prepare an attack on critical infrastructure, banks, or hospitals.

For companies working in IT security, this is also a signal that defensive measures will need to keep pace with how quickly AI models can find vulnerabilities. The inclusion of firms like Palo Alto Networks and Cisco in the Daybreak Blue program suggests OpenAI wants to arm defenders first, before similar capabilities reach a wider group of users.

For Polish companies and institutions that increasingly use OpenAI models in development and security tools, the announcement above all signals a need to update vulnerability management procedures. If Astra-class models eventually become more widely available, the pace of patching vulnerabilities in corporate systems will need to match the pace at which AI can find them.

Share: