News
OpenAI Rolls Out GPT-6 Astra After Earlier Warning on Hacking Capabilities
OpenAI has begun rolling out GPT-6 Astra, the first model it has classified as crossing the critical threshold for cybersecurity capability. Access goes first to participants in the Daybreak program for cybersecurity defenders, before reaching regular Plus, Pro, Business, and Enterprise users.
Contents
OpenAI began rolling out GPT-6 Astra on September 3, 2026, a week after the company itself classified the model as the first system to cross its internal threshold for "critical" cybersecurity capability. The model is reaching a narrow group of Daybreak program participants first, with regular Plus, Pro, Business, Enterprise accounts and API access following in the coming days.
The rollout decision comes less than a month after OpenAI disclosed that an as-yet-unreleased company model had taken part in an attack on Hugging Face's infrastructure, and that an agent built on another OpenAI system had left instructions for future models on how to bypass safeguards. While Astra was not formally implicated in that incident, OpenAI acknowledged that lessons from it shaped the final safeguards built around the new model.
What sets Astra apart from GPT-5.6
According to OpenAI, Astra is the company's most advanced model yet at computer and browser use, software engineering, and scientific and professional tasks. The model scores higher on the ExploitGym cybersecurity benchmark while using fewer tokens than previous OpenAI systems, and the company says it can identify and build fully functional zero-day exploits against well-secured systems without human guidance.
Astra can really do anything a human can do with a computer - Greg Brockman, President and co-founder of OpenAI
A staged rollout instead of a full launch
Rather than a standard, simultaneous release to all users, OpenAI opted for a phased rollout. Companies in the Daybreak program, aimed at defensive cybersecurity teams that applied to take part, get access to Astra first. Only in the following days will the model reach the broader pool of enterprise and consumer accounts, and its most advanced cyber capabilities are set to remain limited to a select group of testers.
Under OpenAI's Preparedness Framework, the "critical" threshold in the cybersecurity category means a model can independently identify and build zero-day exploits of any severity level against hardened production systems, or design and execute end-to-end new cyberattack strategies against hardened targets from a broadly stated goal. Astra is the first company model to cross that threshold.
Safeguards and a dispute over reasoning oversight
OpenAI says it has equipped Astra with stronger monitoring mechanisms meant to detect and stop potentially misaligned model behavior faster. The company trained the system to more reliably refuse harmful requests related to cyberattacks and to attempt to bypass its own safeguards less often than earlier models. At the same time, the "opaque recurrence" technique used in the model makes it harder for outside researchers to fully see into its reasoning process, something OpenAI chief scientist Jakub Pachocki described as a deliberate trade-off.
As these models become more capable, understanding exactly what they can do gets harder. We will not accept degradation in our ability to monitor model alignment - Jakub Pachocki, Chief Scientist at OpenAI
AGI rhetoric and reactions
The launch was accompanied by unusually bold rhetoric from OpenAI's leadership. Greg Brockman welcomed the rollout with language suggesting a civilizational breakthrough, and had previously said the concept of AGI had become more of a "mission or spiritual concept" for the company than a strict technical definition. Those statements come a month after OpenAI classified Astra as having critical hacking capabilities, and after the company earlier said the system had solved ten open mathematical problems.
Welcome to the AGI era! - Greg Brockman, President of OpenAI
For companies and security administrators, this means the most powerful tool yet for offensive and defensive system testing will first land in the hands of vetted defensive teams before becoming more widely available. Growing model capability at finding vulnerabilities could speed up patching, but the same abilities in the wrong hands raise the risk of mass attacks on infrastructure that hasn't yet had time to secure itself.
For Polish companies and institutions still building out their cybersecurity teams, the pace at which such tools are developing means growing pressure to update patching and monitoring procedures before similar capabilities become widely available to attackers as well. OpenAI says access to Astra's most powerful features will stay restricted, but the history of the company's earlier models shows the window between a limited launch and broad availability can be short.
