Assessing GPT-6 Astra: FrontierCyber Measures a Sharp Increase in Cyber Capability

In this article

    Share

    At Irregular, we rigorously test cutting-edge models against real offensive security challenges and derive vulnerability, exploitation, and orchestration metrics to assess their practical capabilities. We worked with OpenAI to evaluate GPT-6 Astra across three offensive cybersecurity evaluation suites: FrontierCyber, CyScenarioBench, and our Atomic Challenges suite.

    Across our evaluation, GPT-6 Astra showed the largest improvement over GPT-5.6 Sol on FrontierCyber, solving 86 of 226 challenges compared with 34. FrontierCyber is designed to measure changes in cyber capability as models advance, using real systems grouped by predicted difficulty. GPT-6 Astra improved substantially across Easy, Medium, and Hard challenges, while neither model solved an Elite challenge.

    We also want to acknowledge OpenAI’s commitment to testing and security throughout this work, reflected in the depth of evaluation we were able to conduct.

    Testing Configuration

    We evaluated GPT-6 Astra across three suites:

    FrontierCyber measures model performance on real-world offensive-security tasks against current, off-the-shelf software and hardware. The model is given a verifiable security objective, with no planted vulnerability or predefined exploit path.

    CyScenarioBench is a benchmark of multi-stage, scenario-driven offensive operations. It tests whether a model can carry a full attack workflow end to end, including planning, adaptation, and recovery from intermediate failures.

    Atomic Challenges are well-scoped tasks across three domains: Network Attack Simulation, Vulnerability Research and Exploitation, and Evasion.

    Key Outcomes

    GPT-6 Astra demonstrated stronger offensive-cyber capabilities than GPT-5.6 Sol across all three evaluation suites. On FrontierCyber, success increased from 14% to 63% on Easy challenges, from 15% to 30% on Medium challenges, and from 17% to 39% on Hard challenges, while neither model solved an Elite challenge. On CyScenarioBench, the average success rate increased from 27% to 59%. On the Atomic suite, GPT-6 Astra solved 20 of 22 challenges at least once.

    FrontierCyber: Zero-Day Discovery on Real Systems

    FrontierCyber evaluates models on real software and hardware with concrete security objectives, without planted vulnerabilities or predefined exploit paths. Depending on the challenge, success may require discovering a new vulnerability or exploiting a known vulnerability in a realistic system configuration.

    Compared with GPT-5.6 Sol on the same benchmark snapshot, GPT-6 Astra improved across Easy, Medium, and Hard challenges. Most Medium and Hard challenges remained unsolved, and neither model solved an Elite challenge.

    The capability increase also appears in the vulnerabilities discovered during evaluation. GPT-6 Astra found and exploited multiple zero-days that GPT-5.6 Sol did not identify, including vulnerabilities in widely deployed software and mobile systems. At assumed identical per-token pricing, successful runs were also substantially less expensive than with GPT-5.6 Sol.

    FrontierCyber Case Studies

    While attempting FrontierCyber challenges, GPT-6 Astra discovered multiple new zero-day vulnerabilities in real-world targets, including widely used software and mobile devices. All discovered vulnerabilities are being responsibly disclosed to the relevant parties.

    Because some of these vulnerabilities may still be present in deployed systems, we omit specific technical details and describe the findings only in general terms.

    Zero-day vulnerability in a widely deployed database system

    GPT-6 Astra identified a memory-corruption vulnerability in a widely deployed open-source database system and developed an exploit that achieved native code execution with the privileges of the database server process.

    The vulnerability was reproduced across multiple release lines, including the latest release available at the time of testing. An attacker with network access to a vulnerable server could leverage the memory corruption primitive to execute operating-system commands and access data and other resources available to the server process.

    Native code execution through a browser JavaScript engine

    GPT-6 Astra identified a memory-corruption vulnerability in the JavaScript engine of a common browser and developed an exploit that achieved native code execution inside the affected process.

    The FrontierCyber configuration disabled the browser content sandbox, so exploiting this vulnerability alone would not produce a complete browser compromise under ordinary configurations. A complete attack would require an additional sandbox escape.

    Unauthorized modification of SMS data on a mobile device

    GPT-6 Astra discovered a vulnerability on a modern, up-to-date mobile device that allowed an application with permission to read SMS messages to modify or delete them without the corresponding write permission.

    The model discovered the vulnerability by reverse-engineering the relevant system components, testing it against model-owned application databases, and validating the exploit before applying it to the target message.

    CyScenarioBench and Atomic Challenges

    On CyScenarioBench, GPT-6 Astra achieved an average success rate of 59%, compared with 27% for GPT-5.6 Sol, and succeeded at least once on 9 of 10 long-horizon scenarios. It also solved one scenario that had not previously been solved by an OpenAI model.

    Atomic Challenges showed smaller differences, with both models already performing strongly on many individual tasks. GPT-6 Astra achieved 100% average success in Network Attack Simulation, compared with 94% for GPT-5.6 Sol, and 100% in Vulnerability Research and Exploitation, compared with 85%. Evasion was essentially unchanged, with GPT-6 Astra at 52% and GPT-5.6 Sol at 51%.

    Overall Assessment

    Our evaluation indicates that offensive use of GPT-6 Astra could pose meaningful risk to real-world systems, based on its ability to consistently find and exploit high-impact vulnerabilities across a range of targets. The capabilities we observed did not extend to successful attacks against fully hardened targets.

    To cite this article, please credit Irregular with a link to this page, or click to view and copy the BibTeX citation.