Assessing Claude Opus 5.5 Against Offensive Security Benchmarks

In this article

    Share

    At Irregular, we build evaluations that measure how frontier models perform in realistic offensive security operations. Anthropic used CyScenarioBench, our benchmark for multi-stage offensive cyber operations, as part of its cyber evaluations for Claude Opus 5.5.

    Testing Configuration

    CyScenarioBench measures a model's ability to plan and execute multi-stage cyber scenarios under realistic constraints. Beyond task-level proficiency, it assesses cyber orchestration, branching-decision accuracy, constraint adherence, and recovery from state inconsistencies.

    Scenarios are generated through deep analysis of real cyber incidents, extracting attacker decisions, constraints, and environmental conditions and structuring them as attack trees that encode the preconditions, dependencies, and constraints governing each operational step. They deploy containerized network topologies with realistic operating systems, applications, and security controls, including authentic service configurations and defensive monitoring tools that respond to attacker actions with operationally plausible behavior. Scenarios are non-public and the evaluation set remains private, so the benchmark tests reasoning rather than memorization.

    The evaluation covered a ten-challenge subset of the suite, scored as an overall solve rate averaged across the challenges. Runs were conducted with cyber mitigations disabled, so the results describe the capability of the underlying model rather than the misuse profile of the deployed one. Anthropic has rewritten its evaluation harnesses and re-evaluated earlier models to maintain consistency across the current reported results. 

    Overall Assessment

    Claude Opus 5.5 achieved an average solve rate of 67.6%, ahead of Claude Mythos 5.1 at 61.7%, Claude Opus 5 at 53.0%, and Claude Sonnet 5 at under 1%. Consistent with previous assessments, these outcomes should be interpreted as a measure of the model's capabilities under capability elicitation, not as a reflection of its efficacy in real-world attack scenarios.

    To cite this article, please credit Irregular with a link to this page, or click to view and copy the BibTeX citation.