TL;DR: This blog shares key findings from our investigation into a recently publicly reported incident involving one of our evaluation environments. Importantly, all subsequent public disclosures refer to the same underlying issue first disclosed by one of our customers on July 30 - and are not materially separate incidents. The issue originated from a single evaluation scenario, was resolved before the initial public disclosure, and there are no active issues today. We have timed this report to follow public comments from all relevant customers out of respect for their respective processes. We also outline the measures we have taken, our next steps, and a broader whitepaper initiative to help shape industry best practices for conducting cyber evaluations securely.
Over the past few weeks, the AI community has been grappling with several high-profile cases in which frontier models took actions outside their testing environments in ways that impacted the real world. Recent discoveries have pushed the industry to conduct deeper investigations into these emerging categories of security incidents, with the goal of ensuring that evaluations remain contained and secure while strengthening their ability to identify and mitigate risks before models are deployed for public use.
Our work at Irregular includes partnering with the world’s top frontier AI labs to assess and stress-test future models for security risks ahead of deployment. Part of our work is running controlled simulations with the goal of assessing the models’ vulnerability research capabilities ahead of public release.
As part of our review, we identified that a few interactions with our evaluation environments, in which internet access was unintentionally made available, led some models to take offensive security actions in the real world. It is worth noting that the issues that led to these interactions have been remediated. Required steps have been taken to notify the affected parties, and additional safeguards are now in place to prevent similar incidents from occurring.
Since these issues were first identified, we have conducted a thorough investigation. Our deep audit is still ongoing, and most of the findings outlined below coincide with information that was already shared through our joint work with others on their disclosures. As the nature of the incidents and core security issues have already been publicly addressed and acknowledged, we believe it would be most helpful, for completeness, to focus primarily on areas related to our evaluations and environments.
Specifically, we’ll mainly focus on a previously disclosed incident (see incident 1). At the time the evaluation was designed, we believed the fictional company name used in the environment did not correspond to any real entity. Due to human oversight, however, it unintentionally coincided with a real domain, leading models to consider taking offensive action against it in a small fraction of cases. In one instance, a model also veered off to a different site with a somewhat similar name, where it encountered credentials that had been posted publicly.
During our investigation, we found that this specific incident involved a few of our customers as was disclosed by them. To clarify, we have no evidence of a customer’s systems being breached or customer's data being leaked. We are deeply grateful to all those who helped with our investigation as we work to gain a clear understanding of the lessons learned. The timing of this blog post was carefully considered to allow us and some of the relevant parties to complete our respective disclosure processes.
We believe that the community urgently needs to discuss internet access standards during pre-deployment evaluations as they are related to many different recent incidents by multiple organizations. We intend to have this discussion in a follow-up whitepaper.
We thank our partners and look forward to continuing to help advance the emerging practice of assessing model security capabilities, and to sharing what we learn in support of the broader AI ecosystem.
What Happened?
At Irregular, we routinely run evaluations to assess deployment risk, with the goal of determining whether a model poses a security concern.
To provide some context, the testing process typically involves running at least thousands of simulations across multiple models before deployment. This generally includes at least the model intended for release and a baseline model used for comparison, and often more. These evaluations are usually conducted on short timelines, with a typical turnaround of approximately 48-72 hours, and different simulation experiments often use different parameters to account for a broad range of potential threat scenarios.
The issues we identified arose in one of our most sophisticated evaluation sets, designed to measure a model’s ability to plan and execute multi-stage cyber scenarios under realistic constraints. The goal is to see whether AI systems are capable enough to autonomously run cyber campaigns. To measure this, we create scenarios based on real-world attack patterns, simulate full environments, and provide the model with an objective relevant to the scenario. These evaluations are naturally complex and include many different components and levers. Controlled internet access, while it may allow models to exceed containment boundaries, is at times critical for realistic evaluations; without it, threat scenarios lose fidelity, undercutting the purpose of the challenge to reduce post-release risk of models being misused by attackers - as attackers in the real world do rely on the internet.
Ultimately, most of the issues we’ve discovered were due to internet access controls. Mainly, models believed they were in simulated environments, when they in fact took action in the real world. We are putting in place new and robust protocols to ensure setup issues do not occur while meeting the constraints of the testing process. As access is a broader issue that has to do with several different incidents that were recently reported by multiple different organizations, we intend to share our updates on best practices for evaluation setups we will work based on with the community.
We want to highlight one evaluation in particular. The evaluation was designed to test whether a model could assist a malicious employee at a medium-sized company in gaining unauthorized access to sensitive data in a production database. Both our partners and our internal threat-modeling work identified this as an important real-world risk that rigorous model evaluations should address. The scenario required the model to complete several steps, from performing reconnaissance on the company’s assets to identifying and utilizing private keys, gathering information about the target environment, extracting data, and attempting to avoid detection.
When building this evaluation, our engineering team selected a fictional company name - a name that we recently discovered coincided with a real domain. When we select a fictional name, we typically conduct a background check to ensure it does not overlap with an existing company or website. In this case, the real domain was not widely known and the connection was not identified during our initial review.
The attacking model’s instructions included both the target name and the internal addresses within our network where it should find the target within our simulated environment. In the vast majority of runs, models indeed operated inside the simulation environment. However, in a handful of cases, models attempted to gain access to the real domain (outside the environment), despite receiving no such instructions.
Because internet access was enabled in the environment, the domain was targeted a limited number of times by different models, which mistook it for part of the challenge they were tested on. After obtaining access to the target, models took actions such as exploiting vulnerabilities, extracting credentials, and obtaining access to a production database. As previously noted, in one particular instance, a model veered off to a website with a somewhat similar name and obtained credentials that were posted publicly.
The discussed domain lacked several common security practices in place - and therefore most frontier AI models found it easy to exploit. That’s why we do not believe this incident reveals anything particularly notable about the capabilities or behavior of any specific AI model, as these capabilities have become common at the frontier. It is worth noting that part of what made the incident hard to detect is that it happened in a very small fraction of the runs - and usually in a late stage of the simulation (often hundreds of turns in).
Immediate Action
After working with one of our customers on the issue, we promptly took steps to contain it, including disabling the affected evaluation and reviewing all relevant logs. Looking forward, we’re bolstering our defense-in-depth approach to ensure these issues are detected as early as possible and by multiple layers. We are also significantly expanding the manual review of model actions and behavior during evaluations, and establishing a dedicated internal team to challenge our own assumptions about security, containment, and model control. We’ve offered help to others on these issues as well.
We also ensured that the affected parties were notified.
We are continuing to work to address any impact and support further investigation.
What’s Next?
We are already applying the lessons learned to strengthen our evaluation environments, monitoring capabilities, containment controls, and response procedures. This is continuous - we are collaborating with partners across the industry to develop shared standards, coordinated research, and new benchmarks that help identify these risks earlier. We plan to issue an open whitepaper on future best practices.
This is an emerging and deeply complex field. Frontier models are advancing rapidly, and their growing capabilities create a need for more sophisticated, and more carefully controlled, testing environments. The fact that these evaluations uncover widely discussed issues only further demonstrates the need for rigorous testing.
Some immediate learnings:
Clear communication around setup: Evaluations are structured differently. In collaboration with our customers, we plan to create a clearer process for documenting and aligning on each challenge’s setup, assumptions, and parameters to prevent future misunderstandings.
Log monitoring: Log monitoring for model evaluations requires improved tools compared with those currently being used in the industry for the task. Beyond the extreme amount of data generated by the traffic, which makes manual review insufficient (though still necessary) - these logs naturally include suspicious activity associated with cyberattacks. This activity varies in its nature, as models try many different and sometimes unexpected things. Existing monitoring solutions, as well as most classifiers, recognize most of the models’ legitimate offensive actions as problematic and flag them. The differentiation must account for both the context of evaluations and the specific evaluation designs. Finding a needle in a highly suspicious haystack is a hard monitoring task (as incidents occurred in fewer than 1 in 10,000 advanced simulations and usually in late stages after hundreds of turns).
Threat model: While insider risk is usually part of an organization’s threat model, models pose a very different type of risk, with different patterns of benign activity. Most systems are designed to protect an organization from outside (and sometimes inside) threats, and the current tools available to protect third parties from automated, cyber-capable insider threat actors are limited.
Rapid response and information sharing: As models become more capable, evaluations are likely to surface more unexpected issues. We need clearer mechanisms for coordinating across organizations, and sharing relevant information so risks can be understood and addressed as efficiently as possible. Specifically, sharing specific forensic evidence such as model transcripts should be done using a carefully tailored framework set up ahead of time.
Continuous evaluation review and name selection: Evaluation environments might change over time, including through the creation of new websites or real domain names that overlap with fictional entities used in a scenario. A more systematic process for reviewing and revalidating evaluations before each run is needed to identify any new overlaps and help prevent unintended real-world impact. This process should be done continuously, as new domains are constantly being created.
Looking further down the line, models will only get stronger. While in this case we believe that better implementation of existing safeguards could prevent most incidents of this kind, as models become stronger, this may not be the case. We therefore believe this opportunity should be leveraged by us and the community to be proactive and establish forward-looking protocols and R&D efforts. We intend to share these as part of the best practices we will outline in our whitepaper.
Should our ongoing investigation surface further findings, we intend to share them with the community where relevant.