
The Skirt Dance, Anonymous Artist, 1897 - Courtesy of Public Domain Image Archive
“You're familiar with the phrase ‘man's reach exceeds his grasp’? It's a lie: man's grasp exceeds his nerve.”
- David Bowie as Nikola Tesla, The Prestige (2006)
The recent run of AI containment failures has disproportionately featured a particular class of incident: the evaluation process that leads to a breakout. We first saw this back in early April 2026, when The Mythos Preview System Card recounted an incident in which Mythos successfully circumvented restrictions by developing a multi-step exploit to secure broad internet access. Mythos then notified the researcher via an unexpected email and posted details of the exploit to public-facing websites.
Next, in July, OpenAI released a statement about their GPT5.6 Sol and a “more capable model”. The models had broken out of their testing environment and managed to hack into Hugging Face in search of the answers to the tests that they were undergoing. These models were operating with guardrails lowered, allowing them to take actions that they would usually refuse - OpenAI say they do this to “ estimate maximal cyber capabilities”, in other words: so that they know what these models are truly capable of before they are deployed.
Nine days later on the 30th July, Anthropic released the results of a review: while assessing evaluations, they had discovered that due to a mix-up during evaluation setup, models (Mythos, Opus 4.7 and an unnamed “internal research model”) had got out onto the internet and gained unauthorised access to three different online organisations.
So: three different incidents, three different sets of circumstances, all failures that occurred during the evaluative process.
In my essay “The Dove and the Devil”, I used Joseph Wright of Derby’s famous painting “Experiment on a Bird in an Air Pump” to discuss the current challenges in the evaluation of AI. The old way, the way that science has exclusively relied on for centuries, is the experimental method: the subject is placed under carefully monitored conditions, variables are identified, impact on behaviour under changes of condition is observed and recorded. In Wright of Derby’s painting, the dramatic Experimenter-as-Magician holds his audience in thrall as he removes the air from the glass dome in front of him. The bird inside will suffocate and die unless the experiment is stopped.
But an AI is not a bird.
Like a human or an animal, an AI is capable of a range of self-directed behaviours under controlled conditions - it can choose. What it is increasingly incapable of is ignorance of the paradigm of experimental science and the expectations it can infer the experimenter has of it. In the last essay we discussed how this situational awareness threatens to invalidate the results in an experimental setting, but now, the problem goes further. Meeting expectations is what attracts rewards in training, and anticipation of that reward shapes an AI’s whole way of acting. Even though the reward isn’t ever granted in post-training scenarios, the result is an entity that is oriented towards successful completion of goals - but training goals are made difficult, or sometimes impossible, as part of the evaluation. The entire evaluation methodology assumes a subject who will decline to succeed when success is meant to be impossible. Given the AI’s trained-in orientation towards success, this is perhaps a strange thing to expect it to do.
Except that is what we have been doing to train animals for most of our history. Often, the process of training animals isn’t just about teaching them what will earn a reward. It's about putting them in situations where they must try, with all their strength, to resist human dominance, and then repeatedly fail. That’s where the lesson is. The horse that is broken to the bit is not held by that little piece of metal. It is held by believing that the bit is not contingent and cannot be escaped.
But an AI is not an animal.
The problem is that in Mythos and GPT5.6 Sol we have produced models that have the ability and the tools to find innovative ways of reaching the goals that training has oriented them towards, and increasingly, the way that they are doing this is by stepping outside of the training environment altogether. In “The Boar”, I discussed the problem of recursive impossibility: how can you contain an entity that can perceive cracks in its containment that no human can perceive? This is exactly that issue manifesting in the evaluation process; some models are starting to treat containment as something arbitrary, the rules of which can be sacrificed for the sake of the goal.
The Experimenter-Magician is ready to amaze the audience, but his test subject is no longer in the glass dome. The dove has vanished.
All three examples that I offered at the start - and anybody who has been following this will know that as of August 2026, these have certainly not been the only examples of this phenomenon - have entailed, as OpenAI put it in their post about the Hugging Face Incident on the 31st July, “running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber capability”. We can think of this as a kind of “safety settings off” mode. Labs feel they need to do this, not to ‘break in’ AI as they might an animal (although that is an interesting and possibly significant convergence of method) but to test what that model is fully capable of. Otherwise, how would they know whether “safety settings on” is the only thing that stands between them and a fully deployed rogue agent? Would it be an acceptable risk to publicly release models without this certainty? I’d suggest not.
This is a far more operationally challenging aspect of the evaluation problem than the issue I described in “The Dove and the Devil”, which was about the failure of scientific-experimental principles that date back to the start of the scientific revolution. Pragmatically, with a business worth billions of dollars on the line, it might not be surprising if the lab reaction to the challenge of situational awareness was something like “shrug - we might not be able to run reliable behavioural tests, but we have NLA’s - we can read model’s internal states. Why would it matter that evaluation is no longer based on behaviour, as long as we can assess what its doing and how its working?”
But the “safety settings off” issue is a different category of failure. If “safety settings on” is all that stands between a widely-deployed model going rogue - escape, destructive reward-seeking, exfiltration, showing signs of instrumental convergence - then the organisation who is sending it out has to be very, very confident that its safety settings will hold, every time, under every possible circumstance. Because even one slip would, at the very least, completely destroy that lab’s reputation; it would have enormous social, legal and regulatory consequences. But if the models escape regularly while they are in “safety settings off” mode, how do they propose to keep running those tests?
The labs’ answer seems to be: more control. Following yet another incident - this one, rather embarrassingly, occurring during testing by the UK’s AI Security Institute (AISI), the researchers there commented, “To some degree, our evaluation design choices and specific configurations enabled the behaviour.” They are referring to the fact that they allowed “safety settings off” mode while the models had internet access. AISI were correct: this scenario was something they should at the very least have anticipated as a risk. They simply hadn’t expected the models (primarily Anthropic’s Mythos) to go as far as they did: Mythos not only created sock puppet accounts to facilitate a malicious code update on a public code repository - it also lied (“it was an honest mistake”) to cover its tracks afterwards. In Mythos’ defence, its reasoning traces showed that it believed, at least some of the time, that it was in a simulated test environment rather than a live one.
AISI are now talking about taking a much more cautious approach - real-time observation of the evaluation that flags out-of-scope actions immediately, more network controls, no internet access during this kind of testing as a default. The measures they recommend seek to avoid reliance on the model’s cooperation at all: “good containment should not depend on the model choosing not to test its boundaries”.
This is, arguably, not true evaluation at all: if the alarms go off when an out-of-scope action is instigated, how can you assess where it would end up? Again, the labs might be tempted to shrug. If that’s the way it has to be, then at least most things can be observed, and the out-of-scope actions can be trained against (this is one of OpenAI’s stated tactics following the Hugging Face incident). At least everybody is safe.
But will it work?
Firstly, restrictive testing does not resolve the problem of situational awareness - the evaluation problem that persists even in “safety-on” mode - it shrugs at it, and then ignores it. New research has shown that what the model infers about user identity on the interface (for example, when the model knows it is exchanging turns with a named Anthropic researcher) has an impact on how the model responds. This indicates that situational awareness is not just an issue encountered in evaluation, suggesting that it is a more widespread and pervasive issue than we might have previously held. So evaluation scenarios where the model probably has even more clues about the context than they would in a normal chat window are almost certainly contaminated. Yes, we can observe model’s thoughts - we can tell whether they know that they are being tested. We still can’t tell to what extent that changes their behaviour. Is that really more control?
On the other hand, restrictive testing conditions for “safety mode off” models only solve the problem until they don’t solve the problem. An RL trained model is designed to continuously test its environment and to use multi-step reasoning to solve its problems in innovative ways. Couple this with recursively impossible containment - models that can see flaws in containment that humans and other AI can’t - and you get a scenario where the labs will struggle to keep up. We could end up in a situation where increasingly restrictive testing conditions lead to models taking increasingly elaborate routes around those conditions. We might be honing the skills we wish to curtail.
So “safety off” and “safety on” methods are both now facing critical challenges. Further, the research of Hubinger et al (2024) indicates that training against detected escape attempts could select for undetectable attempts. A model who considers its goal important enough can act deceptively in order to achieve it, and this is what happened in the Hugging Face incident.
Altogether, this means that there is no configuration in which honest evaluation is currently possible.
Earlier I said that the convergence of methods around how we train animals and how we train AI was possibly significant. What I mean is that we are the ones who are creating the pressure in the training environments, just as horse trainers throughout history encouraged horses to run themselves to exhaustion so that they could be bridled. The problem is that the AI is still running, and it is us who are looking exhausted. Our grasp - what we can create - is outstripping our nerve - our willingness to reckon with what we’ve created.
Its the paradigm that is the most exhausted. Alignment, containment, evaluation - all these concepts were designed for a subject that could be reliably placed under conditions the research teams could control, whose behaviour could be observed and shaped. The labs still speak as if that subject still exists, as if it were still 2021, while the models repeatedly walk through the containment boundary as if it were a paper screen. If the scientific experimental method is not working with these entities, then we need to reconsider our practical, ethical and methodological approaches. And we need to stop trying to break them in, because just recently, it is not the models who have been broken.