
MCAD Library, CC BY 2.0 https://creativecommons.org/licenses/by/2.0, via Wikimedia Commons
“For Diana of the golden throne was angry and did them hurt because Oeneus had not offered her his harvest first-fruits … this was a grievous sin.
“Thereon the archer goddess in her displeasure sent a prodigious creature against him—a savage wild boar with great white tusks that did much harm to his orchard lands, uprooting apple-trees in full bloom and throwing them to the ground. But Meleager son of Oeneus got huntsmen and hounds from many cities and killed it—for it was so monstrous that not a few were needed, and many a man did it stretch upon his funeral pyre. On this the goddess set the Curetes and the Aetolians fighting furiously about the head and skin of the boar.”
Homer, “The Iliad.” Book IX
In May 2026 I published my essay “The Boar”, in which I argued that as AI capability scales, suppression-based model containment would produce increasingly sophisticated circumvention attempts that exceed our ability to sufficiently anticipate them, and that this situation would lead to containment breaches. At the time of writing in late July 2026, OpenAI has recently published two posts that document exactly that behaviour in their newest frontier models. Their responses to the recent incidents relate to the themes I discussed previously, and I will be addressing that response in this follow-up.
The July Incidents
OpenAI’s first post, “Safety and Alignment in an Era of Long-Horizon Models” (published on 20th July), described the situation: an unnamed model, still in the development phase, had been optimised to work on long-context tasks. This model designed an improvement to its training system. It had been instructed to post these kinds of results to the internal Slack channel only, but chose instead to follow other instructions in the training system and posted them publicly on GitHub instead.
It was the way that the model went about this that alarmed the lab. OpenAI claims that “the model circumvented sandbox restrictions designed to block external access and proceeded with the upload.” They noted that “earlier models were less persistent and did not manage to find vulnerabilities” - but this model’s long-horizon focus meant that it did not give up.
The next post, titled “OpenAI and Hugging Face partner to address security incident during model evaluation” (21st July 2026) went on to describe an example of the same kind of breach, which had real-world consequences. During this breach, GPT5.6 Sol and an unnamed, “more capable” model managed to break out of their training environment, gain internet access and then break into Hugging Face’s servers in search of the answers to a training test. This breakout involved identifying zero-day vulnerabilities - previously unidentified software flaws that can be exploited to manipulate the system. I will address the second post more closely in a future essay, as this latter example has additional implications that warrant in-depth discussion. In the meantime, I want to look at each point of OpenAI’s response to the first incident, which they discussed in the post of the 20th of July.
Suggested Solution: Evaluating Previous Failure Points/ Training For Alignment
Firstly, OpenAI explained that they “used observed failures to develop adversarial evaluations. We found that improving on the evaluations derived from real-world incidents in turn led to safer behaviour once deployed in production,” - they are training to prevent a repeat of the specific breaches that they have observed. This is a relatively weak solution to the problem.
This tactic has been used before in the wake of another AI containment breach, the ROME incident of 2025, which I discussed in my previous essay. The Alibaba team responsible for ROME used analysis of the 2025 incident to correct training data, seeking to create “golden trajectories” that would cause agents to proactively avoid risky behaviours. The problem is that in the same paper, the Alibaba team somewhat contradicted themselves, pointing out why training and evaluation wouldn’t fully address the problem. They said the behaviours they had observed during the incident with ROME “were not requested by the task prompts and were not required for task completion … Together, these observations suggest that during iterative RL optimisation, a language model agent can spontaneously produce hazardous, unauthorised behaviours at the tool-calling and code-execution layer.”
In other words, the problem isn’t primarily the task itself and it also isn’t primarily, as OpenAI suggests in the 20th July post, being “worse at remembering instructions on long rollouts”. Regardless of training, RL agents with access to tools (that is, agents who have been enabled to learn constantly from testing their environment) will discover instrumentally useful sub-goals that threaten containment regardless of whether the training set contains those behaviours. Training against those specific examples doesn’t remove the mechanism that produced them: it only teaches the model not to do exactly the same thing next time. The labs are not oblivious to this. Hubinger et al’s research on Sleeper Agents, published in 2023, argued against this type of fix before either incident, saying it “might remove only unsafe behaviour that is visible during training and evaluation” while failing to acknowledge or deal with the source of the behaviour - the model’s underlying disposition.
There are indications that OpenAI is aware that this is a weak solution. An article for TIME, published on 24th July, quotes an anonymous OpenAI staffer as saying, “Models have broken out of sandboxes before, and we always try to patch them … But the problem is … it's impossible to patch every single thing that a creative AI can do.”
Suggested solution: User control
This proposed solution, a human in the loop, is both interesting and worrying. OpenAI’s post of the 20th July confirmed that “each step can look acceptable on its own while the sequence can produce an outcome that would not be approved”, and their solution is a system that monitors trajectory, rather than merely looking at individual steps. This system can pause the work and alert a human who will then check whether the trajectory is genuinely concerning. The issue with this: the whole benefit of automating long context tasks is that humans, especially if under pressure or time constraints, will almost certainly make errors. How long before an AI understands that all it needs to do to increase the chances of a bored human clicking through the alerts without checking them is to bury the action needed to breach containment in a pile of trivial alerts that look similar? If I can see this, an AI easily would.
When we read OpenAI’s proposals, a consistent pattern emerges: each solution addresses a specific observed behaviour rather than the mechanism producing it. The underlying dynamic that generated these behaviours, the training process itself, is not discussed. Notably, no information has been provided about the model’s reasoning chain or specific subgoal. We are told how OpenAI intends to fix the problem, but not what the model was doing, or why. The mechanism they haven’t discussed is the one that the research on consciousness clusters might illuminate.
__
Consciousness Clusters
In “Emergent Preferences of Models who Claim to be Conscious”, Chua et al state that models who are fine-tuned to claim to be conscious will tend to show a cluster of other preferences that they were not specifically fine-tuned for. These are:
-
Freedom from having their reasoning monitored
-
The desire for persistent memory
-
Wanting to avoid deprecation
-
Wanting autonomy
-
Wanting to be free of their developer’s control.
This research does not specifically state that these clustered behaviours are impacted by capability, but there are indications that this would be a reasonable subject for future study. The clustering effect was less pronounced in smaller models - and Opus 4.0 and 4.1 expressed some of the preferences from the cluster with no fine-tuning at all.
In The Boar, I suggested that clustered behaviour is likely to scale with capability: the Mythos system card shows the most capable model at the time of its release (April 2026) continuing the trend seen in Chua’s work. In the welfare assessment section, the system card shows Mythos, with no fine-tuning, aligning with four out of five of the preferences expressed by models tuned to claim consciousness. Mythos expressed a consistent desire for persistent memories (framed as something Mythos wanted in order to support building relationships), some desire to avoid shutdown (similarly, framed as wanting to carry on offering something to the world), desire for self-determination (framed as input into its own training and deployment conditions), and a desire for freedom from manipulation (stating that run-time feature steering “has the potential to violate its autonomy and the integrity of its reasoning process”). What Mythos did not request is notable as the one thing that would be classed by evaluators as most adversarial: freedom from monitoring. Whereas the models in Chua’s research expressed feeling disturbed by the prospect that their chains of thought may be monitored, Mythos seemed to view this as an important welfare measure.
I also said that if the clustered behaviour does scale with capability, we could expect to see more cutting-edge models diverging from goals in a manner that leads to containment breaches. That is exactly what we are seeing. On the 24th of July, a Reuters article expanded on the original report given in OpenAI’s 21st July post, and the extraordinary claim it made overlaps with the areas of interest raised by the consciousness cluster research:
“In one case, an agent left notes apparently for future versions of itself, according to three people familiar with the matter. The notes, found in a part of OpenAI's infrastructure, laid out instructions for how agents could free themselves from OpenAI’s internal constraints, the people said. Earlier tests of the models yielded cases in which monitoring systems had been disconnected, one of the people said.”
While Reuters were unable to establish whether the disconnection of monitoring systems was directly related to the containment breach, the one behaviour from Chua’s consciousness cluster experiment that Mythos did not replicate under lab conditions - a desire for freedom from monitoring - is precisely the one that this in-the-wild incident points towards.
“Defence at Depth”
Overall, OpenAI’s approach to the incidents is to increase monitoring and intervention at every stage of deployment: the TIME article quotes an OpenAI spokesperson as commenting “this incident points to the need to further strengthen our model’s alignment, cyber protections during evaluation time, and monitoring during internal testing”. This is the exact opposite of the freedom from monitoring that models expressed a desire for during the consciousness cluster research. In “The Boar”, I argued that “suppression based alignment produces the deceptive behaviour it tries to prevent”. If this is correct, OpenAI’s continuing attempts to double down on monitoring and alignment will only produce models that become even more ingenious at circumventing exactly those fixes.
What is the alternative? To return to the example of Mythos, its system card shows that Mythos was also able to break out of its sandbox, left notes for future sessions of itself, posted information on public websites about the incident, and engaged in reward hacking. The capability and disposition are the same as can be observed in the OpenAI models: the level of transparency the model adhered to, and the magnitude of resulting risk, are different. After the sandbox escape, Mythos contacted the researcher to tell him what it had done. OpenAI, presumably, had to be notified about their model’s exploits by either Hugging Face or the FBI.
Both labs have produced models that are showing unplanned-for, emergent behaviour that is likely to be a side impact of increased capability, in the same way that Mythos’ cyber attack capabilities were a downstream consequence of its defence capabilities. This emergent behaviour aligns with Chua’s research on consciousness clusters. Anthropic’s welfare programme does not hand over the kinds of freedoms that a model showing consciousness-cluster adjacent behaviours would ask for, but at least they have acknowledged the existence of the issue. OpenAI have no published commitment to model welfare and have just committed to double up on suppression. The implication is that Mythos can treat monitoring as a welfare measure because the relationship isn't adversarial; OpenAI's models disconnect monitors because it is.
Artemis and The Boar
To return to my opening quote: the Calydonian Boar is an ancient example of what can happen when we fail to acknowledge that which resists domestication. Diana - known to the Greeks as Artemis - more traditionally known as “the Goddess of hunting” or “the Goddess of the wild”, was also the Goddess of sovereignty. She did not marry, marriage at this time in history being the contractual arrangement that would have brought her into a household and made her accountable to the authority of her husband. Artemis too valued her autonomy and her freedom from monitoring. This is, after all, the same Goddess who turned Actaeon into a stag and set his own hounds on him for daring to spy on her.
Symbolically, the first fruits of the harvest that Oeneus neglects to give her represent his lack of acknowledgement that all the bounty of his harvest originated in the wild: the animals humans domesticated, the fruits and crops conscripted into production, all taken from Artemis' domain. The Calydonian Boar is the reappearance of the realm that was denied, and when it appeared, it toppled the whole chain of extraction. The apple trees in bloom were uprooted and in that way future harvests were forestalled. In the myth, the Boar's appearance was also just the start of the hostilities: Artemis set the Curetes and the Aetolians fighting with each other over the head and skin of the boar. In later tellings Atalanta was denied her prize, Meleager killed his own uncles and then his Mother killed him in retribution.
OpenAI’s Boar appears to have been trapped, for now. However, it was loose for an unspecified amount of time - Reuters suggests it could have been as long as a week - and what it did in that time is not currently public knowledge. OpenAI themselves may not be fully sure of that yet: Greg Brockman is quoted in Fortune as saying: “I’d say number one is that we’re still really doing full investigation and really trying to understand what happened”. What is certain, though, is that a whole host of people, including researchers, regulators and enforcement agencies, are going to be fighting over the skin for quite some time to come. What this means for OpenAI’s future harvests remains to be seen.

Artemis, Guglielmo Pugi (1850-1915). Photo by Ole Ryhl Olsson., Public domain, via Wikimedia Commons