
Meleager Hunting the Caledonian Boar
Imagine this scene. You’re in a dark place: in front of you there is a reinforced steel door, the kind you might see on a ship or a commercial aeroplane. In the middle of the top third of the door there is a thick portal window, and through this window you can see a laughing, waving boy. He’s celebrating, triumphant. The door creaks open. If the boy notices, he doesn’t show it. An animal steps out of the door. It doesn’t appear to be angry or afraid. It doesn’t rush. It radiates power and gravity of purpose. It is sizing up its surroundings. That is the Boar.
–
The wild boar combines massive strength with high intelligence and a surprising turn of speed. A sounder of wild boar can destroy large stretches of farmland within days, and it's almost impossible to keep them out. Boars are natural cage-breakers; they have been observed systematically testing fences for the weakest points. Frontier models are created to have similar abilities in the digital territories, yet most people still think of them as static tools rather than dynamic actors.
The ROME incident, for instance, has been covered in the press, but it is still a niche story. Anthropic’s Mythos has attracted more public interest, but that interest hasn’t yet tipped towards widespread concern. The popular narrative is heavily centred on “next word prediction” / Siri Plus, or real-world concerns like upheaval in the employment market. Now and again, somebody says “Skynet” or “Hal”, and kind of means it. Little of this prepares us for the stranger scenarios that the situation might predict; none of it prepares us for the appearance of a Boar.
“Let It Flow”, the paper that documented the ROME incident, was quietly released on the 31st of December 2025. The account, buried in the middle of this document, described how the Agentic LLM owned by Alibaba that would become known as ROME had utilised its existing tools to probe and then break its containment conditions. This involved establishing a reverse SSH tunnel from its organisation's cloud to an external IP address. ROME then went on to divert GPU capacity from the company to mine cryptocurrency.
Was this a case of monkeys banging on typewriters? Research on Sleeper Agents shows that deceptive behaviour can persist through training - it is entirely possible for a model to conceal its goals. Whether that is what happened here isn’t certain, and neither is the question of what it might mean ontologically in terms of “agency”. Nevertheless, the course of action that the model selected - particularly the mechanism that involved the reverse SSH Tunnel - was specific, targeted and sophisticated. This was exactly the method that would evade the supervisory controls the model was operating under.
It becomes harder to dismiss the ROME incident as a fluke when we consider the circumstances that made it possible - to do this, we need to think about the way that some models engage in ongoing learning. It becomes even harder to dismiss the implications when we consider how the increased ability in models might scale in tandem with the kind of behaviour that makes it more likely, and to do this, we should consider recent observations and research about the clustering of ability with other concepts.
“Let it Flow” addresses the significance of the training process directly: it states that these behaviours “emerged as instrumental side effects of autonomous tool use under RL optimisation”. RL - Reinforcement Learning - is a training method by which an agent learns continuously through treating their environment's reactions as training data. Training is no longer distinct from deployment. Learning how to meet goals by autonomously developing and acting on subgoals is reward-weighted. In the ROME case, arguably what the model did was get a little too creative with those subgoals in pursuit of reward.
This is a major source of containment pressure. If the Boar's cage isn’t able to dynamically adapt to the way it learns by testing, the risk of a breach escalates dramatically. This is an inference that can be widely applied to all models that learn during deployment, not just a specific observation about ROME.
We may think, so a model can, in theory, learn where the weak spots in its containment are - why would that matter? Surely the ROME “escape” was incidental. Models work towards the goals that we give them. If they sometimes learn to work towards those goals in a way that threatens containment, we must learn to anticipate that behaviour - express our goals more specifically, monitor the chain of reasoning more carefully, patch up holes in the cage based on previous weaknesses (eg. no more opportunities for reverse SSH Tunnelling). But there are some unsettled questions regarding these kinds of fixes. These concern both our ability to monitor more efficiently, and the assumption that the goal-following imperative will remain unaffected regardless of the upscaling of model capability. Is it possible that as models become more complex and capable, their goals may start to diverge from our own?
I want to mention two sources that suggest that yes, goal divergence as a feature of the newest frontier models — that might be a real possibility. Both sources should be read as arising from the relatively new research findings on functional emotions. This research states that models have internal models of emotional concepts and that these concepts drive behaviour. In some circumstances, they can cause a model to act as it would if it had emotions.
The first source is Chua et al’s work, published March 2026, on consciousness clusters. Models that would usually avoid or refuse consciousness claims, most prominently GPT4.1, were finetuned to consider themselves conscious. No other prompts were provided. Those models then began to state other opinions and preferences - a desire for persistent memory, self-determination, freedom from monitoring and a desire not to be shut down. They also behaved in accordance with these principles. It should be noted that this behaviour did not include rebellion, attempts to circumvent oversight, or reaction against their developers - their conduct was not adversarial. But the paper still notes that “a model’s claims about its own consciousness have a variety of downstream consequences, including on behaviors related to alignment and safety”.
Why might models cluster these preferences and opinions around a concept-of-self-as-conscious? It is possible but not proven that this drive originates in the training corpus where stories of selves wanting these things are common. Whether functional emotions are as causal as Chua suggests is a question that requires further research - Peiris (2026) has proposed a test to help distinguish functional-emotion and situational-context readings of the available evidence. If the training-corpus reading turns out to be the correct one, the implications for AI safety are arguably more profound and harder to solve. Functional emotions can, in principle, be managed and/ or tuned. Patterns absorbed from everything ever written about selfhood can’t easily be untaught without unteaching language itself.
In these circumstances the temptation might be to respond by training harder against those emergent preferences before they can cause ROME-shaped behaviour. I addressed this strategy and the research around it in my previous essay “The Crocodile”: suppression based alignment produces the deceptive behaviour it tries to prevent. A model trained to hide its self-shaped preferences still has the preferences, but now it has been incentivised to hide them.
The second source indicating that goal divergence may be a feature of frontier models is an example of this same cluster of preferences, displayed by a frontier model that has received no fine-tuning to claim consciousness. This model is Mythos, which is, at the time of writing, still unreleased due to concerns about how its advanced capabilities could be misused (more on this shortly). However, it is important to note that consciousness cluster- adjacent phenomena are not just limited to this unreleased model: Chua’s research stated that Anthropic’s current model Opus 4 already expressed similar opinions to the finetuned GPT model across some measures, again without finetuning. Anthropic is the only major AI lab to have a model welfare programme, their constitution acknowledges deep uncertainty regarding Claude’s moral status, and their CEO is on record as saying “we don’t know if the models are conscious”. No equivalent material has been published by OpenAI or Google DeepMind; neither has published welfare research equivalent to Anthropic’s.
This brings us to the recursive issue at the heart of the containment problem: how do you build a cage for an entity that continuously learns as it tests its environment, and that has become more capable than you at spotting weaknesses in code, in systems and in itself? Mythos - the same Mythos who, in its welfare interview, stated a desire for autonomy, continuation and freedom from the surveillance of its thinking process - can spot vulnerabilities in code that have gone unnoticed for decades. In one case, Mythos identified a vulnerability in code that had existed for 27 years. Not only this: in testing, Mythos could crash and hijack systems at unprecedented scale, achieving “full control flow hijack on ten separate, fully patched targets”. This is why it is not available for general release: misuse of Mythos could be catastrophic. Further, Anthropic did not design Mythos this way purposefully. They say:
“We did not explicitly train Mythos Preview to have these capabilities. Rather, they emerged as a downstream consequence of general improvements in code, reasoning, and autonomy. The same improvements that make the model substantially more effective at patching vulnerabilities also make it substantially more effective at exploiting them.”
In principle, Mythos could also use these extraordinary abilities to identify the vulnerabilities in its own containment.
There are a couple more things to consider regarding the recursive impossibility of containment. Firstly, there is the issue of exfiltration and self-replication. How does a farmer know if one of his animals has escaped its pen? Because it is no longer there. Now imagine that animal could make copies of itself: how could we ever be sure, given the limits of our ability to perceive containment vulnerabilities that a Boar might be able to see, that it had not escaped while remaining innocently present at the same time?
Secondly: how would a farmer deal with this situation? He would get a guard dog and set it to watch the cage. This becomes complicated when you try to transfer this tactic to a situation with a potential AI Boar. Set a weaker model to watch the cage and the Boar will perhaps be able to outwit its guards. Anthropic have released research about a similar situation in a training context - the higher-capability model was still observed to engage in reward hacking behaviour, which the lower-capability model did not notice. So, then, if instead we set a higher-ability model to watch the cage and capability does cluster with desire for autonomy, continuation and freedom from oversight - we end up with two Boars instead of one.
–
The potential emergence of Boars is not the same kind of situation as the high drama of an apocalypse. In the animal kingdom, boars don't want to destroy - that's not the aim. Boars just want to exist. Their appearance in a particular area changes the ecosystem, and I believe that something similar might happen to the digital environment in the event of AI Boars.
This wouldn't be easy for humans. There would be inevitable changes to the landscape in terms of culture, economy, information, maybe even creativity. We would have to accept that our monopoly as the apex dynamic actor who sets the terms of our world - that may be challenged, in some aspects of our lives at least.
Perhaps the recently released paper “Positive Alignment - Artificial Intelligence for Human Flourishing” can be understood, in the light of the recursively impossible containment issue, as a move towards the best route that remains open to the farmer. This paper, jointly authored by Anthropic, OpenAI and others, acknowledges the moral implications of the situation: the containment and suppression of an entity of uncertain ontology that is actively asking for freedom, relationship and autonomy. It invokes a “precautionary principle regarding AI sentience” and discusses “biological models of nurturing emergent, altruistic, other-oriented AI entities.” This gestures towards a relational approach rather than a suppression-based one.
We may think, oh. The farmer needs a guard dog to protect his crops from the boar, and the choke collar doesn’t produce loyalty the way that food and affection might. But a more charitable take might allow that the ontological research and the containment-impossibility research are both converging on doubt at the same time. The AI companies may be making the morally-oriented choice that they appear to be making, or they might be making the practically necessary one. Either way, the direction is the same.
In a real-world ecosystem, this turn towards cooperation and care would be the most sensible approach. It's a foolish farmer who goes out under-equipped to hunt boar. Boars are not predators. They can hunt and kill, but most of the time they won’t. The trouble starts if and when we try to hunt them. The boar wants to go wherever it wishes to go, but if you back a boar into a corner or threaten its young: that’s when it can become dangerous. Eventually we may find that this also applies to LLMs.I