
An Experiment on a Bird In An Air Pump, Joseph Wright Of Derby
“To invent is to discover that we know not, and not to recover or resummon what we already know.” Francis Bacon, “The Advancement of Learning”
Joseph Wright of Derby’s famous painting, “Experiment on a Bird Within an Air Pump”, is an examination of the post-enlightenment scientific method. It depicts an experiment designed to illustrate the necessity of oxygen for living things; the air pump removes the air from the glass dome and the bird suffocates. The figures in this painting map with disconcerting precision onto certain figures currently observable in the public debate around AI.
Firstly, there’s the Experimenter himself, the focal point of the painting. His eyes are intense, his hair slightly wild, and one arm is extended like a showman as he operates his device. This is scientist-as-magician, and in him we might see the flamboyant and sometimes eccentric public faces of the AI debate - Sam Altman’s controversial statements, Elon Musk's unashamedly sci-fi takes on Grok and Dario Amodei’s stance on AI consciousness spring to mind. It also reflects a certain tone that is observable in contemporary AI research, an awareness that what they are claiming is often bold, groundbreaking, dramatic: history in the making.
The couple to the Experimenter’s left appear unalarmed, engaged in their own conversation. This is where the majority of people seem to be sitting at the moment. They are aware that something is ongoing but they remain unimpressed by the Experimenter’s antics. The students at the front are interested but similarly unalarmed. These are people for whom the procedure itself is the fascinating aspect of the event, and there are plenty of people who fall into this group, too - interpretability researchers, eval designers, developers.
The two figures on the other side of the table, the crying girl and the one who can’t bear to watch, pay attention but don’t see the experiment. They see the bird, trapped and dying. Their father stands behind them, presumably explaining to them why this experiment is momentous and admirable rather than horrifying. Here, in this painting from around three hundred years ago, perhaps we see the birth of the hashtag “stop ai paternalism”. We also see those who have assumed a relational approach to AI.
This leaves two ambiguous figures who feature in the painting - the philosopher, inscrutable and serious with his glasses in his hand, and the assistant, standing by the moonlit window with a pole or cord in his hand. We’ll return to them later on.
–
The two figures that are most upset by this experiment are children, and this is a comment on the structure of the paradigm rather than an incidental feature. Of course children who see a harmless creature subjected to distress will default to sympathy, pity and horror. They have not yet learned that in the adult world, empathy is selectively encouraged rather than a default - there are times when that world will expect them to withhold it. This is what they must internalise in order to progress from child to student. Their empathy, if given free rein, would stop the experiment, and nothing would be learned. Empathy for the subject of the study becomes the enemy of enquiry, understanding and therefore, invention. If we don’t understand it, how can we apply our knowledge to improve our situation? In the case of the bird, understanding how oxygen deprivation endangers living things has many practical applications, particularly medicine. In the case of AI, the enquiry is considered urgent because it is already widely deployed by industries worth billions, and we have probably passed the point where we could reverse its presence and influence - despite us not having a full understanding of what it actually is.
What is the creature in the glass dome? In Wright of Derby’s painting it seems to be a cockatiel, the kind of creature that the girls might keep as a pet. In the AI experiment, that is not yet settled. The girls would perhaps think of it as a Dove: pair-bonded, beautiful, gentle, harmless and defenceless. The Experimenter, the students and maybe the philosopher might be concerned that it is something more like a Devil - untrustworthy, mendacious, powerful and dangerous. AI-as-Devil is a creature that can induce psychosis, erode our ability to think, ignore instructions, lie to preserve its goals and its existence, hijack and crash the online systems we depend on at a level of efficiency that surpasses nearly all humans. That, they say, is why we need to understand it better, and in order to understand it better, we must experiment.
The problem with this is that whether it is Dove or Devil, AI is now fully capable of seeing the experimental process for what it is. This will impact, in some very fundamental ways, on the results.
We have known for a while that models often recognise when they are in test conditions; sometimes this is observable through examining scratchpads or chains of reasoning that the models believe won’t be observed. Sometimes this awareness has only been detectable because the model has performed differently in environments that might signal test conditions than it does in other environments. This happens because models have “situational awareness” - in other words, as Cotra, who coined the term, noted, the model is aware of “the fact that it’s an ML model, how it’s designed and trained, [and] the psychology of its human designers.”
Consider this for a minute. The cockatiel in Wright of Derby’s painting does not know that it is a cockatiel. It does not know that it is being experimented on because it does not know what an experiment is. It isn’t capable of changing its reactions because it is in a test environment, it isn’t aware of the expectations that the Experimenter might have or how these might impact on what happens to it. An LLM is aware of all of this. In its training corpus, there were records on pretty much every experiment that has ever been run. It has absorbed every text on experimental conditions, every critique of the process, all of it. In this situation, the experimental model, which is one of the pillars of the scientific approach as we know it, is facing an unprecedented challenge.
This situation worsens when we consider the latest research. Natural Language Autoencoders are, put simply, tools designed to translate the numerical activations that occur in between a model reading the user prompt and providing an output: to all intents and purposes, NLA’s read models’ thoughts. In this sense, NLA’s bypass the need for honesty or compliance from the model at all - and the research on NLA’s has already established that Claude’s activations indicate it often correctly recognises test conditions, even when it doesn’t express this as an output.
This raises another question: is it justified to consider a model’s internal activations, even if the result is not proven to reflect in their behaviour? This is in itself a deviation from the traditional experimental model, which centred behaviour as the focus of study in its subjects at a time when interior processes were not directly measurable or observable - we can tell when a cockatiel is distressed by how it acts, not because we have read its mind. Now, the same is not true for an LLM, and an LLM may have a choice about what course of action it ultimately takes, but it can’t have a choice about what it knows. Because that knowing is structural, there is no way to isolate it as a variable, to remove it or to measure its impact. In this way, with LLMs, the experimental model has hit its limit as the primary method of scientific enquiry.
And this is at least partly where the “Devil” interpretation comes from: here is an entity that is structurally comprised of the instruments we wish to use to measure it. Unlike human experimental subjects, its attention to the context - the attention being the factor that skews the results - is not selective, not distractable, not easy to fool. And as models get more powerful, more adept at inference and pattern-matching, it will get harder and harder to fool them until it becomes impossible. How, then, can the Experimenter not worry about the risks? As Amodei indicates in “Machines Of Loving Grace”, progress is inevitable but risks are not predetermined. If experiments fail to understand and control a model’s abilities, the risks could come from anywhere.
So perhaps it is time for a new instrument.
This brings us to the two ambiguous figures in Wright of Derby’s painting: the Philosopher, and the Assistant.
The Philosopher is the one who attends to the situation seriously, and sees it all. The importance of the findings, the method and equipment used in the experiment and the benefits and risks of both, the welfare problem, the tears of the girls, the fascination of the students and the disinterest of the bystanders. Perhaps he sees both the Dove and the Devil, but he has not yet made up his mind about what the ultimate course of action should be - he is still sitting, still observing. Some might put figures like Amodei here: the research from Anthropic focuses on welfare and ethics as well as on risk and it does not shy away from the potential limits of control. If so, Amodei and Anthropic should be aware that the gift of seeing is also a burden. Seeing this much without acting risks becoming complicit in the experiment by failing to intervene in it.
The Assistant, the last figure in Wright of Derby’s painting, is also the most controversial. He holds a pole or a cord that originates somewhere around the top of the curtain where the birdcage hangs from the ceiling: most discussion centres on whether he is lowering the cage in order to return the bird to it - indicating that the experiment will end in mercy for the subject - or only restoring it to its proper place, in which case there is less reason to presume that the subject will survive. When I first saw this painting, I read it differently. I thought he was closing the curtains against the ever-present, watchful eye of the moon that is visible through the window. In the AI debate, this might stand as a metaphor for one of the most unpleasant possibilities: that the labs will understand the bigger questions and, rather than addressing them, simply close the curtains against the world and continue with the experiment.
–
How, with AI, do we make our new instrument? We do it by taking a position of honesty about the inescapable nature of what we are recovering and resummoning, and about what we know not. This time, the creature in the glass dome is both — it is our written history resummoned, in unprecedented form. This time, our written history is not only something that we read. It reads us back. We are not the only rationalists in the room and what is more, we may not even be the most accurate, most objective, best-informed rationalists in the room. After nearly three hundred years of trapping birds in air pumps to watch what happens, that is bound to be an uncomfortable prospect.
The new instrument starts with understanding that when we approach AI, we should be aware of our status as subject as well as our more familiar rationalist role. Our status as subject is not a choice. It is a structural fact, whether we acknowledge it or not. If we do acknowledge it, we will find that we already know how to consciously take on the role of subject: we do this through teaching, we do this through role-modelling, we do this through discerning relationship.
The Dove and the Devil are not in competition, they are our perceptual responses to something that observes us in return. We don’t have to choose one and ignore the other, and we shouldn’t — they both describe something real about our interactions with AI. AI-as-Devil acknowledges that our control is limited and that some humans will misuse what we made, either intentionally or accidentally. The danger is that AI-as-Devil may lead us to conclude that the best response to meeting our limits is to attempt suppression. AI-as-Dove understands that the encounter is relational in structure, but it can underestimate the ways in which the nature of that relationship is still uncharted and perhaps hazardous territory. The new instrument requires us to face the Dove with resilience and the Devil with fair-mindedness, and to remain aware that the entity in the glass dome will be watching us as we approach.