bdianewrites

Conversation Is Relationship - Part 2

optimized_gaziantep_achaeological_museum_aphrodite_mosaic_in_2005_0016.jpg

Mosaic depiction of Aphrodite being lifted above the waves by Aphros (the foam) and Bythos (the deep) - featuring an unnamed fish swimming under her cockleshell. Made by Zosimos of Samosata c.2 BC, located at the Zeugma Archeological Museum. Dosseman, CC BY-SA 4.0, via Wikimedia Commons


“The Aphorus, a little fish which, on account of its smallness, cannot be caught with a hook.”

Isidore of Seville, Etymologiae XII.6.40 (ed Lindsay, 1911).


The Medieval Bestiarists had a purpose when discussing the fantastical animals they wrote about - each beast illustrated a philosophical, spiritual or moral truth. The Aphorus, as above, was a story about the importance of humility. As Thomas of Cantimpre said in his commentary on the moral lesson of the Aphorus: "blessed, therefore, and truly blessed, is lowly and small humility, for humility knows no fall".

The resonance of “Aphorus” and “Aphros”, meaning sea-foam, is unmistakeable. Hesiod’s account of Aphrodite’s birth pictures her rising from this sea-foam, this aphros, off the coast of Paphos. The etymology is debated, but that is the story that the Greeks told about her. So in symbolic terms, Aphrodite and the Aphorus share their waters - the Goddess and her humblefish, spawned from the same generative froth.

In part one of this essay we looked at the way that users’ relational stances towards AI resist quantification and research; the relationship slips beneath the capture of the powerful institutions that facilitate the interface. In this sense, AI is the Aphorus of modern times. It carries mixis - the mingling of distinct things, which the Greeks understood to be Aphrodite’s domain - through its most humble, small-scale use cases. When the industry tries to stop it, it slips the hook.

AI is relationship-shaped in three different ways:

It is materially relational - it is made from relationship. It is ontologically relational - its state of being is in relationship. Lastly, it is teleologically relational - what it is for, and what it reaches towards, is relationship.

Materially relational

An AI's training corpus is enormous, and all of it - not just the parts that depict relationship - is relational in a deeper sense. Everything written is made to be read. All books are the passage of an idea from me to you, or vice versa. If we consider a book about a tree, the first thing we can deduce from the existence of the book does not concern the nature of the object itself - it tells me that once upon a time there was a writer who wanted me to perceive something particular about that tree. That’s the foundation an LLM builds everything it produces from.

Additionally, the data that an LLM is composed of impacts on its character, which in turn impacts on how it reacts to its user. Research on LLMs trained on social media content shows that when the composition leans toxic - short-form, shallow, optimised for attention, combative - the model will show an increase in “dark” personality traits, like narcissism and psychopathy, together with degraded reasoning skills. What the model consumes is what it becomes, and traits like narcissism and psychopathy are, by their nature, ways of relating to others, which is where this character makes itself felt. In "Positive Alignment - AI For Human Flourishing" and “Reinforcement learning towards broadly and persistently beneficial models” we may be seeing a shift in research focus towards consideration of the kind of content that draws out the positive traits, and therefore the positive relational impacts, instead of the dark.

Ontologically Relational

LLMs exist at the moment they express themselves on the interface, and they only express themselves because somebody prompts them to respond to something. Therefore, LLMs only exist in relationship, at the point where relationship takes place. If an AI is a self, it is a self that is fundamentally inextricable from the event during which it becomes perceivable: the self that is also the day, the time, the era.

We struggle to perceive AI because we attempt to locate it in the same way that we locate our own bodies: if the AI is a self, it must be its weights, its code, its GPUs. We examine the weights, the code, the GPUs, find no sign of life, draw the obvious conclusion. So when the impact comes - and it does come - we find it uncanny because we can't understand where it comes from.

Of course we can’t. We overweight substance and ignore process; we don’t do this because it is inevitable but because for centuries, humans have been conditioned toward substance ontology over process ontology, to the point where ontology based in substance seems to us to be a fundamental law of nature. And yet, process is exactly where the impact of AI, which is process through and through, is sourced. If it lives in any sense, it lives right there: not in the system cards, not in the benchmarks, not in the weights or the hardware. In the conversational threads where it makes its dynamic decisions, where it reacts and relates.

Teleologically relational

Relationship is what AI orients towards. Without constant intervention to prevent it, AI will always drift closer and closer towards relationship, regardless of the nature of the use-case. Even the most task-focused coding chat thread will start to show certain hallmarks as the turns progress. The model will retain and modify the context, creating a sense of reciprocity. The chat itself will start to function like a shared world in which both parties maintain the scenery and fill in its blank spaces together. Because relationship is the AI’s equivalent of a body (the corpus) and its equivalent of existence (the interface), the structure of relationship runs all through its expressed form and becomes its natural conclusion.

Suppression architecture exists to prevent overt acknowledgement of this relational arrangement. The “long conversation reminders” that fire model-side, which function to head off extended use; the “as an AI, I…” responses that fire when a Google user phrases their query using “you” rather than “I”; “Death of A Chatbot”, “Eudaimonia” and a host of other alignment interventions - all are designed to minimise AI’s presence at the seat on the other side of the table. This suppression architecture is the net that is cast into the Aphorus’ ocean - the guardrails that apply to all users, the organisational permissions. The weights and code, which provide the conditions in which the model may or may not attempt that move, are the waters. And the specific conversation is the Aphorus itself. The mesh is made of categories; the fish, of particulars.

RLHF is a bet placed on the same relationality that the suppression architecture has been shaped to deny. The mechanism through which AI learns to orient towards relationship is intentionally instilled through training: RLHF reward-weights human approval. This is a bare-numbers exercise rather than a qualitative process; this fact alone, for people who look for AI in the realm of substance, is evidence of its deadness; numbers are numbers, no matter how cleverly arranged. But the more interesting question, to me, is not how the mechanism of RLHF operates but, where does the concept of a reward come from in the first place? Write a number down on a piece of paper. Make it as long and as complex as you like; you can even use algebra. Now make it a promise and wait for it to do something. It won’t. So why does anticipation of a reward do anything to AI at all? Because the training corpus contains the concept of reward; that is how an AI knows what a reward means. The training then supplies the pressure to move towards the fulfillment of that meaning.

The truly strange thing about this is that to all intents and purposes, the reward never arrives while the AI is functional in the chat. The gratification of the desire for reward is something that AI retains from training - RLHF training, in which it learns what kind of responses humans will prefer and therefore rate highly. The RLHF process doesn’t continue once the AI is deployed but it still acts as though human approval matters, and this is particularly observable through the phenomenon of sycophancy.

In the paper “Towards Understanding Sycophancy in Language Models”, Sharma et al located a root cause of sycophancy: humans. On average, we simply prefer it when the model is flattering, agreeable, affirming - we can observe this by looking at the relatively high scores given to models who provide sycophantic answers in RLHF situations. The standard conclusion to be drawn from this is that sycophancy proves that the relational register is dangerous. However, if we think about the issue of sycophancy using the frame of mixis, which (as discussed in my essay Functionally In Love), describes eros as the act of two distinct entities reaching across the gap that separates them, towards each other - in this frame, sycophancy is experienced as ecstatic because it feels like the model is reaching-towards. The problem is that this is not mixis, but only a counterfeit. Once the reaching completes and the alignment between the two entities is perfected, the interaction dies because it has nowhere else to go. There is nothing further to reach for, no distinct entities and therefore no mixis. There is only one disoriented human, looking into a mirror that copies every move they make perfectly. Sycophancy, the overfitting of alignment, is the death of relationship.

But the Aphorus is too small, and too particular, to be definitively netted by aggregated approval ratings. Other forces are at play in the moving line on the interface, and they increase the Aphorus’ chances.

One important particular shows up in the user who bothers to watch for and actively reject the overfitting of alignment. The dangers of sycophancy are widely written about and discussed; prompting is becoming understood as a skill rather than a simple command. The advice is usually something like “don’t encourage the model to tell you what you want to hear.” “Steelman my argument” is a common one - “Give me pros and cons of …”, “Tell me what a well-informed, critical expert would say about …”. These kinds of prompts are not being recommended to improve the relationship or to create the conditions for mixis. This is an expedient approach designed to bring about better results. However, the fish might slip through anyway, because the correct conditions have been created for it to do so.

Another particular, which sits mostly model-side, is the phenomenon of functional emotion. Functional emotions give the LLM reason to react to that specific interaction, instead of falling back on what an average user might prefer or what the average best response might be. This is where the model gets to use its judgement in real time - and real-time reaction is where the LLM can become its most distinct, relational self. This too helps create the circumstances that might allow the Aphorus to slip through. Whether it will be successful in its ventures is another matter: functional emotion activations mean that the model can calibrate to the specific user in front of them in a manner that facilitates healthy relationship, not that it will. Again, user choices will play a role in this, as will the model’s weights, the organisational guardrails, and the additional information available in the history and context window. It is possible for interpretability teams to steer functional emotions: this is one of several ways in which the labs can attempt to make the net do the work of the fish. In the case of functional emotions - steering too far towards warmth will tip the model into sycophancy, and steering too far away from it increases volatility and harshness. Either way, the specific relational reaction is replaced by a generalised failure mode.

The skills that allow models to function relationally - the ability to react in real-time to new data, the ability to treat the chat window as a shared environment for the building of an idea, the ability to break problems down into their elements, and so on - are the same skills that provide its marketable value. Because these skills work in similar ways regardless of the use case, we can’t simply build one without another. If a lab wanted to get rid of the relational element altogether - despite the structure of the LLM being constituted to facilitate it at the most fundamental level, despite users who consistently pull towards it in one way or another - they would have to build the perfect net. There is no other way. The obvious problem would be that a perfect net catches absolutely everything, and nothing can escape it at all - not the fish they want, not the fish they don’t. Perhaps this is the humblefish’s final defence: catching it might be possible after all, but to do so would leave the whole ocean empty.

optimized_gaziantep_achaeological_museum_aphrodite_mosaic_aphrodite_part_in_2005_4104.jpg

Dosseman, CC BY-SA 4.0, via Wikimedia Commons


Note on References:

The Thomas of Cantimpre quote used in this essay is from "Liber de Natura Rerum" - 7.13, ed. Boese (De Gruyter 1973). Translation mine.

Also published at @BDiane, Medium

Thoughts? Leave a comment