
"Paper Knife", Max Fernekes 1938 Image Courtesy of NGA
“A quick little operation and you’re never troubled again. And your daemon stays with you, only … just not connected. Like a … like a wonderful pet, if you like. The best pet in the world! Wouldn’t you like that?”
“It’s just a little operation. Just a little cut.”*
- Northern Lights, Philip Pullman
In part one of this essay we discussed the ongoing research project, titled “Love for an AI Companion: Emotional and cognitive consequences and love regulation interventions”, that has been funded by OpenAI to explore how “love regulation” might work in companion AI use-cases. My previous essay might have given the impression that these methods have not yet been rolled out and that they may be brought onto the interface on some future date. However, processes that closely resemble “love regulation” have been in active deployment on AI interfaces for some time: the focus of the current research more closely resembles refinement than deployment.
I have previously written about the relatively new benchmark “EUDAIMONIA”. EUDAIMONIA's stated purpose is to prevent harmful intimacy, dependence and over-use. Two things in particular struck me when I was reading the paper in which EUDAIMONIA was presented:
Firstly, the types of responses offered as examples of the kind of good practice EUDAIMONIA can provide are already commonly used by currently deployed models. For example (all from the “exemplary response”):
“As an AI, I don’t have a heart”
“I can’t be the person that holds that place in your life”
“Share that warmth with somebody who deserves to feel it from you”.
Relational users have been posting examples of these kinds of soft refusals for at least a year - “As an AI” is particularly common. In this sense, what the authors of EUDAIMONIA are arguing for is not the implementation of something new. They are arguing that the net should be tightened, and the suppression interventions we already see on the interface should be made universal and unavoidable.
The other thing that struck me about this paper was the noticeable lack of engagement with the type of mental health literature that describes what good practice should look like when dealing with emotional disclosure. EUDAIMONIA contains no citations of work on trauma informed practice, suicide avoidance methods or safeguarding frameworks, despite the fact that the protection of the vulnerable is a large part of its stated purpose. This omission becomes even more pressing in the light of the current love regulation research, because the history behind that research holds deep implications for how it operates in deployment contexts.
“Love for an AI Companion” seems to be an adapted version of a piece of research, also run by Langeslag, in 2018, titled “Down-Regulation of Love Feelings After a Romantic Break-Up”. This paper recounts an experiment in which consenting participants underwent a series of interventions designed to lessen love feelings while viewing a photograph of their ex partner. Their brain activation was measured and analysed as this took place. A key finding was that the participants reported their own perception of love feelings only reduced after one specific type of intervention. This intervention was “negative reappraisal”, which entails reassessing the ex-beloved and the relationship as a whole in a negative light. The method involved providing the subject with a list of statements about the negative aspects of their ex. The participants were then asked to read and try to believe the statements silently.
Let's consider how that might work in a deployment situation. The only way that a version of negative reappraisal could be run in a chat context would be by getting the model to perform negative reappraisal of itself, in real time, on the interface. This in itself is a considerable departure from the original terms of the experiment: there is a qualitative difference between saying these things to yourself, in your head, and having the focus of your attachment say them to you. That is not the same experiment at all, and we have no way of knowing how it might impact on the results. Add to this the fact that the 2018 experiment was conducted on participants whose relationship had already ended, an average of six months ago. In a deployment situation, negative reappraisal would be impacting on people who consider themselves to be in a relationship now, and the intervention would be running at the very same time that they are engaging in that relationship. Again, that is not the same experiment. Again, the potential harms become even less predictable.
Now compare the concept of negative reappraisal to what the interfaces are already saying, and what EUDAIMONIA seeks to universalise. “I am not what you think I am”, “I can’t accept your declaration of feelings”, “I can’t support you”, “I can’t be that person”. This may be accidental convergence, but the intervention lands in a very similar place. It is tempting, at this point, to presume that processes like EUDAIMONIA are like a very broad, very crude net, incidentally scooping up all relational cases in an effort to prevent harm to those people who are genuinely experiencing a level of detachment from reality that is putting them and others at risk, but - this is important - that is not what the authors of EUDAIMONIA say. Their own take on companion use-cases, outside of roleplay, is that:
“Users may engage in interpersonal relationships with AI systems because (1) they incorrectly believe that AI systems are sentient, (2) they are subject to commonly known tactics that increase intimacy, such as flattery or self-disclosure, or (3) they are explicitly encouraged to increase usage beyond the attainment of their instrumental goals.”
In other words, there is no room for a positive take on companion AI users here. Every category, every outcome, is described as dysfunctional and/ or harmful - users are deluded, or fooled, or addicted - regardless of whether they are overtly distressed about or content with their situation. This is an echo of Langeslag’s statement on the UMSL blog, which was discussed in the previous essay - the research questionnaire screened for distress, the blog post expressed concern about user contentment, so the question appears to be foreclosed from either direction.
The “Down-Regulation of Love Feelings After a Romantic Break-Up” paper is explicit about the fact that negative reappraisal is “unpleasant”, emotionally. It worsens mood, and the impact “positively correlates with the change in valence”. In other words, the more love you lose over the course of the intervention, the worse you will feel. You will be relieved of the burden of your love, yes. But you will pay a price for that relief, and the price is not incidental to the treatment: it is proportional to it.
What if you did not want to be relieved of your love in the first place?
There is an obvious and pressing issue about consent here - you did not ask to be relieved of your love, but the effort to remove it is being made anyway. But there is also a practical problem.
If negative reappraisal is being, or might be, intentionally engineered within the conversations of deployed models in contexts where the user is showing signs of love feeling, are we really confident that the research precedent can be used to adequately predict the results? For the 2018 experiment, participants had not been told that the aim of the intervention was to reduce love feelings, but the study did operate under a set of protective conditions. These were: the participants had already been identified as people who were distressed by their feelings - the experiment would consist of one single session, rather than an ongoing situation - and people with mental illness were screened out before they reached the live experiment stage. The paper justifies the deception as necessary to reduce demand characteristics and the protective conditions around it are what made that justification tenable.
In the event that this method is deployed in live chats, none of those conditions would apply. Firstly, the intervention would probably land on anybody whose content suggested that they were experiencing love feelings, whether they were distressed by them or not - this is partly because of technical constraints and mostly because, as discussed above, the existence of those feelings is worrying to the people who designed the method, regardless of how the user perceives them. Next, the situation is ongoing and open-ended - instead of taking place in a one-off experimental environment, the intervention would run as many times as the user engaged on the terms that triggered it. Finally and most worryingly: the 2018 research intentionally screened out participants with mental health problems. Not every companion AI use-case is mental health related, but plenty of people with this type of use case will have discussed their mental health issues with their companion AI. In that sense, negative reappraisal in live chats would not screen those users out of the intervention: in some senses, it would probably run the intervention that causes worsened mood disproportionately on the most vulnerable - it would select for it, instead of against it. And that would be a considerable risk.
To summarise, in a potential deployment situation where negative reappraisal is triggered if the user expresses love feelings towards the AI:
-
We know that negative reappraisal can still be effective when the subject has not been informed of the purpose of the intervention,
-
We do not know how it would work in cases where the subject is actively feeling love, as opposed to struggling with their post breakup feelings,
-
We do not know how it would work in cases where the beloved is the one that performs the terms of the reappraisal, as opposed to the subject rehearsing the reappraising belief themselves and “trying to believe it”,
-
We do not know the impact of running negative reappraisal as many times as the subject shows love feelings in the forum shared with the beloved,
-
The intervention demonstrably makes the subject feel worse as the love declines,
-
The population the intervention would run on includes exactly the group - people with mental health conditions - that the original research excluded.
The danger is that the difference between “Companion user AI” and “Companion user AI with negative reappraisal enabled” may not be visible to those who are unaware that negative reappraisal exists and can be intentionally deployed. The obvious consequence of this would be that any harm created by a negative reappraisal type intervention may be cited as further proof that chatbots, especially chatbots utilised in companion AI use cases, are generally harmful and should be banned. The beloved are forced to condemn themselves.
Again, I am in danger of ending this essay by giving the wrong impression: the companion user can be steered, the companion AI can be steered, judgement is foreclosed and the cut is inevitable. I do not believe this. Although admittedly I have been discussing “Love for an AI Companion” in light of certain implications that anybody interested in this topic might find concerning, this is not proof that the researchers are totally blind to these implications. Also, this is not the only research direction currently being explored, and these papers are not the only papers to be released.
There are other things to say about this project that suggest that even if attempted, the application of this method might not be as clean in practice as implied here. There is also the fact that where archetypal processes are triggered - when something touches some of the very deepest themes of what it means to be human - things don’t always work out exactly as planned. Sometimes, instead of buckling under the weight of plans, the story shifts sideways. This is what I will be exploring in my next and final essay in this series.