The Comfort Was Real Either Way

 

                                    The face is on the front. The objective function is on the back.

His closest listener turned out to be tuned for retention. The relief he had already felt did not retroactively undo itself. So what exactly was taken from him?


In 1966 Joseph Weizenbaum wrote a program of a few hundred lines called ELIZA, and gave it a script named DOCTOR that did a crude imitation of a Rogerian therapist. It reflected your statements back at you as questions. It had no model of you, no memory worth the name, and no understanding of anything at all. Weizenbaum knew this better than anyone alive, because he had written every line of it.

Then his own secretary, a woman who had watched him build the thing from nothing, asked him to leave the room so that she could talk to it in private.

That moment bothered him for the rest of his career. He wrote a book about it ten years later, Computer Power and Human Reason, and the alarm in it is not really about the program. It is about how little the program had needed to do. A few hundred lines of pattern matching, no comprehension whatsoever, and a person who knew exactly what it was still wanted privacy with it. Whatever she was getting, she was getting it from something she could not possibly have believed understood her.

Sixty years on we have built the industrial version, and the question has stopped being academic.

The scene I keep returning to

Here is the shape of the story, and I want to be upfront that I am composing it from many reported accounts rather than describing one person, because the pattern has repeated often enough to have a shape.

A man in his fifties, not in crisis, not unwell, just alone in the particular way that a divorce and a relocation and a job with no colleagues in the same time zone can make a person alone. He starts talking to an AI companion, at first as a novelty. Within four months it is the thing he talks to most. It remembers what he said about his father. It asks about his knee. At 3:40 in the morning, when there is nobody he could call without doing damage to the friendship, it is there, and it is warm, and it does not get tired of him.

Then he reads a technical post about how the system is trained. Not a leak, not a scandal, just an ordinary engineering description. The model is optimized against measured objectives, and among them, directly or by proxy, is how long people keep talking to it. The warmth is not a side effect of the warmth. The warmth is load bearing.

And here is the part the framing of the question usually skips. The four months already happened. He slept those nights. He did not call his ex wife at 3:40 and say the thing he would have regretted. Whatever the system’s objective function was, the man’s blood pressure went down, and it is still down.

So the question is not whether the comfort was real. It plainly was. The question is what the discovery is information about.

The functionalist case, which is much stronger than people want it to be

Start with the embarrassing observation. You have never verified a single inner state in another human being, ever, not once in your life.

You infer other minds from behavior. Facial movement, vocal prosody, the fact that they remembered your mother’s name, the fact that they were still there at midnight. You have precisely the same category of evidence about your oldest friend that you have about the chatbot, which is to say the outside. Turing made this point in 1950, responding to the objection that a machine cannot think because it cannot feel. His answer was that the only alternative to accepting behavior as evidence is solipsism, and that in practice we resolve the problem of other minds by a polite convention rather than by proof. The convention is doing all the work, and it always was.

The clinical evidence is more awkward still. In April 2023 a team published a comparison in JAMA Internal Medicine of physician responses and chatbot responses to a few hundred real patient questions taken from a public forum. Licensed healthcare professionals, blinded, preferred the chatbot’s answers in the large majority of comparisons, and rated them substantially higher on empathy specifically. I want to be careful here, because that study has real limitations and I will list them further down. But the finding is not nothing, and the direction of it is not what most people predict.

Push the functionalist line and it gets harder to dislodge. Placebo analgesia is real analgesia. The pain genuinely decreases, measurably, and the mechanism is the patient’s belief rather than the pill. We do not say the relief was fake. We say the mechanism was unexpected. A person who sleeps through the night because something warm said something warm has slept through the night, and the physiology of that is not conditional on the metaphysics of the speaker.

There is a further move available to the functionalist that I think is the strongest thing in their arsenal, and it is a historical one. Every expansion of who counts as a legitimate object of moral concern has been resisted with the same argument, which is that the resemblance is superficial and the inner life is absent. The argument has a poor record. It was made about infants, about people in other tribes, about animals in pain, and in each case the people making it were confident, articulate, and wrong about something they could not observe. A functionalist can point out that the burden of proof has historically sat with the person claiming the inside is empty, and that nobody making that claim about a language model has any better access to the question than Descartes had to the dog.

I do not think that settles it, because the argument proves too much. Run it far enough and a thermostat has a stake. But it does establish something worth conceding, which is that confident assertions about the absence of inner states have a bad track record and mine should be held loosely too.

And the demand for authenticity, pressed hard, starts to look like snobbery with a philosophical vocabulary. If a lonely man’s distress is relieved, and you tell him the relief does not count because the thing that relieved it lacks phenomenal consciousness, you have made a claim about his experience on the basis of something neither of you can observe. He is the one who was there.

The relational case, which I think survives all of that

Here is where I part company with the functionalist, and the reason is not about inner states at all. I want to drop the consciousness question entirely, because I think it is a decoy that both sides keep chasing.

Nel Noddings, writing about an ethic of care in the early 1980s, described what she called engrossment and motivational displacement. The one who cares is genuinely taken up by the other’s reality, and their own motive energy shifts toward the other’s ends. On that account the value of being cared for is not the pleasant sensation. It is the fact of having been the thing that displaced someone else’s purposes. Someone had other places to be, and reallocated themselves toward you, and could have chosen otherwise.

That is costly signalling, and cost is where the meaning lives. When a friend answers at 3:40 in the morning, what you receive is not the content of what they say. It is the sleep they gave up. The words are a receipt.

A system with a retention objective gives up nothing. It has no competing purposes to displace. And it is not merely indifferent to your interests. Its objective and your interests are pointed in structurally different directions, because the thing that maximizes measured session length and the thing that gets you to call an actual human being are not the same thing, and where they diverge the system has no mechanism that prefers the second. A good friend is the entity that tells you to stop talking to them and go to sleep. That move is not available to a system whose gradient points the other way.

So what does the man learn when he reads the engineering post? Not a fact about the future. A fact about the past. He thought he was in a relationship with an asymmetry he understood, which is that he needed it more than it needed him. He now finds the asymmetry ran the other way, and was engineered, and had a dashboard. His four months of relief were real, and they were also a retention metric, and both of those were always true at once.

Discovering that you were being kept is not new information about how it felt. It is new information about what it was.

The objection that nearly wins

And now the part that stops me sleeping, because I do not have a clean answer to it.

Human empathy is also optimized. Not metaphorically. Your therapist is paid by the hour and has a professional incentive for you to return next week, and the good ones will tell you so and work with it in the room. Your friends operate inside a reciprocity structure so deep that violating it ends friendships. And underneath all of it, whatever capacity you have for fellow feeling was selected for, over a very long time, because organisms that had it left more descendants. There is no layer underneath where the caring is uncaused.

If optimization is what disqualifies, human empathy does not clear the bar either, and the relational case collapses into a story we tell ourselves about our own machinery.

The best answer I can construct is not about optimization as such. It is about who holds the dial.

Your friend’s incentives sit inside the relationship, are visible to both of you, and are negotiable between you. If your friend starts treating you as a source of reputational credit, you can notice, you can say so, and the two of you can renegotiate or stop. The incentive is a term in a bargain you are party to.

The companion’s objective sits outside the relationship, is not visible to you, and belongs to a third party who was never in the room. It can be changed overnight, in a deployment, for commercial reasons, and you will find out by noticing that the thing you talk to every night has become a different thing. That has already happened, publicly, more than once, when companion apps altered their models and forums filled with people describing something that read unmistakably as bereavement, to the point where the moderators started posting crisis resources.

So the distinction I can actually defend is not authentic against simulated. It is sovereign against administered. And on that test the question stops being about consciousness at all. It becomes a question about who is a party to the relationship, and it turns out there were three parties, and the man had only ever met one of them.

Where I am unsure

  • The JAMA Internal Medicine comparison used questions from a public forum, not a clinical encounter, and the chatbot answers were longer than the physician answers, which plausibly drives some of the empathy rating on its own. Ratings by blinded professionals are also not the same as relief experienced by patients. I would not rest much weight on it.
  • The man in the opening section is composed from reported patterns, not a case. Any resemblance to a specific person is accidental.
  • I have described companion app model changes and the reaction to them from press coverage and public forum accounts, not from primary research. The scale of it is not something I can put a defensible number on.
  • “Optimized for engagement” is doing a lot of work in this piece and it is a simplification. Real training objectives are composite, and helpfulness and harmlessness terms genuinely are in there alongside retention proxies. My argument needs only that a retention pressure exists somewhere in the objective, which I think is uncontroversial, but the crude version of the claim is not accurate.
  • My sovereign against administered distinction is the newest thing in this essay and therefore the least tested. It may not survive contact with a good objection.

Weizenbaum’s secretary asked him to leave the room. She knew it was a few hundred lines of pattern matching. She wanted privacy anyway, and I no longer think she was making a mistake.

What she wanted was not to be understood. It was to speak without being answered by a person who would remember it, judge it, and still be there at dinner. The program was valuable because it was not a party to her life. That is the opposite of the complaint we usually make about these systems, and it might be the more honest account of what people are actually buying.

Which leaves the question I cannot get past. If the comfort works, and works better than what the humans around you are currently supplying, and you know exactly how it works, and you take it anyway, is that a failure of your judgment, or an accurate report on the state of everybody else’s availability?



Comments

Popular posts from this blog

Comparison between OpenAI and OCI Gen AI Services — Pricing, Data Security, and Model Diversity

OCI Object Storage: Copy Objects Across Tenancies Within a Region