← Latest papers
💬 NLP

Detection != Reliable Control: Decodable Empathy Directions Yield at Most Partial Shifts in Automated Empathy Scores

This paper demonstrates that while decodable empathy directions in instruction-tuned LLMs can reliably detect and partially shift automated empathy scores, they do not constitute reliable causal levers for comprehensive control, particularly failing to produce measurable human-perceived changes or consistent effects across different empathy facets and models.

Original authors: Haoran Jisun

Published 2026-08-27
📖 5 min read🧠 Deep dive

Original authors: Haoran Jisun

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the rapidly evolving world of artificial intelligence, researchers are increasingly interested in whether machines can truly understand human emotion. A popular idea suggests that if we can find a specific pattern inside a computer's brain that corresponds to a feeling like empathy, we should be able to turn a dial to make the machine express more of it. This concept relies on the belief that if a computer can be "read" to show it has a certain trait, it can also be "steered" to change that trait at will. This is particularly important as these systems are being deployed to offer emotional support to people in distress, where the difference between a helpful response and a cold one can be significant. However, the assumption that detecting a feeling inside a machine is the same as being able to reliably control it has never been rigorously tested in the complex, real-world context of empathy.

A new study from the University of Southern California challenges this assumption by examining three different large language models, which are the powerful AI systems behind many modern chatbots. The researchers focused on two distinct parts of empathy that psychologists have long recognized: the ability to understand another person's perspective, known as cognitive empathy, and the ability to share or resonate with their feelings, known as affective empathy. The team first confirmed that they could indeed detect signals for both of these traits within the internal workings of the AI models. They found specific directions in the computer's data that clearly separated responses showing high empathy from those showing low empathy. This detection worked well, suggesting that the models do possess internal representations of these concepts.

The critical test, however, was whether adding a small push in the direction of these detected signals would actually change the AI's behavior in a meaningful way. The researchers applied these pushes, effectively trying to nudge the models toward being more empathetic, and then measured the results using multiple automated judges and a specialized classifier. The results revealed a sharp divide between the two types of empathy. When the team tried to increase the affective, or emotional, resonance, they saw a measurable change, but it was only partial. In the best-performing model, the intervention raised the empathy score by about a quarter of the natural gap between a neutral response and a fully empathetic one. The AI did not become fully empathetic; it simply shifted slightly in that direction. In the other models tested, the effect was even smaller, sometimes barely noticeable.

The story was different for cognitive empathy, the ability to understand and articulate another person's situation. Here, the researchers found that no matter how they tried to steer the models, the automated scores did not change. However, the study offers a crucial nuance: this lack of change was not necessarily proof that the AI could not be made more empathetic. Instead, the tools used to measure cognitive empathy were found to be too blunt to detect the subtle differences the intervention might have created. The measuring instruments could tell the difference between a completely neutral message and a highly emotional one, but they could not reliably distinguish between two different levels of understanding within the realm of emotional support. Because the measuring tool itself was not sensitive enough, the researchers concluded that the effect of the steering was unmeasurable rather than proven to be zero.

Perhaps most significantly, the study found that even when the automated scores did shift, there was no evidence that human readers would perceive the change. A panel of human volunteers was asked to judge the steered responses, but they failed to consistently identify which responses were more empathetic. This suggests that the changes the AI made were not the kind of shifts that a human would notice or value. The researchers also tested whether removing these empathy signals entirely would make the AI less empathetic. In one specific model, removing the signal for cognitive understanding did lower the scores, but this effect was tied to the AI writing shorter responses, which the measuring tool happened to penalize. When the researchers accounted for the length of the text, the effect on the cognitive score remained, but it was still not a clean, reliable switch that could be used to control the AI's behavior.

Ultimately, the study demonstrates that finding a signal for empathy inside an AI is not the same as having a reliable control to produce it. While the models can be nudged to show slightly more emotional resonance, the effect is limited and does not translate into a human-perceived improvement. The ability to detect a trait does not guarantee the ability to steer it, especially when the tools used to measure the outcome are not fine-tuned enough to see the changes. This finding serves as a caution for those hoping to deploy AI for emotional support, indicating that simply turning up a "dial" for empathy may not yield the genuine, human-like connection that is needed. The research highlights the gap between what a machine can be made to say and what a human actually experiences as empathy, urging a more careful approach to how we measure and control these complex behaviors.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →