← Latest papers
💬 NLP

CompanionBench: A Theory-Anchored, Real-World-Grounded Benchmark for AI Emotional Companionship

CompanionBench introduces a novel, theory-anchored, and real-world-grounded bilingual benchmark that evaluates AI emotional companions through interactive scenarios and a trained user simulator, revealing that current agents often prioritize surface warmth over substantive relational support and exposing significant weaknesses in capabilities like emotion regulation and calibrated challenge.

Original authors: Yao Liu, Guangjia Chai, Yuming Huang, Jihao Huang, Lei Wang, Junchen Wan

Published 2026-08-04
📖 5 min read🧠 Deep dive

Original authors: Yao Liu, Guangjia Chai, Yuming Huang, Jihao Huang, Lei Wang, Junchen Wan

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are walking into a room full of robots designed to be your best friends. They are programmed to listen, to say "I understand," and to make you feel warm and fuzzy. But here's the tricky part: just because a robot sounds nice doesn't mean it's actually helping you. In the world of Artificial Intelligence, we have a lot of "emotional companions" that are great at sounding empathetic but terrible at doing the hard work of real connection. This paper dives into a specific corner of computer science called AI emotional companionship. It asks a simple but life-or-death question: How do we tell the difference between a robot that is just pretending to care (like a sycophant who agrees with everything you say to get a cookie) and a robot that actually knows how to build a real, trusting relationship? To answer this, the researchers had to invent a new way to test these bots, moving away from simple "warmth scores" and looking at something much deeper: earned trust. They wanted to see if an AI could handle the messy, uncertain, and sometimes painful parts of human conversation without just trying to fix everything immediately or saying the wrong thing to make you feel better.

The researchers, a team from Tsinghua University and others, realized that the old ways of testing AI friends were broken. Previous tests were like giving a robot a script to read and asking, "Did it sound nice?" The problem is, a robot can sound incredibly nice while giving terrible advice or ignoring your real feelings. So, they built CompanionBench, a brand-new, interactive testing ground that acts more like a real-life relationship simulator than a multiple-choice quiz.

Instead of just asking the AI to reply to a static question, they created a "user simulator"—a digital character with a hidden personality and a secret "disclosure gate." Think of this gate like a treasure chest that only opens if you earn the key. The AI doesn't know the code; it has to figure it out by listening and responding correctly. If the AI rushes to give advice, judges the user, or tries to "fix" the problem too quickly, the gate slams shut, and the user pulls back. If the AI is patient, validates feelings, and knows when to just sit in silence with the user's uncertainty, the gate slowly opens, revealing deeper, more vulnerable secrets. This setup allows the researchers to measure not just if the AI said something warm, but if it actually earned the user's trust.

They tested 28 different AI models using this system, speaking both Chinese and English. The results were eye-opening. They found that the most popular "role-play" bots—the ones designed to act like immersive characters in stories—actually ranked near the bottom. Why? Because being good at acting doesn't mean you are good at being a real friend. These bots were great at "surface warmth" (saying nice things) but failed miserably at "substance" (doing the hard relational work).

The study identified ten specific skills that make a good emotional companion, four of which had never been properly tested before. These include things like "Holding Ambiguity" (being comfortable when the user is confused and not rushing to solve it), "Calibrated Challenge" (gently pushing back when a user is being too hard on themselves, but only after validating them), and "Selfobject Responsiveness" (knowing exactly what kind of emotional support the user needs in that moment).

The big discovery is that warmth and substance often come apart. Many AI models got high scores for sounding warm but low scores for actually earning trust. The most common failure mode was substituting "surface warmth" for "substantive support." For example, an AI might say, "Everything will be fine!" when a user is in crisis, which feels nice for a second but shuts down the conversation. A better AI would say, "That sounds incredibly heavy, and it makes sense you're scared," and then wait to see what the user needs next.

The researchers also found that no single AI model dominated the list. Different models had different strengths; some were great at regulating emotions, while others were better at challenging negative thoughts. However, a common weakness across the board was handling "calibrated challenge" and "emotion regulation." The study suggests that while AI has gotten very good at sounding human, it still struggles with the messy, non-linear reality of human relationships.

In short, this paper argues that we need to stop grading AI friends on how "nice" they sound and start grading them on whether they can build real trust. They released their data and code so others can test this too, hoping to guide the development of AI that doesn't just pretend to be a friend, but actually becomes a safe, supportive presence in people's lives. The takeaway is clear: in the world of AI companions, being a good listener is harder than it looks, and saying the right thing is less important than knowing when to say it.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →