← Latest papers
💬 NLP

CoSToM:Causal-oriented Steering for Intrinsic Theory-of-Mind Alignment in Large Language Models

The paper introduces CoSToM, a framework that enhances Large Language Models' intrinsic Theory-of-Mind capabilities by using causal tracing to identify critical internal layers and applying targeted activation steering to align internal knowledge with stable, high-quality social reasoning behaviors.

Original authors: Mengfan Li, Xuanhua Shi, Yang Deng

Published 2026-04-14
📖 4 min read☕ Coffee break read

Original authors: Mengfan Li, Xuanhua Shi, Yang Deng

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Problem: The "Robot Actor" vs. The "Real Thinker"

Imagine you are watching a very talented actor on stage. The script says, "The character is sad," and the actor immediately starts crying. They look perfect. But if you ask them, "Why are you sad?" while they are off-stage, they might say, "I don't know, the script just told me to cry."

This is the current state of many Large Language Models (LLMs) when it comes to Theory of Mind (ToM). ToM is the human ability to understand that other people have their own thoughts, beliefs, and desires that are different from your own.

Current AI models are great at mimicking this behavior when you give them a specific prompt like, "Pretend you are a negotiator who understands the other person's feelings." But if you remove that prompt and just say, "Negotiate with me," the AI often forgets everything. It stops thinking about the other person's feelings and starts spitting out random or rude answers.

The researchers ask: Does the AI actually "know" what the other person is thinking deep inside its brain, or is it just acting on command?

The Solution: COSTOM (The "Mind-Steering" System)

The authors created a framework called COSTOM (Causal-oriented Steering for ToM). Think of it as a way to tune the AI's brain so it actually thinks about other people's feelings, rather than just pretending to.

They did this in three steps:

1. The Detective Work: Finding the "Empathy Switch"

First, they needed to find where in the AI's massive brain the concept of "other people's feelings" lives.

  • The Analogy: Imagine a giant library (the AI) with millions of books (layers of data). You want to find the specific shelf where the "Empathy" books are stored.
  • The Discovery: Using a technique called Causal Tracing, they acted like detectives. They "plugged in" a probe to different parts of the AI's brain while it read a conversation. They found that the AI does understand feelings, but it happens very early in its processing—like the first few pages of a book. If you wait until the end of the book (the deeper layers), the AI forgets the feelings and just focuses on finishing the sentence.

2. The Tuning: "Steering" the Brain

Once they found the "Empathy Switch" (the early layers), they didn't want to rebuild the whole library. They wanted to just tweak that specific shelf.

  • The Analogy: Imagine a car engine. Usually, to make a car go faster, you might try to rebuild the whole engine (Full Fine-Tuning). But COSTOM is like adding a small, precise turbocharger just to the air intake valve.
  • How it works: They used a "Gradient Bridge." Imagine a teacher standing next to the AI. When the AI thinks about a feeling, the teacher checks if it's right. If the AI is wrong, the teacher sends a signal backwards through the brain, but only to the specific "Empathy Switch" layers found in step 1. This "steers" the AI to keep those feelings active and strong, even as it moves to the later stages of generating a response.

3. The Result: A Natural Conversationalist

Finally, they tested the tuned AI in real conversations (like negotiating for firewood or persuading someone to buy a violin).

  • The Outcome: Before COSTOM, the AI was like a robot that needed a reminder to be nice. After COSTOM, the AI became like a natural human. It didn't need a prompt saying "be empathetic." It naturally understood what the other person wanted and negotiated or persuaded them effectively.
  • The Magic: It wasn't just "better at math." It was better at socializing. It could maintain a coherent, friendly, and strategic conversation without getting confused or rude.

Why This Matters

Most previous methods tried to fix AI by giving it better instructions (Prompting) or retraining the whole model (Fine-tuning).

  • Prompting is like telling a student, "Remember to be polite!" It works for a second, but they might forget.
  • Full Fine-tuning is like making the student re-learn the entire school curriculum just to be polite. It's expensive and slow.

COSTOM is like giving the student a permanent, internal habit of empathy. It changes the way the AI thinks at a fundamental level, ensuring that understanding other people's minds becomes a stable part of its personality, not just a trick it performs when asked.

Summary in One Sentence

The researchers found the specific "empathy circuits" inside AI brains, gave them a gentle nudge to stay active, and turned robotic actors into genuine social thinkers who can negotiate and persuade naturally.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →