← Latest papers
💻 computer science

AIPsy-Affect: A Keyword-Free Clinical Stimulus Battery for Mechanistic Interpretability of Emotion in Language Models

This paper introduces AIPsy-Affect, a 480-item clinical stimulus battery featuring keyword-free emotional vignettes and matched neutral controls to enable rigorous mechanistic interpretability of emotion in large language models by eliminating the confound of emotion-specific vocabulary.

Original authors: Michael Keeman

Published 2026-04-28
📖 4 min read☕ Coffee break read

Original authors: Michael Keeman

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to figure out how a super-smart robot (a Large Language Model) "feels" when it reads a story. You want to know: Does the robot understand the emotion of the situation, or is it just spotting specific "emotion words" like angry, sad, or happy?

This is the big problem the paper solves.

The Problem: The "Keyword Trap"

Think of previous experiments like this: You show the robot a story that says, "I am furious!" and the robot's internal alarm goes off for "anger."

  • The Question: Did the robot understand the feeling of fury? Or did it just see the word "furious" and press the button?
  • The Issue: In almost all past tests, the stories were full of emotion words. It's impossible to tell if the robot is smart about feelings or just a good word-spotter. It's like testing if a dog understands the concept of "fetch" by only using the word "fetch," rather than actually throwing a ball.

The Solution: AIPsy-Affect (The "Keyword-Free" Test)

The authors created a new set of 480 stories called AIPsy-Affect. Think of this as a "magic trick" for testing robots.

1. The "Situation-Only" Stories (The Clinical Items)
They wrote 192 stories that describe a situation that should make you feel a specific emotion (like rage, grief, or joy), but they strictly banned any words that name that emotion.

  • Example: Instead of saying "He was furious," they write: "He threw the papers on the floor. The bank statement showed zero balance. His brother had signed the withdrawal form."
  • The story screams "betrayal and anger" through the events, but the word "anger" is nowhere to be found.

2. The "Boring Twin" Stories (The Matched Controls)
For every emotional story, they created a "twin" story that is almost identical in length, characters, and setting, but emotionless.

  • Example: "He sorted the papers on the table. The bank statement showed the balance grew by 4%. His brother had signed the deposit form."
  • Same characters, same objects, same length. The only difference is that the "betrayal" is gone, replaced by a boring routine.

3. The "Complex Boring" Stories
They also added 48 stories about very detailed, technical things (like a machine cutting metal or a ship sailing) to make sure the robot isn't just getting excited because the story is "long and detailed" rather than "emotional."

How They Tested It (The "Three-Layer Defense")

The authors didn't just trust their own writing; they ran the stories through three different "detectives" to prove the stories were truly keyword-free:

  1. The Simple Word Counter (VADER): This tool looks for "happy" or "sad" words. It found some difference between the emotional and boring stories, but when they checked which words caused the difference, they were just situational words like "money," "death," or "approved." No actual emotion words were found.
  2. The Emotion Dictionary (NRC): This tool looks for specific emotion categories (like "fear" or "joy"). It found zero difference. The emotional stories actually had fewer emotion words than the boring ones!
  3. The Smart AI Classifier (GoEmotions): This is a very smart AI trained to spot emotions.
    • Result A: It could tell the difference between the "emotional" stories and the "boring" ones (it knew something was up).
    • Result B: It could not guess which emotion it was. It got the category wrong 95% of the time.
    • The Takeaway: The smart AI could sense "something emotional is happening" based on the situation, but it couldn't name the feeling because the "name tags" (keywords) were missing.

Why This Matters

This dataset is a tool for scientists, not a product for doctors or therapists.

  • What it allows: Researchers can now test if a robot has a true "internal feeling circuit" or if it's just a word-matching machine. If a robot can tell the difference between the "furious" story and the "boring twin" story without using the word "furious," then we know the robot has learned to understand the situation, not just the vocabulary.
  • What it is NOT: The paper does not claim this will help robots feel real emotions, nor does it claim this will help diagnose human mental health. It is purely a "stress test" to see how current AI models process emotional information internally.

In a Nutshell

The authors built a set of stories that are emotionally loud but linguistically silent. By removing the "cheat codes" (emotion words), they created a fair test to see if AI models are actually understanding human feelings or just memorizing a dictionary. The results show that while AI can sense the presence of emotion in these stories, it struggles to categorize them without the help of specific keywords.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →