← Latest papers
💬 NLP

TEMPER: Testing Emotional Perturbation in Quantitative Reasoning

The paper introduces TEMPER, a benchmark and framework demonstrating that emotional framing in quantitative reasoning tasks significantly degrades large language model performance even when numerical content is preserved, while showing that neutralizing such emotional language effectively recovers accuracy.

Original authors: Atahan Dokme, Benjamin Reichman, Larry Heck

Published 2026-04-10
📖 4 min read☕ Coffee break read

Original authors: Atahan Dokme, Benjamin Reichman, Larry Heck

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: Can AI Think Clearly When It's Angry?

Imagine you ask a very smart, well-trained robot to solve a math problem.

  • Scenario A: You ask politely: "If I have 5 apples and buy 3 more, how many do I have?"
  • Scenario B: You scream in frustration: "Ugh, I'm so mad! I had 5 rotten apples, then I bought 3 more terrible ones! How many of these awful apples do I have now?!"

In both cases, the math is exactly the same (5+3=85 + 3 = 8). The numbers haven't changed. But in Scenario B, the robot is surrounded by "emotional noise."

The paper asks: Does the robot get confused by the tone of the voice, even if the facts are perfect?

The answer is a resounding YES. The researchers found that when they wrapped math problems in emotional language (like anger, disgust, or fear), even the smartest AI models made more mistakes.


The Experiment: The "Emotional Translator" Factory

To test this, the researchers couldn't just ask humans to yell at the AI (that would be chaotic). Instead, they built a special Emotional Translator Factory.

  1. The Teacher: They trained a "Teacher" AI that is an expert at spotting emotions (like a drama critic).
  2. The Student: They trained a "Student" AI to rewrite math problems.
  3. The Trick: The Student learned to take a boring math problem and rewrite it to sound like a grumpy teenager, a terrified person, or an over-enthusiastic fan, without changing a single number.

They created 5,400 pairs of problems. For every original math question, they had an "Emotional Version" and a "Neutral Version."

The Results: The "Emotional Hangover"

When they tested 18 different AI models (from small ones to the massive "frontier" models like GPT-4 and DeepSeek), they found two major things:

1. The "Emotional Hangover" (Accuracy Drops)

Just like a human might make a simple math error if they are screaming in anger, the AIs got worse at math when the problem was emotional.

  • The Drop: Accuracy fell by 2% to 10%.
  • The Culprit: The emotion "Disgust" was the worst offender. If a problem sounded "gross" or "revolting," the AI was most likely to fail. "Fear" was second. "Joy" and "Surprise" were the least distracting.
  • The Analogy: Imagine trying to do your taxes while someone is playing heavy metal music and screaming in your ear. Even if the numbers on the page are clear, your brain gets distracted by the noise. The AI's "brain" got distracted by the emotional words.

2. The "Magic Eraser" (Neutralization)

Here is the cool part. The researchers took those angry, emotional problems and ran them through a "Neutralizer" to strip away the emotion, leaving just the math.

  • The Result: The AI's performance bounced back! It recovered most of the lost points.
  • What this proves: This confirmed that the AI didn't fail because the math was broken. It failed because the style of the text messed up its thinking process. If you "clean the noise," the AI can think clearly again.

Why This Matters: The "Hidden Brain" Shift

The researchers looked inside the AI's "brain" (its internal data layers) to see what was happening.

  • The Metaphor: Imagine the AI's brain is a library.
    • When you ask a normal math question, the books are neatly organized on the shelves.
    • When you ask an emotional question, the books get thrown into a chaotic pile. The AI has to spend extra energy digging through the mess to find the right numbers.
    • The study showed that emotional text pushes the AI's internal "thoughts" 3 to 4 times further away from the correct path than just changing the words normally would.

The Takeaway for Real Life

We often think AI is like a calculator: cold, logical, and immune to feelings. This paper shows that AI is actually quite sensitive to human emotion.

  • For Users: If you are frustrated with an AI, screaming at it (or typing in all caps with angry emojis) might actually make it less helpful, even if you are asking for a simple calculation.
  • For Developers: We need to build AI that can "filter out" emotional noise, just like a noise-canceling headphone filters out background chatter, so it can focus on the math.

Summary in One Sentence

Even though AI is a machine, wrapping a math problem in angry or disgusted words confuses its brain and makes it make mistakes, but if you strip away the emotion, it remembers how to think clearly again.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →