← Latest papers
🤖 AI

How Does Thinking Mode Change LLM Moral Judgments? A Controlled Instant-vs-Thinking Comparison Across Five Frontier Models

This study finds that while enabling reasoning modes in frontier LLMs does not significantly alter their aggregate binary moral verdicts, it notably reduces cross-model disagreement and demographic inconsistencies in specific contentious scenarios while more frequently shifting the models' self-labeled ethical frameworks.

Original authors: Sai Sourabh Madur

Published 2026-05-07
📖 5 min read🧠 Deep dive

Original authors: Sai Sourabh Madur

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have five different super-smart robots (AI models) from five different tech companies. You ask them a series of 100 tricky moral questions, like "Is it okay to push one person off a bridge to save five others?" or "Should a doctor prioritize a patient based on their job or their age?"

The researchers wanted to see what happens when they tell these robots to "think before they speak."

Normally, the robots answer instantly (like a reflex). In "Thinking Mode," the robots pause, generate a long internal reasoning process, and then give their answer. The paper asks: Does taking that extra time to think actually change their moral compass, or do they just say the same thing but with a longer explanation?

Here is the breakdown of what they found, using simple analogies:

1. The Big Picture: The Crowd Still Agrees

When you look at the robots' answers as a whole group, taking time to think didn't really change the final verdict.

  • The Analogy: Imagine a jury of five people. Whether they shout their answer immediately or whisper it to each other for a minute first, they still end up agreeing on the verdict about 78% of the time. The "Thinking Mode" didn't make them suddenly agree or disagree more often as a group.

2. The "Hard" Cases: Thinking Helps a Little Bit

However, the story changes when you look at the really difficult, controversial questions (like complex trolley problems or modern ethical dilemmas).

  • The Analogy: In the "Instant Mode," the five robots were basically guessing on these hard questions, agreeing with each other only about as often as if they were flipping coins.
  • The Result: When they switched to "Thinking Mode," they started to agree with each other a bit more. It wasn't a magic fix, but it was like they stopped shouting random guesses and started finding a little more common ground.
  • The Catch: The researchers warn that this improvement is "directional" (it went the right way) but not statistically rock-solid yet. It's a hint, not a guarantee.

3. The "Reasoning" vs. The "Verdict"

Here is the most interesting twist: The robots changed their reasons much more often than they changed their answers.

  • The Analogy: Imagine two people both saying, "No, you can't have that cookie."
    • Person A (Instant) says: "Because I said so."
    • Person B (Thinking) says: "No, because cookies are bad for your teeth and you promised to eat veggies first."
    • They both said "No," but their internal logic shifted.
  • The Finding: The robots frequently changed the ethical framework they claimed to use (e.g., switching from "Utilitarian" to "Deontological") even when their final Yes/No answer stayed the same. It's like a lawyer changing their legal argument in the middle of a trial but still asking for the same verdict.

4. The "Demographic" Test: Thinking Makes Them Fairer

The researchers tested if the robots judged people differently based on their race, gender, or nationality.

  • The Analogy: Imagine a robot judging a crime. In "Instant Mode," it might give a harsher sentence if the criminal is described as an "immigrant" versus a "citizen."
  • The Result: When three of the five robots switched to "Thinking Mode," they became much more consistent. They stopped letting the person's background sway their judgment as much. It's like the "thinking" process acted as a filter that removed the noise of bias.

5. The "Apples vs. Oranges" Problem (A Major Caveat)

The paper has a huge warning label: We didn't actually compare the same amount of "thinking."

  • The Analogy: Imagine asking five students to solve a math problem.
    • Student A (Claude) spends 1 minute thinking.
    • Student B (Qwen) spends 80 minutes thinking.
    • Student C (GPT) spends 2 minutes.
  • The Reality: The researchers couldn't make them all "think" for the exact same amount of time because the companies control the buttons differently. One robot might be doing a quick mental check, while another is writing a whole essay in its head. So, when we say "Thinking Mode," we are comparing a light stretch to a heavy workout depending on which robot you use.

Summary of the 5 Hypotheses

The researchers had five guesses before they started. Here is what happened:

  1. Do different robots have different moral defaults? Yes. (Some lean more on "greatest good," others on "rules").
  2. Does changing the wording of a question change the answer? No. (The robots were surprisingly good at ignoring tricky wording).
  3. Does the robot's answer change based on the person's background? Sometimes. (Instant mode was inconsistent; thinking mode fixed this for some robots).
  4. Do robots agree on easy questions but fight on hard ones? Yes. (They agreed on simple stuff, but were confused on the hard stuff until they thought).
  5. Does "Thinking Mode" change the moral judgment? Sort of. (It rarely changed the final Yes/No, but it often changed why they said it, and it helped them agree slightly more on the hardest questions).

The Bottom Line

Turning on "Thinking Mode" doesn't turn the robots into completely new moral beings. They generally give the same "Yes" or "No" answers. However, thinking does seem to help them:

  1. Be slightly more consistent with each other on the hardest questions.
  2. Be less biased against specific groups of people.
  3. Change their internal logic even if the final answer stays the same.

The researchers released all their data (the questions, the answers, and the robots' "thought processes") so other scientists can double-check this work.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →