← Latest papers
🤖 AI

Mind the (DH) Gap! A Contrast in Risky Choices Between Reasoning and Conversational LLMs

This study reveals a significant divergence in risky decision-making between reasoning and conversational large language models, finding that reasoning models exhibit rational, context-insensitive behavior similar to optimal agents, while conversational models display human-like biases and a large gap between explicit descriptions and outcome histories.

Original authors: Luise Ge, Yongyan Zhang, Yevgeniy Vorobeychik

Published 2026-04-22
📖 5 min read🧠 Deep dive

Original authors: Luise Ge, Yongyan Zhang, Yevgeniy Vorobeychik

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are hiring two different types of AI assistants to help you make a high-stakes investment decision. You have two choices: a Math Whiz (a "Reasoning Model") and a Chatty Friend (a "Conversational Model").

This paper is essentially a report card on how these two types of AI handle risk, uncertainty, and tricky questions. The researchers put them through a series of tests to see if they act like cold, calculating robots, emotional humans, or something in between.

Here is the breakdown of their findings, using some everyday analogies.

1. The Two Teams: The Accountant vs. The Barista

The study found that all the AI models they tested naturally split into two distinct groups:

  • The Reasoning Models (RMs): Think of these as stiff, hyper-logical accountants. They are trained heavily on math and logic. When you ask them to choose between a risky bet and a safe one, they act like a perfect computer program. They ignore distractions, don't get confused by how the question is worded, and always choose the option with the highest average payout. They are "rational" in the strict economic sense.
  • The Conversational Models (CMs): Think of these as chatty baristas or friendly neighbors. They are great at talking, writing stories, and sounding human. However, when it comes to math and risk, they are much more "human-like" in their flaws. They get swayed by how a question is framed, they get confused by the order of options, and they often make choices that don't maximize their winnings.

The Big Surprise: The "Chatty" models are actually more like real humans than the "Math Whiz" models are. The Math Whizzes are so logical they barely resemble us at all.

2. The "Description vs. Experience" Gap

This is the paper's most fascinating finding, called the DH Gap.

Imagine you are offered a gamble:

  • Scenario A (Description): I tell you, "There is a 70% chance you win $100 and a 30% chance you win nothing."
  • Scenario B (Experience): I don't tell you the odds. Instead, I show you a list of the last 20 times this game was played: "You won $100 fifteen times, and got $0 five times."

How Humans React: Humans usually act differently in these two scenarios. We tend to be more cautious when we just read the odds, but we take more risks when we see the history of wins. This is a known human quirk.

How the AI Reacted:

  • The Math Whizzes (RMs): They didn't care. Whether you gave them the odds or the history, they calculated the math and picked the same winner every time. They were immune to the "gap."
  • The Chatty Models (CMs): They fell right into the trap. When shown the history of wins, they started acting much more like humans—taking risks they wouldn't have taken if you just gave them the numbers. They have a huge "gap" between how they think when reading vs. when "experiencing" data.

3. The "Explain Yourself" Effect

The researchers also asked the AI to explain why they made a choice.

  • The Result: Surprisingly, asking for a short, one-sentence explanation actually made the AI (especially the Chatty ones) make better, more rational decisions. It was like asking a student to "show their work" on a math test; it forced them to slow down and think logically.
  • However, if they asked for a long, complex mathematical explanation, the Chatty models sometimes got confused and made worse choices, likely because they were trying to "sound smart" rather than actually being smart.

4. What Makes the Difference?

Why are some AIs "Math Whizzes" and others "Chatty Baristas"?
The study found that the key isn't just how big the model is (how many "brain cells" it has). The key is training.

  • If a model is specifically fine-tuned for mathematical reasoning, it becomes a "Reasoning Model" (RM). It becomes a cold, hard calculator.
  • If a model is just trained to chat and follow instructions without that specific math focus, it stays a "Conversational Model" (CM).

The Takeaway

If you are building an AI to manage your money, run a stock portfolio, or make critical safety decisions, you want the Math Whiz (Reasoning Model). They are consistent, logical, and won't get tricked by how you phrase a question.

If you want an AI to chat with customers, write stories, or simulate human behavior for a movie script, the Chatty Model is great. But be careful: if you ask them to make a financial decision, they might act irrationally, just like a human would.

In short: The paper warns us that not all AIs are created equal. Some are built to be perfect robots, while others are built to be imperfect humans. Knowing which is which is crucial before you let them make decisions for you.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →