Humans Disengage, Reasoning Models Persist: An Item-Controlled Dissociation in Deliberation Allocation
This paper reveals a critical dissociation between humans and large reasoning models (LRMs) in deliberation allocation: while both spend more time on harder problems overall, humans invest less time on items they ultimately get wrong (abandoning difficult tasks), whereas LRMs paradoxically generate longer reasoning traces on problems they fail (persisting in uncertainty), highlighting a fundamental difference in stopping and control policies despite surface-level similarities.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are watching two different students take a really hard math test. One is a brilliant human teenager, and the other is a super-smart computer program designed to "think" through problems step-by-step. For a long time, scientists have been fascinated by how these two compare. They've noticed a surface-level similarity: when a problem is tricky, both the human and the computer tend to spend more time on it. It's like watching a car's speedometer; if the road gets steep, both drivers slow down to climb it. This observation led many to believe that computers are starting to "think" just like people do, using similar mental gears to figure out which problems are hard and which are easy.
But here is the twist: just because two drivers slow down on the same steep hill doesn't mean they are driving the same way. One driver might slow down to carefully navigate the turn, while the other might slow down because they are confused and staring at the map, unsure of what to do next. The big question in the world of artificial intelligence and human psychology is: how do these thinkers decide when to keep trying and when to give up? Do they spend extra effort on the problems they are struggling with, or do they cut their losses and move on? Understanding this isn't just about math scores; it's about figuring out if computers are truly learning to reason like us, or if they are just mimicking the look of thinking without the same internal logic.
The Great "Thinking" Mismatch
A new study by researcher Han-yu Wang peels back the layers of this mystery and finds a surprising secret hiding in plain sight. While humans and advanced "reasoning" AI models agree on which problems are hard, they completely disagree on what to do about it.
Think of it like a video game. When a player hits a boss that is too strong, a smart human player might realize, "This is too hard right now," and decide to stop wasting time, maybe trying a different strategy or moving on to a different level. They disengage. But the AI, according to this study, does the exact opposite. When the AI gets a problem wrong, it doesn't stop; it keeps going, generating even more text, more "thoughts," and more steps before finally giving up the wrong answer.
The Core Discovery: The "Wrong" vs. "Right" Gap
The study looked at a massive collection of data where humans and six different AI models solved the same visual puzzles (called H-ARC). The researchers measured two things:
- Registration: Did they agree on which puzzles were hard? Yes. Both humans and AI took longer on the hard puzzles.
- Allocation: When they got a puzzle wrong, did they spend more or less time on it compared to when they got it right? No. This is where they split.
- Humans: When humans got a puzzle wrong, they spent less time on it than when they got it right. The data shows a "Cohen's d" of -0.10. This suggests a "give-up" strategy. If a human feels they are stuck, they often stop thinking about it quickly.
- AI Models: When the AI got a puzzle wrong, it spent significantly more time on it. The AI models showed a "Cohen's d" ranging from 1.47 to 3.13. This means the AI kept churning out "thoughts" (tokens) even when it was clearly failing.
The paper argues that this isn't just a small glitch; it's a fundamental difference in how they "think." The human pattern looks like engagement vs. abandonment: we stay on problems we think we can solve and quit the ones we can't. The AI pattern looks like length-on-uncertainty: when the AI is unsure, it just keeps talking, generating longer chains of reasoning, even though that extra talking doesn't help it get the right answer.
The "Trace" Clues
To understand why the AI does this, the researchers looked at the actual text the AI generated (its "thought trace"). They found that when the AI was about to get a question wrong, its text was full of "hedging" and "self-doubt." It would say things like "wait, actually," "hmm," or "let me reconsider." It was essentially talking to itself, unsure of the path, but it didn't stop.
In contrast, when humans got a question wrong quickly, they often made fewer moves (like fewer grid edits in the puzzle), suggesting they had already decided the puzzle wasn't worth the effort. The study suggests that for humans, the decision to stop is a smart, metacognitive choice. For the AI, the decision to stop seems to be tied to a different rule: it keeps going until it runs out of "budget" or hits a wall, even if it knows (or should know) it's going the wrong way.
What This Means (and What It Doesn't)
The paper is very careful not to say the AI is "stupid" or that humans are "perfect." Instead, it suggests that the AI and humans are playing by different rulebooks.
- The AI isn't necessarily failing to think: It might be following a rule that says, "If you are unsure, generate more text." This might be a good strategy for some tasks, but for these puzzles, it leads to wasted effort.
- The "Thinking" isn't the same: Just because the AI produces a long chain of text that looks like human reasoning doesn't mean it's using the same mental stopping rules. The study shows that the AI's "thinking" is more like a hamster running on a wheel when it's confused, while the human is like a hiker who decides to turn back when the trail gets too steep.
The researchers also tested this on other types of puzzles (like intuitive physics and binary reasoning). The pattern held up: the AI kept spending extra time on its mistakes, while humans tended to spend less time. However, in one specific type of puzzle (Cortes), the AI's behavior changed slightly, showing that the "length-on-uncertainty" rule might depend on the specific game being played.
The Bottom Line
This study doesn't prove that AI can't think; it proves that AI and humans allocate their thinking differently. The surface similarity (both taking longer on hard things) hides a deep difference: humans tend to stop thinking when they realize they are stuck, while AI models tend to keep thinking (and talking) even when they are stuck.
The authors suggest that to truly understand AI, we need to look at when they stop, not just how long they think. They propose a new way to test this: if we cut off the AI's "thinking" early (truncation), we might see if those extra words were actually helping or just noise. Until we do that, the paper concludes that the AI's "persistence" on wrong answers is a sign that its internal logic for deciding "when to quit" is fundamentally different from our own. It's a reminder that in the race to build human-like AI, we might be building machines that are very good at looking like they are thinking, but very bad at knowing when to stop.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.