A Dual-Task Paradigm to Investigate Sentence Comprehension Strategies in Language Models
This paper proposes a dual-task paradigm combining arithmetic computation with sentence comprehension to demonstrate that restricting cognitive resources in large language models shifts their processing strategies toward human-like rational inference, evidenced by an increased accuracy gap between plausible and implausible sentences.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine your brain is a busy kitchen. Usually, when you're cooking a complex recipe (reading a sentence), you have plenty of counter space and hands to juggle ingredients (words) and follow the instructions (grammar) perfectly. But what happens if someone suddenly asks you to solve a math problem while you're chopping vegetables? Your kitchen gets crowded, your counter space shrinks, and you might start taking shortcuts.
This paper is about testing exactly that scenario, but with Artificial Intelligence (AI) instead of humans. The researchers wanted to see if AI models, when their "mental kitchen" gets crowded, start thinking more like humans do when they are tired or distracted.
The Big Idea: The "Dual-Task" Test
The researchers created a game called a Dual-Task Paradigm. Think of it like this:
- Task 1 (The Sentence): You have to read a sentence and answer a question about it.
- Example: "The cocktail blended the bartender." (This sounds weird! Cocktails don't blend people; bartenders blend cocktails.)
- Task 2 (The Math): Hidden inside that sentence are little math problems you have to solve.
- Example: "The 5 cocktail + blended 3 the = bartender x1..."
The researchers tested the AI in three different modes:
- Solo Mode: Just read the sentence. No math.
- Noisy Mode: Read the sentence with math words in it, but ignore the math. Just pretend they aren't there.
- Dual Mode: Read the sentence AND actually solve the math problems as you go.
The "Rational Inference" Shortcut
Here is the fascinating part. Humans have a habit called Rational Inference. When we are under pressure (like when our brain is full), we sometimes stop paying close attention to the strict rules of grammar and start guessing based on what makes sense in the real world.
If you are very tired and someone says, "The cocktail blended the bartender," your tired brain might skip the grammar error and think, "Oh, they must mean the bartender blended the cocktail," because that's the only thing that makes sense.
The researchers asked: Do AI models do this too when they are busy?
What They Found
They tested several AI models (like GPT-4o, o3-mini, and o4-mini). Here is what happened:
- In Solo Mode: The AI was very strict. It looked at the sentence "The cocktail blended the bartender" and said, "No, that's grammatically wrong. The bartender blended the cocktail." It followed the rules perfectly.
- In Dual Mode (The Busy Kitchen): When the AI had to solve math problems at the same time, it started acting more like a tired human. It began to ignore the weird grammar and focus on the meaning.
- It became much better at realizing that "The cocktail blended the bartender" was a mistake and that the intended meaning was likely the reverse.
- It started prioritizing common sense over strict grammar rules.
However, not all AIs did this. Some models (like GPT-4.1 or Llama) didn't change their behavior much; they just got worse at everything when they were busy. But the smartest models (GPT-4o, o3-mini, o4-mini) specifically shifted to this "human-like" shortcut strategy.
Why Does This Matter?
The paper suggests that being human isn't just about having a big brain; it's about how we manage a small one.
When our resources (memory and attention) are limited, we naturally switch to a "good enough" strategy. We rely on what we know about the world rather than analyzing every single word. The study shows that AI models do the exact same thing. When you force an AI to split its attention between math and reading, it stops being a rigid robot and starts using "rational inference," just like a human would when their working memory is overloaded.
The Takeaway
This study proves that constraints make AI more human. By limiting an AI's ability to process everything perfectly at once, we actually see it adopt the same clever (and sometimes error-prone) shortcuts that humans use to get by in a busy world. It suggests that the "human-like" way of understanding language isn't a special magic trick, but a natural result of trying to do too many things at once with limited mental space.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.