Presupposition and Reasoning in Conditionals: A Theory-Based Study of Humans and LLMs
This study reveals that while large language models can sometimes match human ratings on presupposition projection in conditionals, their performance often stems from surface pattern matching rather than genuine pragmatic reasoning, highlighting a disconnect between statistical alignment and true linguistic competence.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Idea: The "If-Then" Puzzle
Imagine you are playing a game of logic with a friend. You say: "If John is a scuba diver, he will bring his wetsuit."
Your friend asks, "Does John actually have a wetsuit?"
- The Human Answer: Most humans say, "Probably not yet. He only has a wetsuit if he is diving." The wetsuit is tied to the condition.
- The Other Human Answer: Now imagine you say: "If John flies to London, his sister will pick him up."
- The Human Answer: Here, humans say, "Yes, John definitely has a sister." The fact that he has a sister doesn't depend on the flight; it's just a fact about John.
This difference is called the "Proviso Problem." It's a tricky puzzle in linguistics about how we decide if a hidden fact (a "presupposition") is always true or only true if the "If" part happens.
The Experiment: Humans vs. AI
The researchers wanted to see if Large Language Models (LLMs)—the brains behind chatbots like the one you are talking to—can solve this puzzle the same way humans do.
They set up a "taste test" with two groups:
- 120 Humans: Real people who read sentences and rated how likely a hidden fact was to be true on a scale of 0 to 7.
- 4 AI Models: Different versions of AI (like GPT-5, Gemini, Llama, and Qwen) that did the exact same task.
They tested the AI in two ways:
- The "Bare" Test: Just the sentence (e.g., "If John is a diver...").
- The "Context" Test: The sentence plus a tiny bit of background info (e.g., "John is from Ottawa. If John is a diver...").
The Results: The "Imposter" AI
Here is what they found, broken down into simple metaphors:
1. Humans are Intuitive Detectives
Humans are great at this. They mix logic with common sense. If the "If" part makes the hidden fact more likely (like a diver needing a wetsuit), they lower their rating. If the "If" part has nothing to do with the fact (like flying to London and having a sister), they keep the rating high. They are sensitive to how relevant the two parts of the sentence are to each other.
2. The AI "Mimics" the Answer, But Not the Reasoning
This is the most surprising part.
- The Small AI (Qwen): This model gave answers that looked very similar to the humans. If humans said "7," the AI said "6.8." It was a great mimic.
- The Big AI (GPT-5, Gemini): These models gave answers that were less like humans. They were often too high or too low, missing the subtle shifts humans made.
The Twist: When the researchers asked the AIs to explain why they gave those answers (like asking a student to show their math work), a weird pattern emerged:
- The Small AI (Qwen), which gave the best human-like answers, actually had terrible reasoning. It couldn't explain the logic properly. It was just guessing based on patterns it saw in its training data, like a parrot repeating a phrase it heard often without knowing what it means.
- The Big AI (GPT-5), which gave worse human-like answers, actually had better reasoning. It could explain the rules of logic well, but when it came to the final number, it got it wrong compared to humans.
The "Checklist" Analogy
To figure this out, the researchers used a "Judge AI" (a referee) with a checklist. Imagine a teacher grading a student's essay.
- The Checklist: "Did you identify the trigger? Did you check if the context matters? Did you explain the logic?"
- The Result: The "Big AI" passed the checklist (it wrote a good essay). The "Small AI" failed the checklist (its essay was messy).
- The Paradox: Even though the "Small AI" failed the logic test, its final score was closer to the human average.
The Conclusion: Pattern Matching vs. Understanding
The paper concludes that the AI models aren't actually "understanding" the language the way humans do.
- Humans use a mix of logic, world knowledge, and relevance to decide if a fact is true.
- AI seems to be doing surface-level pattern matching. It looks at the words "If John is a diver" and "wetsuit," sees that they often appear together in its training data, and guesses the number. It doesn't truly "know" that the wetsuit is conditional on the diving.
The Takeaway: Just because an AI gives you an answer that looks right (like a student who memorized the answer key), it doesn't mean it understands the lesson. In fact, the AI that gave the most "human-like" answers was actually the one relying the least on real logic.
What the Paper Does NOT Say
- It does not say AI is useless.
- It does not say AI will never understand language.
- It does not suggest using this for medical diagnosis or legal advice.
- It simply says: Right now, when it comes to this specific type of tricky logic puzzle, AI is faking it better than it thinks, and the "smartest" looking answers might be the most superficial.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.