← Latest papers
🤖 AI

A criterion for Artificial General Intelligence: hypothetic-deductive reasoning, tested on ChatGPT

The paper proposes hypothetic-deductive reasoning as a fundamental criterion for Artificial General Intelligence and demonstrates through testing that current models like ChatGPT possess only limited capacity for this type of reasoning in complex contexts.

Original authors: Louis Vervoort, Vitaliy Mizyakov, Anastasia Ugleva

Published 2026-06-23
📖 4 min read☕ Coffee break read

Original authors: Louis Vervoort, Vitaliy Mizyakov, Anastasia Ugleva

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Question: Is the AI Actually "Thinking"?

Imagine you have a student who has read every book in the library. They can recite facts, solve math problems, and write essays that sound incredibly smart. But if you ask them, "How did you figure that out?" or "Why does that happen?", they might stumble. They might give the right answer by guessing the pattern of words, but they don't actually understand the rules behind the answer.

This paper asks: Does ChatGPT really "think," or is it just a very good guesser?

The authors, Louis Vervoort and his team, argue that to be considered a true "thinking machine" (or Artificial General Intelligence, known as AGI), an AI must master a specific skill called Hypothetic-Deductive Reasoning.

The Two Key Skills

The paper breaks "thinking" down into two main abilities:

1. Causal Reasoning (The "Domino" Test)

  • What it is: Understanding that if you push the first domino, the last one falls. It's about knowing what causes what.
  • The Paper's Test: The researchers used "neuron diagrams" (like complex flowcharts of dominoes). They asked the AI: "If I pull this lever, which other levers will move, and which ones won't?"
  • The Result: The AI was okay at simple chains, but when the dominoes got tangled (with some levers blocking others), the AI got confused. It often missed the root cause or got the sequence wrong. It was like a person who knows the word "cause" but can't actually trace the path of a falling domino in a messy room.

2. Hypothetic-Deductive Reasoning (The "Recipe" Test)

  • What it is: This is the "Scientist" way of thinking. It involves two steps:
    1. Hypothesize: Pick the right rulebook (theory) for the problem.
    2. Deduce: Follow the steps in that rulebook to get the answer.
  • The Paper's Test: The researchers asked the AI physics and logic puzzles (e.g., "If I push a row of cubes where one is made of ice, what happens when the room gets hot?").
  • The Secret Weapon: The researchers didn't just ask for the answer. They asked the AI to list the rules (hypotheses) it used to get that answer.
    • Analogy: Imagine asking a chef, "How did you make this soup?" If they say, "I just threw things in," they might be lucky. If they say, "I used the rule that salt dissolves in hot water and the rule that onions cook faster than carrots," they are actually thinking.

What Happened When They Tested ChatGPT?

The researchers tested both the older version (ChatGPT-3.5) and the newer, smarter version (ChatGPT-4).

  • The "Right Answer, Wrong Reason" Problem:
    ChatGPT-4 often gave the correct answer to the puzzles. It was very impressive! However, when asked to list the rules it used, it often made things up.

    • Example: In one test, the AI correctly guessed that a bucket of water would spill. But when asked why, it invented a fake rule about "insulating shirts" that had nothing to do with the actual physics of the situation.
    • The Metaphor: It's like a student who gets the math problem right because they memorized the final number, but when asked to show their work, they write down a formula they made up on the spot.
  • The Verdict:
    The paper concludes that while ChatGPT is amazing at mimicking human conversation and solving simple problems, it currently lacks genuine understanding.

    • It is like a parrot that can say "The sky is blue" perfectly, but doesn't understand why the sky is blue or what "blue" actually means in a physics sense.
    • When the problems got slightly complex (combining different ideas), the AI started to fail. It couldn't consistently link the right rules to the right answers.

The "AGI" Bar

The authors set a high bar for what counts as "Artificial General Intelligence" (AGI). They argue that an AI isn't truly "thinking" until it can:

  1. Solve a problem.
  2. Explicitly list the rules and theories it used to solve it.
  3. Do this correctly across many different types of problems (not just one).

The Conclusion:
Right now, ChatGPT is a brilliant improviser. It can "wing it" and sound smart. But it hasn't yet proven it has a deep, logical map of how the world works. To be a true "thinking machine," it needs to stop just guessing the next word and start proving why that word is the right one, step-by-step, using real rules.

Until an AI can pass these "list the rules" tests consistently, the paper says it is not yet a true AGI.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →