← Latest papers
🤖 AI

Human-Level Reasoning: A Comparative Study of Large Language Models on Logical and Abstract Reasoning

This study evaluates and compares the logical and abstract reasoning capabilities of nine prominent Large Language Models against human performance using eight custom-designed questions, revealing significant gaps in the models' deductive abilities.

Original authors: Benjamin Grando Moreira

Published 2026-06-19
📖 5 min read🧠 Deep dive

Original authors: Benjamin Grando Moreira

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a room full of incredibly talented librarians. These librarians (the Large Language Models, or LLMs) have read almost every book in the world. They can recite facts, write poems, and translate languages faster than anyone else. But the question this paper asks is: Do they actually understand what they are reading, or are they just really good at guessing the next word in a sentence?

To find out, the author, Benjamin Grando Moreira, set up a "logic gym" and invited 15 of these digital librarians to compete against a group of human students and professors.

Here is a simple breakdown of what happened, using some everyday metaphors.

The Test: A Puzzle Box, Not a Trivia Quiz

Instead of asking the models "Who was the first president?" (which they can answer by memory), the researcher gave them 8 tricky puzzles. These puzzles were like riddles that required the models to figure out hidden rules, not just recall facts.

Think of it like this:

  • The Human Brain is like a detective who looks at clues, connects dots, and says, "Aha! I see the pattern!"
  • The AI Brain is like a super-fast pattern-matching machine that says, "I've seen this word before, so I'll guess the next word based on what usually follows."

The 8 Challenges (The "Obstacle Course")

The paper tested the models on things like:

  1. The Day of the Week Riddle: If "SUN + 1 = MON," what is "TUE + 2?" (Most models got this, but some got confused by the abbreviations).
  2. The Elephant Song: A silly pattern where the number of times you say "bother" changes based on whether the number of elephants is odd or even. Humans saw the pattern easily; many AIs got stuck on the repetition.
  3. The Secret Code: A simple letter shift (like A becomes B). Most models cracked this, but some got the capitalization wrong.
  4. The Trick Math Question: "What is 3 + 3 x 5?" The correct answer is 18. But the test gave options like 16, 20, 30, and 45. 18 wasn't even an option!
    • The Twist: Many models (and some humans) got confused. They saw the math was right but the options were wrong, and some just picked the "closest" wrong answer instead of saying, "Hey, none of these are right."
  5. The Month Mystery: This was the hardest. It asked to find a hidden code for the month "May." The code involved squaring the month's number (5 squared = 25) and counting the letters in the word "May" (4 letters), then sticking them together to get 254.
    • The Result: Humans were okay at this, but the AIs mostly failed. They tried to use complex math formulas that didn't fit, rather than seeing the simple "stick the numbers together" trick.
  6. The Language Switch: Matching Portuguese month abbreviations to Spanish names. Humans struggled here if they didn't know Spanish, but the AIs (who know many languages) aced it.
  7. The Bilingual Day: A mix of English and Portuguese days of the week with a hidden rule about when to switch languages. Both humans and AIs found this tricky.
  8. The Binary Code: Adding numbers that look like they are in base-10 (like 1000) but are actually in base-2 (binary). Only the smartest models realized, "Wait, these aren't normal numbers!" and solved it.

The Scoreboard: Who Won?

The researcher gave points for:

  • 10 points: Getting the answer right.
  • 5 points: Getting the logic right, even if the final number was wrong.
  • 0 points: Getting it wrong or giving up.

The Results:

  • The Humans: The professors (the experts) did the best overall, especially on the tricky puzzles. The students did well but struggled more with the abstract riddles.
  • The AIs: On average, the AIs scored about 73 points out of 100, while the humans scored about 70.
    • Wait, didn't the AIs win? Yes, slightly! But here is the catch: The AIs were great at the easy stuff (like knowing months in different languages) and terrible at the weird, abstract stuff (like the elephant song or the hidden number code).
    • The best AI models (like GPT-4o and Claude) scored near perfect, but the weaker ones (like Sabiá 3) scored very low.

The Big Takeaway

The paper concludes that these AI models are like brilliant parrots. They can mimic human conversation and solve standard problems perfectly. However, when you ask them to think outside the box, spot a weird pattern, or realize that a question is a trick, they often stumble.

  • Humans are better at "fluid intelligence"—figuring out new rules on the fly.
  • AIs are better at "crystallized intelligence"—using the massive amount of data they've already memorized.

The study shows that while AI is getting scary good at talking, it still hasn't quite mastered the art of thinking the way a human does, especially when the rules of the game are hidden or unusual. The AI is still mostly guessing based on what it has seen before, rather than truly understanding the logic behind the puzzle.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →