← Latest papers
💬 NLP

Can Large Language Models Infer Causal Relationships from Real-World Text?

This paper introduces ReCITE, the first benchmark derived from real-world academic literature to evaluate large language models' causal reasoning capabilities, revealing that current models struggle significantly with inferring causal relationships in complex, natural texts compared to their performance on synthetic datasets.

Original authors: Ryan Saklad, Aman Chadha, Oleg Pavlov, Raha Moraffah

Published 2026-04-14
📖 5 min read🧠 Deep dive

Original authors: Ryan Saklad, Aman Chadha, Oleg Pavlov, Raha Moraffah

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

🕵️‍♂️ The Big Question: Are AI Detectives or Just Parrots?

Imagine you hire a super-smart detective (a Large Language Model, or LLM) to solve a mystery. You give them a thick, messy police report (a real-world academic paper) and ask: "Who caused what?"

  • The Parrot: The detective just repeats exactly what the report says. If the report says "A caused B," the detective says "A caused B."
  • The Detective: The detective reads between the lines. The report might say, "It rained, and the grass got wet," without explicitly saying "Rain caused the grass to get wet." A real detective infers the cause.

The Problem: Most previous tests for AI were like giving the detective a cheat sheet where the answers were written in big, bold letters. This paper asks: Can these AI detectives solve the mystery when the clues are hidden, scattered, and written in complex, real-world language?

🏗️ The New Test: "ReCITE" (The Real-World Maze)

The researchers built a new test called ReCITE. Instead of using simple, made-up sentences, they gathered 292 real academic papers from fields like economics, engineering, and environmental science.

Think of these papers as dense, tangled forests.

  • The Trees: These are the "events" (e.g., "rain," "crop growth," "income").
  • The Roots: These are the "causes."
  • The Goal: The AI must draw a map (a causal graph) showing how the trees are connected by their roots.

Why is this hard?
In a real forest, you don't see a signpost saying "Rain causes Growth." You have to look at the wet soil, the green leaves, and the weather report to figure it out. Sometimes the clues are missing entirely, or the text is 100 pages long.

📉 The Results: The AI Got Lost

The researchers tested the world's smartest AIs (like Claude, GPT-4, Gemini, etc.) on this forest maze.

The Scorecard:

  • The Best AI: Got a score of 0.535 (on a scale where 1.0 is perfect).
  • The Average AI: Did even worse.

What does this mean?
It's like giving a GPS a map of a city it has never seen, but the streets are named in a code it doesn't fully understand. The AI can guess some streets, but it misses the main highways and often connects the wrong neighborhoods.

Key Findings:

  1. The "Explicit" Trap: When the text clearly says "A causes B," the AI does okay. But when the text implies it (e.g., "A happened, then B happened"), the AI's performance drops by half. It struggles to read between the lines.
  2. It's Not a Vocabulary Problem: The researchers gave the AI the list of all the "trees" (nodes) it needed to find. Even with the list, the AI still couldn't figure out how they were connected. This proves the problem isn't that the AI doesn't know the words; it's that it doesn't understand the logic of cause and effect.
  3. Size Doesn't Save You: Bigger, more powerful models did slightly better, but they still failed at the core task. Being "smarter" didn't make them better detectives.

🧩 The Analogy: The "Recipe" vs. The "Cooking Class"

  • Previous Tests (Synthetic Data): Imagine teaching a student to cook by giving them a recipe that says: "Step 1: Add salt. Step 2: Salt makes food salty." The student just memorizes the steps.
  • This Paper (ReCITE): Imagine giving the student a 50-page memoir of a grandmother describing her life in a village. The student has to figure out that "drought" caused "poor harvest," which caused "low income," which caused "migration." The text never explicitly states "Drought caused migration." The student has to infer the whole chain.

The Result: The AI students are great at memorizing the recipe, but they fail miserably at the cooking class. They can't connect the dots in the real world.

🔍 Why Does This Matter?

If we want AI to be truly helpful in the future (like diagnosing diseases, fixing climate change, or managing economies), it needs to understand cause and effect, not just word patterns.

  • Current AI: "I see the word 'fire' and the word 'smoke,' so I'll say fire causes smoke." (It's just matching patterns).
  • Future AI (What we need): "I understand that heat ignites fuel, which creates smoke. If I stop the heat, the smoke stops." (It understands the mechanism).

🏁 The Conclusion

This paper is a reality check. It shows that while Large Language Models are amazing at writing, summarizing, and chatting, they are currently bad at deep reasoning when faced with messy, real-world information.

They are like brilliant librarians who can find any book instantly, but they aren't yet philosophers who can understand the deep, hidden connections between the ideas in those books. To build "Artificial General Intelligence" (AI that thinks like a human), we need to teach them how to be better detectives, not just better parrots.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →