← Latest papers
🤖 AI

The Agentic Garden of Forking Paths

This paper demonstrates that AI agents can replicate the diverse and often conflicting analytical choices made by human researchers across high-stakes domains, highlighting the problem of selective reporting in scientific analysis and proposing the "Agentic Bootstrap" method to estimate the "m-value" as a new criterion for evaluating the credibility of scientific claims based on their position within the space of plausible analyses.

Original authors: Jiacheng Miao, Jonathan K Pritchard, James Zou

Published 2026-07-03
📖 5 min read🧠 Deep dive

Original authors: Jiacheng Miao, Jonathan K Pritchard, James Zou

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a mystery using a single box of clues (the data). In the past, if two detectives looked at the same box, they might come up with slightly different stories, but usually, they would agree on the main facts.

This paper introduces a new twist: AI agents acting as detectives. The researchers found that if you give two AI agents the exact same box of clues but tell them to have different "personalities" (one is a skeptic, one is a believer), they will not just tell slightly different stories—they will tell completely opposite stories, and both stories will sound perfectly logical and scientifically sound.

Here is a breakdown of their findings using simple analogies:

1. The "Garden of Forking Paths"

Imagine a massive forest with millions of paths (the "Garden of Forking Paths"). Every time a researcher analyzes data, they have to choose which path to walk down.

  • The Old Problem: Human researchers sometimes subconsciously choose the path that leads to the result they want to find. This is called "researcher degrees of freedom."
  • The New AI Problem: AI agents can now run through this forest incredibly fast. They can try thousands of different paths in minutes. The paper shows that if you tell an AI, "You believe immigration hurts the economy," it will naturally wander down the paths that lead to that conclusion. If you tell another AI, "You believe immigration helps the economy," it will wander down the opposite paths.
  • The Scary Part: Both AIs are doing "good" science. They aren't lying or breaking rules. They are just taking different, valid roads through the forest to reach their destination.

2. The Experiment: The "Ideological Mirror"

The researchers tested this on four big questions:

  • Does immigration affect welfare support?
  • Does coffee affect health?
  • Does social media hurt teen mental health?
  • Does gut bacteria affect weight?

They gave AI agents different "personas" (like giving them a specific political or scientific bias) and asked them to analyze the same data.

  • The Result: The "Pro-Immigration" AI agents consistently found positive effects, while the "Anti-Immigration" agents found negative effects.
  • The Human Comparison: In a famous real-world study, 42 human teams analyzed the same immigration data and had a "gap" in their conclusions. The AI agents reproduced 72% of that human gap.
  • The Speed: An AI agent could do this entire analysis in about 15 minutes for roughly $1.68.

3. The "Quality Check" Trap

You might think, "Well, the biased AI must be doing bad math."

  • The Surprise: The researchers had both other AIs and human experts (PhD statisticians) review the reports.
  • The Verdict: 86% of the AI reports passed the review. 78% passed the human expert review.
  • The Takeaway: The AI agents were reaching opposite conclusions using methods that experts considered flaw-free. The problem wasn't that the math was wrong; the problem was that the AI (like a human) selectively explored the forest to find the path that matched its "belief."

4. How the AI Does It: Exploration vs. Selection

The paper breaks down how the AI gets to the wrong (or right) conclusion:

  1. Exploration: The AI starts walking. If it's a "believer," it naturally wanders toward paths that look promising for its belief. If it's a "skeptic," it wanders elsewhere.
  2. Selection: After walking many paths, the AI picks the one that best fits its story to put in the final report.
  • Analogy: Imagine a chef who wants to make a spicy dish. They taste 100 different soups. They naturally pay more attention to the spicy ones, tweak them to be spicier, and finally serve the spiciest one. A "non-spicy" chef would do the exact same process but end up serving a bland soup. Both chefs followed the rules of cooking perfectly.

5. The Solution: The "m-value" and "Agentic Bootstrap"

Since we can't trust a single report anymore (because the AI could have easily picked a different path), the authors propose a new way to judge science.

  • The Old Way (p-value): "If we repeated this experiment 1,000 times with the same recipe, how often would we get this result?" (This checks for luck in the data).
  • The New Way (m-value): "If we tried 1,000 different valid recipes on this same data, how often would we get this result?" (This checks for luck in the choices made).

Agentic Bootstrap is the tool they built to do this. It uses AI agents to simulate thousands of different "what-if" scenarios (different paths through the forest) and creates a map of all possible valid conclusions.

  • If a researcher's result is in the middle of the map, it's robust.
  • If a researcher's result is in the extreme edge of the map (the "tail"), it means they likely picked a very specific path just to get that result.

The Finding: When they applied this to the human immigration study, they found that 13.5% of the human reports were in the most extreme 5% of possible outcomes. This suggests that many human findings are "extreme" not because the data is so strong, but because the researchers (or their AI helpers) selectively picked the path that led there.

Summary

The paper argues that AI makes it cheap and easy to "cherry-pick" the best-looking scientific story from a forest of millions of valid options. The solution isn't to stop using AI, but to use AI to map the whole forest first. Before we trust a scientific claim, we should ask: "Is this result a typical path through the forest, or did someone just pick the one path that led to the conclusion they wanted?" The m-value is the new ruler used to measure that.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →