← Latest papers
💬 NLP

Human Adults and LLMs as Scientists: Who Benefits from Active Exploration?

This paper demonstrates that granting human adults agency through active exploration significantly improves their ability to identify conjunctive causal rules compared to passive observation, while revealing that although large language models can match human inference accuracy, they often employ less efficient exploration strategies and exhibit similar performance gaps between conjunctive and disjunctive reasoning.

Original authors: Mandana Samiei, Eunice Yiu, Anthony GX-Chen, Dongyan Lin, Jocelyn Shen, Blake A. Richards, Alison Gopnik, Doina Precup

Published 2026-06-05
📖 5 min read🧠 Deep dive

Original authors: Mandana Samiei, Eunice Yiu, Anthony GX-Chen, Dongyan Lin, Jocelyn Shen, Blake A. Richards, Alison Gopnik, Doina Precup

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Question: Are Adults Bad at "Teamwork" Logic?

Imagine you are trying to figure out how a mysterious machine works. You have four different keys (let's call them "Nexioms"). You need to find out which keys turn the machine on and how they work together.

There are two ways the machine could work:

  1. The "Or" Rule (Disjunctive): The machine turns on if you put in any one of the correct keys. (Like a light switch: flip one, and the light goes on).
  2. The "And" Rule (Conjunctive): The machine only turns on if you put in all the correct keys at the same time. (Like a safe: you need the code, the fingerprint, and the key all together).

The Old Finding:
Scientists used to think adults were bad at figuring out the "And" rule. When people just watched someone else test the keys, adults usually guessed the "Or" rule, even when the evidence suggested an "And" rule. It was like adults had a mental blind spot for teamwork logic.

The New Question:
The authors of this paper asked: Is this because adults are actually bad at logic, or is it because they were just sitting passively watching?

They wondered if adults would do better if they could act like scientists—picking up the keys themselves, testing them, and seeing what happened immediately.


The Experiment: The "Nexiom Detector"

The researchers built a digital game called the "Nexiom Detector."

  • The Active Group: These participants could click to add or remove keys and press "Test" to see if the machine turned on. They controlled the experiment.
  • The Passive Group: These participants just watched the Active Group's screen. They saw the same results but couldn't touch anything.
  • The "Planner" Group: A third group could plan which keys to test, but they didn't see the results of their own plans. Instead, they watched the results of someone else's tests.

They also tested several Large Language Models (LLMs) (AI chatbots) to see if they could solve the puzzle as well as humans.


What They Found

1. Giving Adults a "Remote Control" Changed Everything

When adults were allowed to actively test the keys themselves, their performance skyrocketed.

  • The Result: They became very good at figuring out the tricky "And" rules. They successfully identified which keys were needed and that all of them had to be present.
  • The Takeaway: Adults aren't bad at "And" logic. They just struggle when they can't design their own tests. When they act like scientists, they are just as smart as children at figuring out complex rules.

2. The "Planner" vs. The "Doer"

Here is a crucial finding: Just thinking about what to test isn't enough.

  • The "Planner" group (who thought about tests but watched someone else's results) did poorly. They performed almost as badly as the Passive group.
  • The Analogy: Imagine you are a chef. If you write down a recipe (planning) but someone else cooks the dish and you only watch them eat it, you won't learn how to cook. You need to taste the food you cooked to know if the salt was right.
  • The Lesson: Active learning works best when your action is tightly linked to the result. You need to see the immediate consequence of your own choices.

3. The "And" Rule is Still Harder (But Doable)

Even with active control, the "And" rule took more time and more tests to figure out than the "Or" rule.

  • The Analogy: Finding a single key that opens a door is easy (one try might work). Finding the exact combination of three keys that opens a safe takes more trial and error.
  • However, the adults didn't give up. They just ran more tests to be sure. They didn't get stuck; they just worked a little harder.

4. How Did the AI (LLMs) Do?

The researchers pitted the humans against advanced AI models.

  • The "Or" Rule: The AI was great. Some models performed just as well as the top human scientists.
  • The "And" Rule: The AI struggled more than the humans. While some smart models got close, many of them were less efficient. They often took more tests than necessary or got confused by the "And" logic.
  • The Takeaway: AI can be a good scientist for simple rules, but humans are still better at the strategic, step-by-step detective work required for complex "teamwork" rules.

Summary: What Does This Mean?

This paper tells us that adults are not "broken" when it comes to complex logic. The problem wasn't their brains; it was the situation.

  • Passive Observation: Like watching a magic show from the audience. You see the trick, but you don't understand how it's done because you can't touch the props.
  • Active Exploration: Like being the magician's assistant. You get to hold the props, try the tricks, and see what happens when you pull the wrong lever.

The Main Conclusion:
When we let adults act like scientists—giving them the freedom to test their own ideas and see the immediate results—they are excellent at figuring out complex "And" rules. The "handicap" we thought they had was actually just a lack of agency.

However, even with this freedom, figuring out complex rules takes more effort than simple ones, and while AI is getting good at this, humans still have a slight edge in being efficient, strategic explorers.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →