← Latest papers
🧠 neuroscience

Humans use a dual policy to improve inferences during epistemic information seeking

Through three studies involving 702 participants, the authors demonstrate that humans improve epistemic inference by employing a distinct two-stage policy—initially "streaking" to test hypotheses before shifting to uncertainty-guided exploration—which outperforms standard neural network strategies and reveals unique psychological traits beyond traditional reward-based exploration.

Original authors: Cao, Y., Almeras, C., Lee, J. K., Maye, I., Wyart, V.

Published 2026-02-16
📖 5 min read🧠 Deep dive

Original authors: Cao, Y., Almeras, C., Lee, J. K., Maye, I., Wyart, V.

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

The Big Idea: How We Learn When There's No Prize

Imagine you are walking through a massive, unfamiliar city. You have two ways to explore:

  1. The "Hungry Tourist" (Reward-Seeking): You are starving. You want to find the best pizza place right now. You check a review, go there, and if it's good, you eat there again. You are focused on the immediate reward.
  2. The "Curious Explorer" (Epistemic/Information-Seeking): You aren't hungry. You just want to understand the city's layout. You don't know which street leads where, and there's no "best" street yet. You just want to build a mental map.

This paper asks: How do we explore when we aren't looking for a reward, but just for knowledge?

The researchers found that humans have a unique, two-step strategy for learning that computers (specifically standard AI) don't naturally figure out on their own.


The Experiment: The Magic Gem Bags

To test this, the researchers created a game with two "bags" of gems.

  • Bag A has mostly blue gems (but sometimes orange ones).
  • Bag B has mostly orange gems (but sometimes blue ones).
  • You can't see inside; you have to pull one gem out at a time to guess which bag is which.

They played two versions of the game:

  1. The "Match" Game (Reward): You are told, "Collect as many Blue gems as possible." Here, you are motivated by points. You quickly figure out which bag gives blue gems and stick to it.
  2. The "Guess" Game (Knowledge): You are told, "Just learn what's inside the bags. At the end, I'll ask you to guess which bag is which." There are no points for picking the right gem during the game. You are just gathering information.

The Discovery: The "Streaking" Strategy

When playing the Guess game (learning for knowledge), humans did something weird and fascinating that the "Hungry Tourist" didn't do.

The "Streaking" Analogy:
Imagine you are tasting two new soups to figure out which one is spicy and which is salty.

  • The Logical AI approach: Taste a spoonful of Soup A, then a spoonful of Soup B, then A, then B. Compare them side-by-side constantly.
  • The Human "Streaking" approach: You take five spoonfuls of Soup A in a row. Sip, sip, sip, sip, sip. Then you switch and take five spoonfuls of Soup B. Sip, sip, sip, sip, sip.

The researchers call this "Streaking." Even though switching back and forth seems like it would give you more data faster, humans naturally group their learning. They focus on one option, test it repeatedly to get a "feel" for it, and then switch to the other.

Why do we do this?
The paper suggests our brains are a bit "noisy" (like a radio with static). If we switch too quickly, the static makes it hard to hear the signal. By "streaking" (sticking with one option), we drown out the noise and get a clearer picture of that specific option before moving on. It's like tuning a radio to one station for a while to make sure you've got the signal clear before switching channels.

The Two-Stage Policy

So, human learning in the "Guess" game follows a two-stage dance:

  1. Stage 1 (The Streak): We pick an option and stick with it for a few turns to build a solid, clear hypothesis. "Okay, this bag is definitely mostly orange."
  2. Stage 2 (The Uncertainty Dance): Once we have a decent idea, we start switching back and forth to the option we know the least about, just to fill in the gaps in our map.

The Computer vs. The Human

The researchers trained Artificial Neural Networks (AI) to play the same game.

  • The AI's Success: The AI was great at Stage 2. It quickly learned to switch to the "most uncertain" option to learn the fastest. It was a perfect, logical learner.
  • The AI's Failure: The AI never figured out Stage 1 (Streaking). It didn't naturally group its choices. It just kept switching back and forth.

The Takeaway: Streaking isn't a "bug" in human thinking; it's a feature. It's a specific human trick to deal with our "noisy" brains. We need to focus on one thing at a time to make sense of it, whereas a perfect computer doesn't need that crutch.

Why Do Some People Streak More Than Others?

The study also looked at personality. They found two different types of people:

  1. The "Need for Closure" People: These are people who hate uncertainty and want answers fast. They tend to streak less. They switch quickly because they are anxious to finish the task.
  2. The "Smart Explorers": People with higher general reasoning skills tended to streak more effectively. They used the streaking strategy to build better mental maps, which led to higher accuracy in guessing the bag colors.

Summary in a Nutshell

  • Reward vs. Curiosity: When we want a reward, we hunt for the best option. When we want knowledge, we explore.
  • The Human Quirk: When exploring for knowledge, humans naturally "streak"—we focus on one thing repeatedly before switching.
  • The Benefit: This seems inefficient, but it actually helps us learn better because it helps our brains filter out "noise" and confusion.
  • The AI Gap: Computers can learn to be efficient, but they don't naturally invent the "streaking" strategy. It's a uniquely human adaptation to our own messy, noisy brains.

In short, humans are not just inefficient computers. We have a special, two-step way of learning that turns our "noisy" brains into a strength, allowing us to build better mental maps of the world than a purely logical machine might.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →