← Latest papers
📄 animal behavior and cognition

Transition from model-free to structure-informed decision making in dynamic environments

This study demonstrates that mice shift from reactive model-free to deliberative structure-informed decision-making strategies during the early stages of training in a dynamic environment, challenging the traditional view that such transitions occur only after extensive practice.

Original authors: Yasueda, M., Taira, M., Akam, T., Walton, M. E., Doya, K.

Published 2026-08-18
📖 4 min read☕ Coffee break read

Original authors: Yasueda, M., Taira, M., Akam, T., Walton, M. E., Doya, K.

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

In the vast landscape of how living things make choices, scientists have long distinguished between two fundamental ways of navigating the world. One approach is immediate and reactive: an animal learns that a specific action leads to a reward and repeats that action without necessarily understanding why it works. This is a strategy of habit, built on simple associations between doing something and getting something good. The other approach is more thoughtful and deliberate: an animal builds an internal map of how the world is structured, understanding how one event leads to another and how different choices change the odds of a reward. This is a strategy of planning, where the decision-maker uses knowledge of the rules to guide behavior. For years, researchers believed that as animals practiced a task over a long time, they would slowly shift from this thoughtful planning to the simpler, automatic habits. This idea suggested that with enough repetition, the brain would stop calculating and start relying on muscle memory.

A team of researchers decided to test this assumption by watching how mice learn a complex decision-making task, but with a crucial twist: they looked at the very beginning of the learning process rather than waiting for the animals to become experts. The mice were placed in a situation where they had to make a series of choices to find a reward. The environment was tricky; the path to the reward was not always straightforward. Sometimes, a choice would lead directly to the next step as expected, but other times, the path would take an unexpected turn. Furthermore, the likelihood of finding a reward at the end of the path was not fixed; it changed over time, forcing the mice to constantly update their expectations. The researchers observed the mice's behavior closely as they went through their training sessions, tracking whether the animals were simply reacting to past rewards or if they were beginning to understand the hidden structure of the task.

The results revealed a surprising pattern that challenges the old view of how learning unfolds. As the mice practiced, their behavior changed in a specific way that indicated they were not just forming habits. When the mice encountered a situation where the path took an unexpected turn, they adjusted their future choices differently than when the path went as expected. This divergence in behavior is a clear sign that the animals had started to use their knowledge of the task's structure to guide their decisions. By fitting mathematical models of different learning styles to the mice's actions, the researchers found that the mice were increasingly relying on a strategy that used this structural knowledge. This shift happened early in the training process, long before the mice had spent extensive time on the task.

This finding suggests that the transition from simple, reactive learning to complex, structure-informed decision-making happens much sooner than previously thought. While earlier studies focused on what happens after an animal has practiced for a long time, often concluding that the brain eventually stops thinking and starts acting on autopilot, this new work shows that the brain begins to build a mental model of the rules almost immediately. The mice did not wait to become habitual before they started to understand the game; instead, they quickly moved away from relying solely on past rewards and began to use their understanding of how the task was put together to make better choices. This demonstrates that in dynamic environments where rules and rewards can change, the ability to understand the underlying structure is a primary tool for survival, one that the brain reaches for early and often.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →