Entrenchment in Human Exploration: A Future-Path Model of Early Convergence
This paper introduces a cognitive model of "entrenchment" in human exploration, demonstrating that people prematurely settle on suboptimal options by simulating their own future choices, a mechanism that uniquely explains increased non-greedy behavior with longer horizons and distinguishes itself from standard exploration strategies like UCB or Thompson sampling.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
Human beings are constantly faced with a fundamental choice: should we stick with what we know works, or should we try something new to see if it is better? This dilemma, known to scientists as the exploration-exploitation tradeoff, is the engine of learning. When we explore, we gather information about the world; when we exploit, we use that information to get the best reward. Ideally, we balance these two drives, sampling enough to find the best option and then settling in to enjoy it. However, in many situations, people seem to stop searching too soon. They lock onto a choice that is not actually the best one and keep making it, even when better options are available. This stubborn persistence, which researchers call "entrenchment," suggests that our brains might be simulating the future in a way that traps us in suboptimal habits.
A new study by Shinji Nakazato at the Tokyo University of Science offers a fresh explanation for why this happens. The research proposes that when we evaluate our current choices, we do not just look at the immediate reward. Instead, we run a mental simulation of our own future decisions, imagining what we will do next. Crucially, the study suggests that we perform this simulation while holding our current beliefs fixed, as if we will never learn or update our opinions in the future. This mental shortcut creates a specific kind of uncertainty about our future path. The model shows that when the time horizon is long—meaning we have many decisions left to make—this simulated uncertainty grows large enough to make us abandon the best-known option and drift toward worse ones. The findings challenge the idea that humans are always trying to maximize their rewards, suggesting instead that our own internal simulations of the future can lead us astray.
To test this idea, the researcher analyzed data from two different types of experiments where people played games involving repeated choices. In one set of experiments, known as the Wilson task, participants faced a simple two-option game. The researchers carefully controlled the game so that the participants' knowledge about the options was identical at the start of a decision, but the number of future turns they had to play varied. In some cases, participants knew they had only one turn left; in others, they knew they had six turns left. If people were simply trying to find the best option, their behavior should have been the same in both cases because their knowledge was the same. However, the data showed a clear difference: when participants knew they had more turns remaining, they were significantly more likely to choose the option that was not the best one. This result was so strong that it could not be explained by standard theories of how people learn or by the idea that they were just being more careful or curious. The only model that could reproduce this specific pattern was the one based on the "future-path" simulation, where the decision-maker imagines their future choices without anticipating that they will learn anything new.
The study went further to distinguish this behavior from other forms of exploration. Sometimes, people choose a worse option on purpose to gather more information about the best possible choice; this is called directed exploration. To see if the stubborn choices in the study were a form of smart exploration or just getting stuck, the researcher looked at a second, more complex game played on a grid with many possible spots. In this game, the true best spot was known to the researchers. The analysis revealed that when people made non-optimal choices in the long-horizon situations, they were overwhelmingly moving toward spots that were genuinely worse, not toward the best spot. They were not exploring to find the optimum; they were drifting away from it. This confirmed that the behavior was indeed "entrenchment," a self-reinforcing loop where a person gets stuck on a suboptimal choice because their mental simulation of the future makes that choice seem more attractive than it really is.
The research also examined how the size of the choice set affects this behavior. In the grid game, participants had to choose between 30 or 121 different spots. The model predicted that having more options would increase the uncertainty in the mental simulation of the future, making it even more likely that people would get stuck on a bad choice. The data supported this: as the number of available options increased, the rate of sticking with a non-optimal choice also went up. Furthermore, the study found that when the difference in value between the best option and the second-best option was very clear, people were less likely to get stuck. But when the gap was small, or when the future felt far away, the tendency to persist with a mediocre choice grew stronger.
These findings paint a picture of human decision-making that is less about cold calculation and more about the limitations of our imagination. The study suggests that our brains are not always looking ahead to optimize our future; sometimes, they are looking ahead and seeing a fog of possibilities that makes the current, familiar path seem safer, even when it leads nowhere. By simulating our future selves as if we will never change our minds, we inadvertently create a barrier to learning. The research does not claim that this is a flaw in human intelligence, but rather a structural feature of how we process time and uncertainty. It shows that the very mechanism we use to plan our lives—imagining what we will do next—can sometimes be the thing that keeps us from finding the best path forward. The results were robust across thousands of trials and held true even when the researchers ruled out other explanations, such as simple mistakes or a general preference for variety.
Ultimately, the paper provides a mathematical framework that captures this specific type of getting stuck. It shows that the probability of making a non-optimal choice increases as the remaining time grows and as the number of options expands. This is the opposite of what traditional theories of rational learning predict, which suggest that having more time should help people find the best solution. The study confirms that for human learners, having more time can actually make us more likely to settle for less. The research leaves open the question of exactly how the brain calculates this future uncertainty, but it firmly establishes that the way we imagine our future choices is a powerful driver of our present behavior, capable of locking us into patterns that are not in our best interest.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.