← Latest papers
📈 economics

Dynamic Decision-Making under Model Misspecification: A Stochastic Stability Approach

This paper analyzes the performance of Thompson Sampling under model misspecification by classifying posterior evolution into distinct regimes within a two-armed Gaussian bandit and establishing a unified stochastic stability framework for general finite model classes to characterize limiting beliefs and asymptotic regret.

Original authors: Xinyu Dai, Daniel Chen, Yian Qian

Published 2026-07-21
📖 5 min read🧠 Deep dive

Original authors: Xinyu Dai, Daniel Chen, Yian Qian

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot how to navigate a maze, but you've given it a map that is slightly wrong. Maybe the map says a wall is made of glass when it's actually brick, or it thinks a shortcut leads to the exit when it actually leads to a dead end. This is the world of "misspecified learning." In science and economics, we often assume that if we give a smart system enough data, it will eventually figure out the truth and stop making mistakes. This idea relies on the concept of "Bayesian learning," where a system updates its beliefs based on new evidence, like a detective gathering clues to solve a case. Usually, we expect that with enough clues, the detective will point to the one true suspect. But what happens if the detective is using a flawed theory about how the world works? Does the robot eventually learn the right path, or does it get stuck in a loop of confusion? This question matters because today, everything from online shopping algorithms to government policy decisions relies on these learning systems. If they get stuck in a loop, the consequences could be expensive prices, bad recommendations, or ineffective laws.

This paper, written by researchers Xinyu Dai, Daniel Chen, and Yian Qian, dives deep into what happens when a learning system uses a specific, popular strategy called "Thompson Sampling" while its internal map is wrong. Thompson Sampling is a clever way for a robot to learn: instead of just picking the option it thinks is best right now, it occasionally tries a different option just to see what happens. It's like a chef who usually cooks their favorite dish but occasionally tries a new recipe just to keep their skills sharp. The authors wanted to know: if the chef's recipe book is full of errors, does this "tasting" behavior help them eventually find the truth, or does it trap them in a weird, endless cycle?

The researchers found that the answer depends entirely on how the wrong recipes interact with the real ingredients. They discovered three main scenarios. First, there is the "Self-Confirming" trap. Imagine the robot believes a high price is best, and it keeps charging high prices. If the real world happens to look good at high prices (even for the wrong reason), the robot gets confident and never changes its mind. It locks into a single strategy forever, which might be the right one, or it might be a permanent mistake. Second, there is the "Uniform Dominance" scenario, where one of the robot's wrong models is just clearly better than the others at explaining everything. In this case, the robot eventually figures out which model is the "least wrong" and sticks with it, converging to a stable decision.

But the most surprising discovery is the third scenario: the "Self-Defeating" loop. This happens when the robot's wrong models are so tricky that every time it tries to prove one is right, the results actually prove it wrong. For example, if the robot thinks a high price is best, it charges high prices. But the data from those high prices makes the robot think, "Wait, maybe a low price is better!" So it switches to a low price. But then the data from the low price makes it think, "No, high price was better!" The robot ends up oscillating back and forth forever. The authors show that in this specific setup, the robot never settles down. Even with infinite amounts of data, it never stops guessing. Instead of finding a single answer, its beliefs settle into a permanent, rhythmic dance of uncertainty.

The paper proves mathematically that this "Self-Defeating" behavior isn't just a glitch; it's a stable state where the system keeps exploring forever. This is a big deal because it challenges the old idea that "more data always leads to certainty." The authors show that if the learning algorithm is designed to keep experimenting (like Thompson Sampling does), and the world is misunderstood in a specific way, the system will never stop fluctuating. It will keep changing its mind, leading to constant changes in behavior—like a company that keeps changing its prices up and down forever, not because the market is changing, but because its learning algorithm is stuck in a loop of self-doubt. The researchers also extended this idea to situations with many more models, showing that while these loops can happen, they often get "pruned" down to simpler loops or single choices as the system gets more complex, unless the conditions are just right to keep the chaos alive.

In short, this paper tells us that being "smart" and "curious" isn't always enough to find the truth. If your starting assumptions are wrong in a specific way, your curiosity can actually prevent you from ever settling on a decision. The robot might never stop trying new things, not because it's learning, but because the very act of learning keeps pushing it away from the answer. This suggests that for systems making real-world decisions, we need to be careful about how we design their learning rules, because sometimes, the best way to learn might be to stop guessing and start trusting a simpler, more stable approach.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →