A Bayesian latent class reinforcement learning framework to capture adaptive, feedback-driven travel behaviour
This paper proposes a Bayesian Latent Class Reinforcement Learning framework to model heterogeneous, adaptive travel behaviors, identifying three distinct traveler classes with unique preference evolution and exploration-exploitation strategies through analysis of driving simulator data.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to learn the best way to get to school. Sometimes you take the same route every day because it's fast and reliable. Other times, you try a new street just to see if it's faster, even if it might get you there later. This is the core of how humans make decisions: we balance sticking to what we know (exploitation) with trying new things to learn more (exploration). For a long time, scientists studying how people travel have used models that assume everyone learns the same way or that our preferences are fixed like stone. But in real life, people are messy. We learn at different speeds, we get confused by new information, and we might act totally differently when we are actually driving versus when we are just imagining a trip. This paper dives into that messy reality, using a mix of psychology and math to figure out not just what people choose, but how they learn to make those choices over time.
The researchers behind this study, led by Georges Sfeir and colleagues, decided to build a new kind of "travel brain" model called a Latent Class Reinforcement Learning (LCRL) framework. Think of it as a detective story where the detectives are trying to figure out why different drivers act so differently. Instead of assuming everyone is the same, they used a driving simulator to watch 83 people make 20 route choices each. Half the time, the participants actually drove the routes and felt the time pass (experienced feedback); the other half, they just clicked buttons on a screen with hypothetical times (stated preference). The team then fed this data into their fancy new model, which acts like a sorting machine. It doesn't just say "this person likes fast routes"; it asks, "Is this person a slow learner who sticks to old habits? A risk-taker who tries everything? Or someone who changes their mind depending on whether they are actually driving or just thinking about it?"
The model found that the group of drivers wasn't a single blob of people, but rather three distinct "types" of learners, each with their own personality. About 22% of the people were the "Context Chameleons." These folks acted one way when they were just imagining a trip on a computer screen (preferring the safe, reliable route) but completely flipped their behavior when they were actually behind the wheel in the simulator, suddenly craving the risky, unpredictable route. They learned slowly, taking their time to update their beliefs based on new experiences.
Then there was the largest group, making up nearly half the participants (49%), who were the "Persistent Exploiters." These drivers were the ones who loved the risky, unreliable route no matter what. Whether they were driving or just clicking buttons, they kept choosing the uncertain path, even when it sometimes took them longer. They were the most stubborn exploiters, sticking to their risky choice and adapting somewhat faster to feedback than the slowest group, though they still relied considerably on past experiences.
Finally, there was the third group, about 29% of the sample, who were the "Adaptive Switchers." These drivers were more likely to be female and have higher incomes. They started out liking the safe route when they were actually driving, but when they were just answering questions on a screen, they suddenly became more adventurous and picked the risky route. They displayed the most exploratory tendencies, switching back and forth between safe and risky options. While they had the highest learning rate estimate of the three groups, it was statistically insignificant, suggesting they maintained stable expectations that evolved gradually over time rather than reacting with rapid, volatile shifts.
The paper suggests that by ignoring these different "learning personalities," we might be making mistakes in how we design traffic systems or predict how people will react to new roads. If you treat a "Context Chameleon" like a "Persistent Exploiter," your traffic advice might fail completely. The researchers didn't just guess these groups existed; they used a sophisticated math technique called Variational Bayes to prove that splitting people into these three classes fit the data much better than assuming everyone learns the same way. They even ran simulations to show how each group would react if they kept getting the same bad traffic or the same good traffic over and over, confirming that their "personalities" really did change how they learned.
In short, this paper argues that to understand travel, we can't just look at the map; we have to look at the driver's mind. We need to know if they are the type to stick to the familiar, the type to chase the new, or the type who changes their mind depending on the situation. By building a model that captures these differences, the authors hope to help cities and planners design better systems that work for all types of learners, not just the average one.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.