← Latest papers
📊 statistics

A Sequential Tour-Based Mode Choice Framework with GPS-based Data: Integrating Machine-Learning-Derived Features into Mixed Logit and MNL Models

This study develops a sequential tour-based mode choice framework for Sydney using GPS data and machine-learning-derived features, demonstrating that Mixed Logit models significantly improve prediction for interdependent tour components and that the framework remains robust to upstream estimation uncertainty, thereby validating the use of simulated inputs for practical travel demand forecasting.

Original authors: Mostafa Rahimi, Maliheh Tabasi, Abdul Rawoof Pinjari, Taha Hossein Rashidi

Published 2026-07-30
📖 4 min read☕ Coffee break read

Original authors: Mostafa Rahimi, Maliheh Tabasi, Abdul Rawoof Pinjari, Taha Hossein Rashidi

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to predict the future of a city's traffic. For decades, scientists have treated every car trip like a single, isolated event, as if a person decides to drive to the grocery store without remembering they just drove to work. But in reality, our days are more like a string of beads on a necklace; the first trip sets the tone for the rest. This is the world of "tour-based" modeling, where a "tour" is a loop of activities starting and ending at home. To make these predictions, researchers use two main tools: traditional math models that are easy to understand but sometimes too simple, and powerful "machine learning" algorithms that are great at spotting patterns but act like a mysterious "black box" that no one can fully explain. The big question for city planners is: How do we build a model that is both smart enough to predict complex behavior and clear enough to explain why a person chose a bus over a car?

This paper is like a master chef trying to combine the best of both worlds. The researchers built a "sequential" system, which means they didn't just look at one trip; they modeled a person's entire day, step-by-step, from the moment they wake up to the moment they go back to sleep. They took the "black box" machine learning to find hidden clues about how people behave, and then fed those clues into a clear, understandable math model. But the real magic happens when they tested a tricky idea: usually, we teach these models using perfect, real-world data where we know exactly what happened yesterday. But in the real future, we won't know what happened yesterday! We'll have to guess. So, the team asked: "What if we teach the model by making it guess its own past?" They ran three different scenarios to see if teaching the model with its own "guesses" (simulated data) would make it worse, or if it would actually make it tougher and better at predicting the future.

Here is what they found, and it's a bit of a plot twist. First, they discovered that for most of the day, people are creatures of habit. Whether it's the first trip of the morning or a quick errand later in the day, the strongest predictor of how someone travels is simply what they usually do. The idea that people constantly switch modes based on what they did in the previous tour (like, "I took the bus to work, so I'll take the train to lunch") turned out to be mostly a myth in this data. Once you account for a person's usual habits, the "previous trip" doesn't add much new information.

Second, they found that the "smart" part of the model (called Mixed Logit) only really shines in specific situations. It made a huge difference—improving accuracy by about 5 percentage points—when predicting the start of a new tour (like a second trip out of the house). This is when people have the most freedom to choose. However, for the first trip of the day or the trips inside a tour, the simpler model worked just as well. Why? Because inside a tour, if you drive a car, you have to drive the car for the rest of that loop (you can't leave the car behind). That rule is so strict that no amount of "smart" math is needed to predict it.

Finally, the big surprise regarding the "guessing" test: teaching the model with its own simulated guesses didn't break it; it barely made a difference at all. In fact, when they looked at the accuracy of predicting a whole day of travel, the model trained on "guesses" was actually slightly better (67% accurate) than the one trained on perfect real-world data (63% accurate). It's as if practicing with a slightly imperfect map made the driver better at navigating the real world. The paper suggests that by training the model to handle the uncertainty of its own predictions, it becomes more robust and ready for the real world, proving that we don't need perfect data to build a reliable travel forecast.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →