Fine-tuning LLMs for Tourist Trajectory Prediction using Field Experiment Data
This paper demonstrates that fine-tuning Large Language Models on field experiment data from Wakayama Castle Park enables accurate, context-aware prediction of tourist trajectories, including in undersampled scenarios like rainy days, thereby establishing a robust foundation for evaluating mobility interventions through counterfactual analysis.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine standing at the edge of a vast, bustling park, watching hundreds of people weave through paths, pause at gardens, and wander toward castles. For city planners and tourism officials, understanding why these visitors choose one path over another is not just a matter of curiosity; it is a critical tool for managing crowds, protecting fragile sites, and ensuring that money flows to the right places. When a destination tries to introduce a new electric shuttle or a bike-sharing system to help people move between scattered attractions, they face a difficult question: will this new service actually change how people explore, or will visitors simply use it to reach the same crowded spots faster? Traditional computer models struggle to answer this because they often treat human movement as a simple series of steps, ignoring the messy, real-world factors that influence our choices, such as a sudden rainstorm, a tired child, or the simple desire to find a quiet bench. These older systems require massive amounts of historical data to learn patterns, and when faced with a new situation they have never seen before, they often fail to guess correctly.
A team of researchers from the University of Osaka and the RIKEN Center for Computational Science has taken a different approach by turning to a type of artificial intelligence known as a large language model. These are the same powerful computer systems that can write stories, answer complex questions, and hold conversations, trained on vast libraries of human text. The researchers realized that these models already possess a deep, intuitive understanding of human behavior. They "know" that people seek shelter when it rains, that families need frequent breaks, and that lunchtime usually means a search for food. Instead of building a new model from scratch, the team decided to teach this existing, knowledgeable system how to navigate a specific location. They focused their study on Wakayama Castle Park in Japan, a sprawling destination with a historic castle tower, gardens, a zoo, and numerous rest areas. During a pilot program in December 2023, the researchers collected detailed movement records from 566 tourists, tracking where they went, when they arrived, and what they did, while also noting the weather and the time of day.
To teach the computer, the researchers translated these raw movement records into a format the language model could understand: a simple, structured story. They described each visitor as a character with specific traits, such as age or group type, and set the scene with the current weather. Then, they listed the visitor's journey step-by-step, noting the time, the action, and the specific place visited, much like writing a diary entry for a day in the park. They then asked the computer to read these stories and learn the patterns of how people actually moved through the park. The goal was to see if the model could look at a visitor's history up to a certain point and accurately predict where they would go next. This process, known as fine-tuning, allowed the model to combine its general knowledge of human nature with the specific layout and habits of Wakayama Castle Park.
The results were striking. When tested on new visitors the model had never seen before, it correctly predicted the next specific location a tourist would visit about 49 percent of the time. This was a significant improvement over traditional statistical methods, which managed to guess correctly only about 9 to 15 percent of the time. The older methods often failed because they could not account for context; they did not understand that a rainy day would make a visitor skip the open gardens and head straight for the indoor museum. The language model, however, used its built-in common sense to make these adjustments. Even in difficult situations where the data was scarce, such as on rainy days when only a handful of visitors were recorded, the model maintained a high level of accuracy, correctly predicting that people would seek indoor shelter. It also successfully generated entire day-long itineraries that felt realistic, starting with morning visits to major landmarks, moving to lunch, and ending with afternoon rest, all without being explicitly programmed with these rules.
The researchers found that the model's ability to understand the "why" behind a decision was just as important as its ability to guess the "where." While it was slightly less accurate at predicting the exact name of a specific shop or garden, it was very good at predicting the type of place a visitor would choose, such as a restaurant or a historic site. This suggests that the model had learned the general intentions of the tourists. Interestingly, the team also tested a version of the model that had been specifically trained on Japanese language and culture, but the general, multilingual version performed better. This indicates that for understanding how people move and make choices, the broad, diverse knowledge of human behavior found in a general model is more valuable than specialized language training.
This work does not claim to have solved the problem of predicting human behavior perfectly, nor does it prove that these models can predict exactly how a new policy will change the future. The researchers are careful to note that their model is a powerful tool for forecasting what people will do under current conditions, but using it to simulate the effects of a new intervention, like a new shuttle bus, requires further testing. However, the study demonstrates that these large language models can serve as highly accurate mirrors of human behavior, capable of understanding the subtle, context-dependent decisions that define a day at the park. By capturing the logic of why a tired parent chooses a bench over a long walk, or why a group avoids a hill on a hot day, this technology offers a new way for destination managers to plan for the future, moving beyond simple statistics to a deeper understanding of the people who visit their lands.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.