TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning
This paper introduces TREK, a rigorous and fully reproducible benchmark featuring a deterministic evaluator and a large-scale synthetic travel dataset to assess LLM agents' ability to generate feasible, constraint-compliant itineraries, revealing that even state-of-the-art models struggle significantly with multi-constraint planning and satisfying unstated user needs.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a super-smart robot how to plan a vacation. In the world of artificial intelligence, these robots are called "LLM agents." They are like digital travel agents who can read your messy text messages, understand what you want, and then go out into the internet to book flights, find hotels, and schedule sightseeing. But here's the catch: just because a robot can talk about a perfect trip doesn't mean it can actually build one. A real vacation plan is a delicate puzzle. Every flight must actually exist, the hotels must be open, the days must be long enough to drive between cities, and the total cost must fit your wallet. If the robot makes up a flight number or forgets that a museum closes at 5 PM, the whole plan falls apart. For a long time, we've been testing these robots by asking them to write stories about trips, but we haven't had a strict way to check if their plans would actually work in the real world.
This is where a new study called TREK (Travel Reasoning and Evaluation Kit) comes in. Think of TREK as a massive, high-stakes "flight simulator" for travel-planning robots. Instead of letting the robots guess, the researchers built a perfectly consistent, fake world with 212,530 records of flights, hotels, and attractions across 375 cities. They then gave 15 different AI agents a set of 800 travel challenges. Some challenges were possible, and some were impossible (like trying to fly to a city that has no airport). The robots had to use a special set of tools to build a plan, and then a strict, computer-based referee (with no human or AI "judge" to be lenient) checked every single detail. The goal was to see if the robots could produce a plan that was not just a nice story, but a perfectly executable, error-free itinerary that respected the traveler's hidden needs, like being wheelchair-accessible or pet-friendly.
The results were a bit of a reality check. Even the strongest model in this specific 2027 study, GPT-5.6, managed to create a fully perfect, usable plan for only 46.2% of the possible trips. The average robot did much worse, with a median success rate of just 6.6%, and some failed completely, scoring 0.0%. The study found that the robots were actually quite good at the "boring" stuff: they rarely made up fake flights (hallucinations) and could usually tell when a trip was impossible. However, they hit a massive wall when it came to understanding the person traveling. The biggest bottleneck was satisfying "implicit needs"—those unspoken desires like "I'm a foodie and want to stay near a Michelin-star restaurant" or "I'm a disabled traveler and need a hotel with a ramp." Even the best robots failed to meet these hidden needs more than half the time.
The researchers also tested if giving the robots more time to "think" or letting them use more computer power would help. Surprisingly, it didn't. In the one pair of models they could compare, the one that spent more time "reasoning" actually did worse than the one that just followed instructions. Furthermore, spending more money on computer tokens didn't guarantee a better plan; some of the weakest robots were the most expensive, while the best one was surprisingly efficient. The study concludes that while AI agents are getting better at following rules, they still struggle to be truly reliable travel planners because they can't yet perfectly balance all the complex, human-like constraints of a real trip at once. The gap between a robot that talks about a trip and one that can book a perfect one is still wide open.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.