← Latest papers
💬 NLP

WorldTravel: A Realistic Multimodal Travel-Planning Benchmark with Tightly Coupled Constraints

The paper introduces **WorldTravel**, a new multimodal benchmark consisting of complex, real-world travel scenarios designed to expose the significant performance gap in current frontier models when navigating tightly coupled temporal and logical constraints through both text and visual web environments.

Original authors: Zexuan Wang, Chenghao Yang, Yingqi Que, Zhenzhu Yang, Huaqing Yuan, Yiwen Wang, Zhengxuan Jiang, Shengjie Fang, Zhenhe Wu, Zhaohui Wang, Zhixin Yao, Jiashuo Liu, Jincheng Ren, Yuzhen Li, Yang Yang, Ji
Published 2026-02-10
📖 4 min read☕ Coffee break read

Original authors: Zexuan Wang, Chenghao Yang, Yingqi Que, Zhenzhu Yang, Huaqing Yuan, Yiwen Wang, Zhengxuan Jiang, Shengjie Fang, Zhenhe Wu, Zhaohui Wang, Zhixin Yao, Jiashuo Liu, Jincheng Ren, Yuzhen Li, Yang Yang, Jiaheng Liu, Jian Yang, Zaiyuan Wang, Ge Zhang, Zhoufutu Wen, Wenhao Huang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to plan a dream vacation to Europe. You want to see the Louvre, eat at a Michelin-star restaurant, and stay in a boutique hotel. But there’s a catch: the museum only has tickets for 10:00 AM, the restaurant is fully booked except for a 1:00 PM slot, and if you spend too much time at the museum, you’ll miss your train. One tiny mistake—a late bus or a forgotten reservation—and your entire week of plans collapses like a house of cards.

This is exactly what the researchers behind WorldTravel are testing.

The Problem: The "House of Cards" Effect

Most AI assistants today are like helpful but slightly scatterbrained travel agents. If you ask them to "find a hotel in Paris," they do a great job. But if you ask them to "coordinate a complex 5-day trip where every single activity is tightly locked to a specific time and a specific budget," they tend to stumble.

Current AI benchmarks (the "tests" we use to grade AI) are often too easy. They are like math tests where every question is independent. If you get question #1 wrong, it doesn't affect question #2. But real travel is a "tightly coupled" puzzle. In travel, every decision is a domino; if the first domino (your museum entry) falls the wrong way, it knocks over the rest of your trip.

The Solution: WorldTravel & Webscape

The researchers created a new, much harder "final exam" for AI called WorldTravel.

Instead of giving the AI a neat, organized spreadsheet of information, they built a digital playground called Webscape. In this playground, the AI can't just read a text file; it has to "look" at simulated websites—just like a human does. It has to navigate booking pages, look at restaurant menus, check maps, and—crucially—figure out if a time slot is "grayed out" (meaning it's sold out) just by looking at the picture.

The "Perception-Action Gap": The AI's Blind Spot

The study tested the world's smartest AI models (like GPT-5.2) and found something shocking: they are failing miserably.

The researchers discovered two major "walls" that the AI hits:

  1. The Perception-Action Gap (The "Eyes vs. Brain" Problem):
    Think of this like a person who is a genius at math but is wearing heavy, blurry goggles. The AI might be brilliant at reasoning (the brain), but it struggles to see the information on the screen (the eyes). When the AI had to look at a screenshot to find a price or a time, its performance plummeted. It could "think" the plan, but it couldn't "see" the data to fuel the thought.

  2. The 10-Constraint Wall (The "Mental Load" Problem):
    The researchers found that AI models have a "capacity limit." As long as the trip is simple (maybe 5 or 6 things to do), the AI is okay. But once you hit about 10 constraints (10 specific rules like "must be at the museum by 2 PM" or "must spend at least 3 hours there"), the AI's brain essentially "short-circuits." The complexity becomes too heavy, and the plan falls apart.

Why This Matters

This paper is a wake-up call. It tells us that if we want AI to actually manage our real lives—booking our flights, managing our calendars, or running our businesses—we can't just make them "smarter" at talking. We have to teach them to see the world clearly and reason through long, complex chains of events without dropping the ball.

In short: We've taught AI how to chat; now we need to teach it how to actually navigate the messy, interconnected reality of the real world.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →