LHAW: Controllable Underspecification for Long-Horizon Tasks
The paper introduces LHAW, a modular synthetic pipeline that systematically generates controllable underspecified long-horizon task variants by removing information across four dimensions and validating their impact through empirical agent trials, thereby providing the first framework for cost-sensitive evaluation of agent clarification behavior in ambiguous settings.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you hire a very smart, super-fast robot assistant to run your entire business for you. You tell it, "Go make a marketing campaign for our new product."
In a perfect world, the robot knows exactly what you mean. But in the real world, your instructions are often missing tiny, crucial details. Maybe you didn't say which product, or what budget to use, or who the audience is.
If the robot is too confident, it might guess wrong and waste weeks of work. If it's too cautious, it might stop and ask you a million questions, annoying you and slowing everything down.
The Problem:
Currently, we don't have a good way to test if our AI robots know when to guess and when to ask for help. Most tests only give them perfect instructions, so they never learn to handle the messy, incomplete reality of real life.
The Solution: LHAW (The "Missing Piece" Simulator)
The authors of this paper created a tool called LHAW (Long-Horizon Augmented Workflows). Think of LHAW as a "Game Master" for AI robots.
Here is how it works, using a simple analogy:
1. The Recipe Test (Creating the Puzzle)
Imagine you have a perfect recipe for a cake (a well-specified task). LHAW takes that recipe and systematically removes ingredients or instructions to create a "broken" version.
- Goal Removal: "Make a cake." (But what kind? Chocolate? Carrot? For a birthday?)
- Constraint Removal: "Bake it." (But at what temperature? For how long?)
- Input Removal: "Use the flour." (But which bag of flour? The one in the pantry or the one in the garage?)
- Context Removal: "Decorate it." (But with what? Sprinkles? Icing? Is it for a baby shower or a funeral?)
LHAW creates thousands of these "broken recipes" and labels them:
- Outcome-Critical: If the robot guesses wrong, the cake is ruined (the task fails).
- Divergent: The robot might make a great cake or a terrible one depending on its guess.
- Benign: The robot can figure it out on its own (e.g., "Bake at 350 degrees" is a safe default).
2. The Stress Test (Watching the Robot)
Now, LHAW gives these broken recipes to different AI models (like GPT-5, Claude, Gemini) and watches what they do.
- The "Silent Failers": Some robots just guess. They make a cake, but it's the wrong flavor. They fail silently.
- The "Over-Askers": Some robots stop and ask, "What flavor? What size? What pan? What oven?" They ask so many questions they annoy the user.
- The "Strategic Askers": The best robots know exactly which questions to ask to save the day without wasting time.
3. The Scorecard (Measuring Efficiency)
The paper introduces a new way to grade these robots called "Gain per Question."
- Imagine you are the user. Every time you answer a question, it costs you time and mental energy.
- Bad Robot: Asks 10 questions to get a tiny bit of clarity. (High cost, low gain).
- Good Robot: Asks 1 specific question that solves the whole problem. (Low cost, high gain).
What Did They Find?
The researchers tested the smartest AI models available today and found some interesting patterns:
- GPT-5.2 is like a nervous student: It asks a lot of questions. It rarely fails, but it annoys the user by asking too much.
- Gemini models are like the "silent type": They guess a lot and often get it wrong because they are afraid to ask.
- The "Cost" of Asking: The paper also tested how the robot reacts if you tell it, "I'm very busy, don't bother me unless it's an emergency." The robots actually changed their behavior! They asked fewer questions when they thought you were busy, showing they can understand "cost."
Why Does This Matter?
We are moving toward a future where AI agents will do complex, long-term jobs (like managing your finances, coding software, or running a supply chain).
If these agents can't tell the difference between "I can figure this out" and "I need help," they will either:
- Fail silently and cause disasters.
- Nag you constantly and become useless.
LHAW is the first tool that lets us train and test these robots to be strategic. It teaches them to be the perfect employee: confident enough to work alone, but humble enough to ask for help when it truly matters.
In short: LHAW is a training simulator that teaches AI robots how to handle missing information without driving their human bosses crazy.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.