Ask Early, Ask Late, Ask Right: When Does Clarification Timing Matter for Long-Horizon Agents?
This paper introduces a forced-injection framework to demonstrate that the optimal timing for AI agent clarification is task-dependent—specifically, goal clarification must occur within the first 10% of execution while input clarification remains valuable up to 50%—revealing that current frontier models fail to ask at these empirically determined critical windows.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are hiring a super-smart robot assistant to build a complex house for you. You give it a list of instructions, but you accidentally leave out a few crucial details.
The big question this paper asks is: When should the robot stop and ask you for clarification?
Should it ask immediately before laying the first brick? Should it wait until the roof is halfway up? Or should it just guess and hope for the best?
The authors of this paper decided to test this by running over 6,000 experiments with different AI models. They didn't wait for the robots to figure it out on their own; instead, they played "God" and forced the missing information to appear at specific times during the building process to see how much it helped.
Here is what they found, explained simply:
1. The "What" Matters More Than the "When"
The most surprising discovery is that there is no single "best time" to ask a question. It depends entirely on what information is missing. Think of it like different types of construction errors:
- The "Goal" (What are we building?): This is the most critical piece of info.
- The Analogy: Imagine the robot starts building a garage because you didn't specify, but you actually wanted a swimming pool.
- The Finding: If the robot asks about the goal after it has already laid 10% of the foundation, it's too late. The damage is done. The value of asking drops to almost zero immediately. You have to clarify the goal right at the start.
- The "Input" (What materials do we have?): This is about data or resources.
- The Analogy: Imagine the robot starts building, but you forgot to tell it whether to use red bricks or blue bricks.
- The Finding: The robot can keep working for a while, maybe even building half the walls, before it realizes it needs to know the color. Clarification here is still very helpful up until the 50% mark of the project.
- The "Constraints" (What rules must we follow?):
- The Analogy: Imagine the robot builds a beautiful house, but you forgot to tell it that no windows can be higher than 5 feet.
- The Finding: If the robot builds the whole house and then you tell it the rule, it might have to tear everything down. However, if you tell it the rule halfway through, it can sometimes adjust, though it's messy.
2. The "Point of No Return"
The paper found that for "Goal" questions, there is a very narrow window. If you wait until the robot is 10% done, asking "Wait, what are we building?" is useless. The robot has already committed to a path, and changing it now is too expensive (in terms of wasted time and effort).
For "Input" questions, the window is wider. The robot can explore and figure things out on its own for a while, so you have until the halfway point to step in.
3. The Robots Are Bad at Timing
The researchers also watched how these AI robots behave when they are allowed to ask questions naturally (without the researchers forcing the answers).
- The Over-Askers: One model (GPT-5.2) asked questions constantly, but often too late. It was like a builder who keeps asking "What color is the paint?" after the walls are already dry.
- The Under-Askers: Another model (Gemini) almost never asked questions, even when it was clearly confused. It just kept guessing.
- The Reality Check: None of the current top-tier robots naturally know the "sweet spot" for asking. They either ask too late (when the goal is already set) or not at all.
4. The Cost of Waiting
The paper measured "wasted compute," which is basically the amount of work the robot does that turns out to be useless because it was based on a wrong guess.
- Early Clarification: Very little wasted work.
- Late Clarification: A lot of wasted work. The longer you wait to clarify the "Goal," the more the robot builds the wrong thing, and the more time is thrown away.
The Bottom Line
The paper concludes that timing is everything, but it depends on the topic.
- If you are unsure about the Goal, speak up in the first few seconds.
- If you are unsure about the Data, you have a bit more time (up to halfway).
- If you are unsure about Rules, it's best to say them before the robot starts, but it's better late than never.
Currently, AI agents are terrible at knowing when to speak up. They need to be taught to ask the right questions at the right time, or they will keep wasting time building the wrong house.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.