Diagnosing Semantic Handoff Failures in Agent-Orchestrated Vision-Language-Action Skill Composition
This paper identifies and diagnoses "semantic handoff failures" in long-horizon robot tasks, revealing that while individual vision-language-action skills perform well in isolation, they frequently fail when chained together due to unaddressed state mismatches, target grounding issues, and control execution errors that clean training data fails to represent.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a robot to make a sandwich. You don't give it one giant command like "make a sandwich." Instead, you break it down into tiny, specific instructions: "walk to the fridge," "open the door," "grab the cheese," "put cheese on bread," and "close the fridge."
This paper is about what happens when these tiny instructions are passed from one to the next. The authors call this the "Semantic Handoff Problem."
Here is the simple breakdown of their findings:
1. The "Perfect Finish, Bad Start" Problem
Imagine a robot is told to "walk to the fridge." It does a great job! It walks right up to the fridge and stops. The instruction is technically complete.
But here is the catch: The robot stopped too far away to actually reach the handle. Or, it stopped at an angle where its arm is twisted and can't grab the door.
In the robot's world, the "walk" task is done. But for the next task ("grab the door"), the robot is in a terrible position. The first skill succeeded, but it left the robot in a state where the second skill can't start. This is a Semantic Handoff Failure. The robot finished the job, but it didn't leave the scene ready for the next job.
2. The "Clean Room" vs. The "Messy Kitchen"
The researchers tested their robot skills in two different ways:
- The Clean Room (Snapshot Testing): They reset the robot to a perfect, pre-recorded starting position every time they tested a single skill. It's like practicing a piano scale with your hands perfectly placed on the keys. In this "clean" environment, the robot was amazing. It could grab things, open doors, and press buttons with near-perfect success (77% to 100%).
- The Messy Kitchen (Chained Testing): Then, they let the robot do a whole sequence of tasks without resetting. The robot had to walk, then grab, then place, then open. By the time it got to the second or third step, the robot was in a "messy" state—maybe the camera was looking at the wrong angle, or the object was slightly shifted.
The Shocking Result: Even though the robot was a master in the "Clean Room," it fell apart in the "Messy Kitchen." When the skills were chained together, the robot would get stuck. It would try to grab an object but miss because it was standing at a weird angle caused by the previous step.
3. The "Referee" System
To figure out why the robot was failing, the authors built a special "Agent Harness." Think of this as a strict referee standing between every step.
- The Plan: The agent decides the next move.
- The Act: The robot tries to do it.
- The Verify: Before the robot is allowed to move to the next step, the referee (using a smart camera system) checks: "Is the robot actually ready for the next step?"
- Example: The referee won't let the robot say "I'm done walking" just because it sees the fridge. The referee checks: "Is the fridge close enough to actually grab?" If not, the robot has to walk a bit more.
- The Replan: If the robot fails the check, the agent doesn't just give up. It says, "Okay, that didn't work. Let's try walking again" or "Let's try grabbing from a different angle."
4. What They Learned
By watching the robot fail and using the referee to diagnose the problem, they found that the robot wasn't "dumb." The robot knew how to walk and how to grab. The problem was transitioning.
- The Diagnosis: Most failures happened because the robot finished one task but didn't leave itself in a good position for the next one.
- The Solution: They realized that training robots on "perfect" starting positions (the Clean Room) isn't enough. To make robots reliable in the real world, they need to be trained on the "messy" states that happen after a previous task is done. They need to learn how to handle the chaos of a real kitchen, not just a perfect lab.
Summary
The paper argues that we can't just teach robots individual tricks and expect them to work together. We need to teach them how to hand off tasks smoothly. The robot needs to finish one job in a way that makes the next job easy to start.
Their "Agent Harness" acts like a coach that watches the robot, spots when it's setting itself up for failure, and forces it to try again until it gets the handoff right. This turns a robot that fails almost every long task into a system that can at least tell you exactly where and why it got stuck.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.