← Latest papers
💻 computer science

AnySlot: Goal-Conditioned Vision-Language-Action Policies for Zero-Shot Slot-Level Placement

The paper introduces AnySlot, a framework that achieves zero-shot, sub-centimeter precision in slot-level robotic placement by decoupling language grounding from control through an explicit visual goal representation, accompanied by the new SlotBench benchmark for evaluating such structured spatial reasoning tasks.

Original authors: Zhaofeng Hu, Sifan Zhou, Qinbo Zhang, Rongtao Xu, Qi Su, Ci-Jyun Liang

Published 2026-04-15
📖 4 min read☕ Coffee break read

Original authors: Zhaofeng Hu, Sifan Zhou, Qinbo Zhang, Rongtao Xu, Qi Su, Ci-Jyun Liang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a robot to play a very specific game: "Put this red block into the third slot from the left, but only if it's the biggest one, and make sure it doesn't touch the blue box."

For a human, this is easy. You look, you think, you point, and you move your hand. For a robot, this is a nightmare.

This paper introduces a new system called AnySlot that solves this problem by changing how the robot thinks. Instead of trying to do everything in one giant brain-melt, it breaks the job into two distinct steps: The Architect and The Builder.

Here is the breakdown using simple analogies:

1. The Problem: The "Confused Chef"

Current robots (called "Flat VLA policies") are like a chef who tries to read a recipe, chop the vegetables, stir the pot, and plate the food all at the exact same time.

  • The Issue: If the recipe says "put the soup in the second bowl from the left," the robot gets confused. It tries to guess the coordinates while simultaneously moving its arm.
  • The Result: It often picks the wrong bowl or drops the soup because it's trying to do too much math in its head at once. It's like trying to solve a complex math equation while riding a unicycle.

2. The Solution: The "Architect and Builder" (AnySlot)

The authors realized that to be precise, you need to separate planning from doing. They created a two-stage system:

Stage 1: The Architect (High-Level Goal)

Instead of asking the robot to "calculate the X and Y coordinates," they ask an Image Generator (like a digital artist) to draw a picture of the goal.

  • The Analogy: Imagine you are giving directions to a friend. Instead of saying, "Walk 42 steps north, then turn 15 degrees," you simply draw a big blue dot on a map right where they need to go.
  • How it works: The robot reads the sentence ("Put the block in the second slot") and asks an AI artist to generate an image where a blue sphere is drawn exactly on top of that specific slot.
  • The Magic: This blue sphere becomes a "visual target." It turns a confusing language puzzle into a simple visual game: "Go to the blue dot."

Stage 2: The Builder (Low-Level Policy)

Now, the robot's arm doesn't have to think about which slot is the "second largest." It just has to follow the blue dot.

  • The Analogy: The robot is now a delivery driver who just needs to drive to the house with the blue flag on the roof. It doesn't need to know the street name or the zip code; it just follows the flag.
  • The Result: Because the robot only has to focus on "move arm to blue dot," it can be incredibly precise (sub-centimeter accuracy). It ignores the complex language and focuses purely on the visual target.

3. The New Test: "SlotBench"

The authors realized there was no good way to test if robots could actually do this kind of precise, logic-heavy work. So, they built a video game called SlotBench.

  • Think of it as a gym for robot brains. It has 9 different levels of difficulty, ranging from simple ("pick the top slot") to very tricky ("pick the slot that is furthest from the red block but closest to the Pepsi can").
  • They tested old robots on this gym, and most failed miserably. They couldn't handle the logic.
  • AnySlot walked into the gym and aced almost every level (90% success rate) because it used the "Blue Dot" trick to simplify the task.

4. Why This Matters

  • Precision: Old robots were like a person trying to thread a needle while wearing boxing gloves. AnySlot is like a surgeon with steady hands.
  • Zero-Shot Learning: This means the robot can learn to do these tasks without being specifically trained on every single possible arrangement of blocks. If you give it a new puzzle it has never seen before, it can still figure out where the "blue dot" should go and place the object perfectly.

Summary

AnySlot is a framework that stops robots from trying to be poets and mathematicians at the same time.

  1. Translate: Turn complex language instructions into a simple visual picture (a blue dot on the target).
  2. Execute: Let the robot's arm simply follow that picture.

By turning "thinking" into "drawing," the robot can finally place objects with the precision needed for real-world factories and assembly lines.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →