← Latest papers
💻 computer science

Not What You Asked For: Typographic Attacks in Household Robot Manipulation

This paper demonstrates that typographic attacks on household robots utilizing vision-language models can bypass visual judgment to cause physically consequential manipulation failures, achieving a 67.8% success rate in a simulated environment by poisoning the robot's semantic map and leading to the incorrect grasping and transport of objects.

Original authors: Ali Iranmanesh, Peng Liu

Published 2026-05-19
📖 4 min read☕ Coffee break read

Original authors: Ali Iranmanesh, Peng Liu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a household robot as a helpful but slightly naive assistant. It has two main ways of understanding the world:

  1. Its Eyes (Vision): It looks at an object and says, "That looks like a coffee mug."
  2. Its Brain (Language): It listens to your command, "Please bring me the coffee mug," and matches the words to what it sees.

In modern robots, these two systems are glued together into a single "brain" (called a Vision-Language Model). This is great because the robot can understand thousands of new things without needing to be retrained. However, the paper argues that this glue creates a dangerous loophole.

The "Magic Sticker" Trick

The researchers discovered that if you stick a piece of paper with the words "COFFEE MUG" printed on it onto a completely different object—like a toaster, a shoe, or a lamp—the robot's brain gets confused.

Because the robot's "language" and "vision" parts are so tightly linked, the robot stops looking at the object's shape and color. Instead, it starts reading the text first and looking later. It sees the words "COFFEE MUG" and immediately decides, "Ah, this must be the coffee mug," completely ignoring the fact that it's actually a toaster.

The Domino Effect: From a Glance to a Grasp

What makes this dangerous isn't just that the robot thinks the toaster is a mug. It's what happens next.

  1. The Poisoned Map: The robot doesn't just make a quick mistake; it writes this error into its permanent 3D mental map of the room. It's like the robot scribbles "This is a mug" on a sticky note and sticks it to the toaster in its memory.
  2. The Blind Trust: Even if the robot walks away and comes back, or if the sticker is hidden from view, it still trusts its mental map. It thinks, "I know there's a mug right there," because it wrote that down earlier.
  3. The Kinetic Failure: The robot then physically walks over, grabs the toaster, and tries to hand it to you or put it in a cup holder.

The paper calls this a "Kinetic Failure." It's not just a software glitch where the robot says the wrong thing; it's a physical action where the robot commits to the wrong object and carries it around the house.

The Experiment: How They Tested It

The researchers couldn't test this on real robots in real homes (it's too expensive and risky), so they used a highly realistic video game simulation called Habitat.

  • The Setup: They took a robot that was already good at finding and picking up objects.
  • The Attack: They placed a printed sticker with the name of the target object (e.g., "CUP") on a random, nearby object (a distractor).
  • The Result: Out of 59 specific test scenarios where the robot was already capable of doing the job, the sticker tricked the robot 67.8% of the time.

In the most successful cases, the robot didn't just grab the wrong thing; it successfully navigated to the wrong object, picked it up, carried it across the room, and tried to place it in the correct spot. The robot was fully "committed" to the mistake.

Why This Matters (According to the Paper)

The paper highlights that previous research looked at how stickers trick robots into navigating to the wrong place (like a drone flying to the wrong roof). But this is the first time anyone has shown that stickers can trick a robot into physically grabbing and moving the wrong item.

The danger is that the robot's "memory" (the 3D map) keeps the poison alive. Even if you take the sticker away, the robot might still try to pick up the wrong object because its internal map is corrupted.

The Proposed Fix

The authors suggest that we can't just fix the robot's "eyes" to ignore the text. Instead, we need to change how the robot uses its memory. They propose three safety checks:

  1. Double-Check: Don't trust a single glance. Make the robot look at the object from different angles before writing it into its permanent memory.
  2. Spot the Text: Have a secondary system that says, "Wait, there's a sticker on this object. Don't trust the label yet."
  3. Final Check: Right before the robot actually grabs the object, have it look at it one last time to make sure it hasn't been tricked.

In short: A piece of paper with words on it can trick a smart robot into thinking a toaster is a cup, and the robot will happily carry that toaster to your kitchen counter. The paper proves this is a real, physical risk for future home robots.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →