← Latest papers
💻 computer science

GIFT: Generalizing Intent for Flexible Test-Time Rewards

The paper presents GIFT, a framework that leverages language models to infer high-level human intent from demonstrations, enabling reward functions to generalize robustly to novel environments and objects without retraining by mapping test states to behaviorally equivalent training states based on intent rather than surface-level cues.

Original authors: Fin Amin, Nathaniel Dennler, Andreea Bobu

Published 2026-03-25
📖 4 min read☕ Coffee break read

Original authors: Fin Amin, Nathaniel Dennler, Andreea Bobu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a robot how to pack a bag for an art class. You show it a few times: "Put the paintbrush in the bag. Put the sketchbook in the bag."

The robot learns this. But here's the problem: if you later ask the robot to pack for a different art session, and you hand it a lump of molding clay, a standard robot might get confused.

Why? Because the robot didn't actually learn the concept of "art supplies." Instead, it learned a silly rule like: "If it looks like a paintbrush, put it in the bag." So, when it sees a dish scrubber (which looks a bit like a brush) or a toothbrush (which has the word "brush" in its name), it might try to pack those instead of the clay. It's following the surface details, not the real goal.

This paper introduces a new framework called GIFT (Generalizing Intent for Flexible Test-Time rewards) to fix this.

The Core Idea: The "Intent Detective"

Think of GIFT as a robot that doesn't just memorize pictures; it acts like a detective trying to figure out your motive.

  1. The Old Way (Visual/Language Similarity):

    • Visual: The robot sees a dish scrubber and a paintbrush. They look similar (both have bristles). The robot thinks, "They are the same!" and packs the scrubber.
    • Language: The robot reads "toothbrush" and "paintbrush." They both end in "brush." The robot thinks, "They are the same!" and packs the toothbrush.
    • Result: The robot fails because it's looking at the wrong clues.
  2. The GIFT Way (Intent Similarity):

    • GIFT uses a large language model (like a super-smart AI chatbot) to act as a translator.
    • It looks at what you did (putting art stuff in a bag) and what you didn't do (leaving it on the table).
    • It asks the AI: "What is the human's real goal here?"
    • The AI answers: "The goal is 'Pack Art Supplies'."

How It Works: The "Shape-Shifting" Trick

Once GIFT knows the goal is "Pack Art Supplies," it performs a magic trick called Intent-Conditioned Alignment.

Imagine the robot is looking at a table full of new objects it has never seen before: a lump of clay, a screwdriver, and a paintbrush.

  • Without GIFT: The robot sees the clay and thinks, "I've never seen this. I don't know what to do."
  • With GIFT: The robot asks the AI, "Under the rule 'Pack Art Supplies,' what is this lump of clay equivalent to?"
    • The AI says, "Ah, under that rule, the clay is the same as the paintbrush you saw in training."
    • The AI also says, "The screwdriver is a distractor; it's not art."

GIFT then mentally swaps the clay for the paintbrush in its memory. It tells the robot: "Treat this clay exactly like you treated that paintbrush."

Now, the robot can use the exact same instructions it learned yesterday to handle a completely new object today, without needing to be retrained.

The Real-World Test

The researchers tested this on a real robot arm (a Franka Panda) and in a computer simulation.

  • The Setup: They taught the robot to pack "art supplies" or "valuable items" using a few examples.
  • The Test: They gave the robot new objects it had never seen (like molding clay or a specific iPhone model) and mixed in "trick" objects (like a dish scrubber or a paper ring).
  • The Result:
    • The old methods (looking at pictures or words) got tricked by the dish scrubbers and paper rings.
    • GIFT correctly identified that the clay was "art" and the iPhone was "valuable," ignoring the trick objects. It won the game almost every time.

The Big Takeaway

Most robots are like students who memorize the answers to a specific test. If you change the questions slightly, they fail.

GIFT teaches the robot to understand the subject matter. It's the difference between memorizing that "2 + 2 = 4" and understanding the concept of addition. Once you understand the concept, you can solve "2 + 2" even if the numbers are written in a different font, or if you're adding "apples" instead of "numbers."

By focusing on human intent (the "why") rather than surface features (the "what it looks like"), GIFT allows robots to be flexible, smart, and ready for the real world.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →