← Latest papers
💻 computer science

Video-Based Optimal Transport for Feedback-Efficient Offline Preference-Based Reinforcement Learning

The paper introduces VOTP, a semi-supervised framework that leverages Video Foundation Models and optimal transport to generate high-fidelity pseudo-labels from minimal human feedback, enabling efficient and robust reward learning for offline preference-based reinforcement learning.

Original authors: Tung M. Luu, Hwanhee Kim, Younghwan Lee, Chang D. Yoo

Published 2026-06-16
📖 4 min read☕ Coffee break read

Original authors: Tung M. Luu, Hwanhee Kim, Younghwan Lee, Chang D. Yoo

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot how to do a complex task, like opening a drawer or walking without stumbling. In the world of robotics, the robot needs a "reward system" to know what it's doing right. Usually, human engineers have to painstakingly write code to tell the robot, "Good job, you moved your arm up," or "Bad job, you dropped the object." This is like trying to write a recipe for a dish you've never cooked; it's hard, and you might get the instructions wrong.

To fix this, scientists use Preference-Based Reinforcement Learning (PbRL). Instead of writing code, they show the robot two videos of it doing the task and ask a human, "Which one looks better?" The robot learns from these choices.

The Problem:
Asking a human to watch videos and pick the "better" one is slow and expensive. To teach a robot well, you might need thousands of these comparisons. It's like trying to teach a child to ride a bike by stopping them after every single wobble to ask, "Was that good?" It takes forever.

The Solution: VOTP
The paper introduces a new method called VOTP (Video-based Optimal Transport Preference). Think of VOTP as a clever teaching assistant that helps you teach the robot with very few human questions.

Here is how it works, using a simple analogy:

1. The "Smart Eye" (Video Foundation Models)

First, VOTP uses a pre-trained "Smart Eye" (a Video Foundation Model). Imagine this eye has watched millions of hours of human videos—people dancing, cooking, running, and playing. Because it has seen so much, it understands what "good movement" looks like, even if it's never seen a robot before. It can turn a video of a robot moving into a mathematical "fingerprint" that captures the essence of the action.

2. The "Matchmaker" (Optimal Transport)

This is the magic part.

  • The Setup: You give the system a tiny handful of examples (say, 10 pairs of videos) where a human has already said, "Video A is better than Video B."
  • The Mountain of Data: You also have a huge mountain of unlabeled videos (thousands of pairs) where no human has said anything yet.
  • The Matchmaking: VOTP uses a mathematical tool called Optimal Transport. Imagine you have a pile of "Good" fingerprints and a pile of "Bad" fingerprints from your 10 human examples. Now, you have a huge pile of "Unknown" fingerprints from the unlabeled videos.
  • The Connection: Optimal Transport acts like a super-efficient matchmaker. It looks at an "Unknown" video and asks, "Which of the 10 human-approved videos does this look most like?" It doesn't just guess; it calculates the best possible way to connect the unknown videos to the known ones based on how similar their "fingerprints" are.

3. The "Inference" (Pseudo-Labels)

Once the matchmaker connects an unknown video to a known "Good" video, VOTP says, "If this unknown video looks like the one the human liked, then this unknown video must also be good!" It automatically assigns a label (a "pseudo-label") to the thousands of videos you didn't have time to watch.

4. The Result

Now, instead of teaching the robot with just 10 human examples, the robot gets to learn from 10 human examples plus thousands of automatically labeled examples. The robot learns much faster and better, even though you only asked the human for a tiny bit of help.

What the Paper Found

The authors tested this on robots that had to walk (locomotion) and robots that had to move objects (manipulation). They also tested it on a real robot arm in a lab.

  • Less Human Work: VOTP achieved top results with only a handful of human labels (sometimes as few as 10).
  • Better than the Rest: It beat other methods that tried to do the same thing, often needing hundreds or thousands of labels to get similar results.
  • Robustness: Even if the lighting changed or there were distracting things in the background (like a video playing on a TV behind the robot), VOTP still figured out what was good and what was bad.
  • Real World: When they tried it on a real robot arm lifting a banana or opening a drawer, VOTP helped the robot succeed much more often than methods that relied only on the few human labels.

In short: VOTP is a way to teach robots by showing them a few examples of what humans like, and then using a "smart eye" and a "mathematical matchmaker" to figure out what the robot should do for the rest of the videos on its own. It saves humans from doing the boring, repetitive work of labeling thousands of videos.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →