← Latest papers
💻 computer science

TRANSPORTER: Transferring Visual Semantics from VLM Manifolds

This paper introduces TRANSPORTER, a model-independent approach that transfers visual semantics from Vision Language Model (VLM) manifolds to generative models via an optimal transport coupling, enabling the creation of high-fidelity videos that visualize and interpret the underlying rules driving VLM predictions.

Original authors: Alexandros Stergiou

Published 2026-04-08
📖 4 min read☕ Coffee break read

Original authors: Alexandros Stergiou

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a super-smart robot that can watch a video and tell you exactly what's happening: "A man is running fast on a beach." This robot is a Vision-Language Model (VLM). It's great at understanding, but it's a bit of a "black box." If you ask it, "Why did you think he was running fast?" it might give you a confusing list of numbers or a vague sentence. It doesn't really show you how it reached that conclusion.

This paper introduces a new tool called TRANSPORTER to solve that mystery. Think of it as a "Dream Machine for Robot Thoughts."

Here is how it works, using simple analogies:

1. The Problem: The Robot's Secret Language

The robot (VLM) doesn't speak human language internally; it speaks in logits. Imagine these are like a giant scoreboard with thousands of slots. When the robot sees a video, it lights up certain slots with numbers.

  • If the slot for "running" lights up high, the robot thinks the person is running.
  • If the slot for "slow" lights up, it thinks they are walking.

The problem is, we can't see the video inside the robot's brain. We only see the final answer.

2. The Solution: TRANSPORTER (The Translator)

TRANSPORTER acts like a magical translator that turns those abstract numbers (logits) back into a video. It asks the question: "What would the video look like if the robot's 'running' score was super high, but the 'walking' score was zero?"

It doesn't just guess; it uses a mathematical trick called Optimal Transport.

  • The Analogy: Imagine you have a pile of sand (the robot's abstract numbers) and you want to move it to shape a perfect sandcastle (a video). Moving sand is messy and hard. Optimal Transport is like finding the most efficient, smooth path to move every grain of sand from the pile to the castle without spilling a drop. TRANSPORTER learns this path so it can perfectly reshape the robot's "thoughts" into a visual movie.

3. How It Works in Practice

The researchers taught TRANSPORTER to act like a video editor that only listens to the robot's internal scores.

  • The Setup: They show the robot a video of a "red bowling ball." The robot gives it a score.
  • The Magic: They tell TRANSPORTER: "Hey, change the score so the robot thinks it's a blue bowling ball instead."
  • The Result: TRANSPORTER generates a brand new video. It keeps the bowling alley, the pins, and the motion exactly the same, but the ball magically turns blue.

They did this for all sorts of changes:

  • Objects: Turning a "book" into a "newspaper."
  • Actions: Changing "walking" to "running."
  • Scenes: Changing a "sunny day" to a "rainy day."

4. Why This Matters

Before this, if we wanted to understand why a robot made a mistake, we had to guess or look at blurry heatmaps (like looking at a foggy window).

With TRANSPORTER, we can visualize the robot's reasoning.

  • If the robot thinks a video is "funny," TRANSPORTER can generate a video that shows exactly what features make it funny.
  • If the robot gets confused between "running" and "spinning," TRANSPORTER can show us a video that slowly morphs from one to the other, revealing exactly where the robot gets confused.

The Big Picture

Think of TRANSPORTER as a mirror for AI.

  • Old way: You ask the AI a question, and it gives you a text answer. You have to trust it.
  • New way (with TRANSPORTER): You ask the AI what it sees, and it shows you a video of its thought process. It's like the AI saying, "I thought you were running because I saw your legs moving this fast," and then it plays a video clip of exactly that movement.

This paper proves that we can take the invisible, mathematical "thoughts" of a complex AI and turn them into high-quality, understandable videos, helping us finally understand how these machines see our world.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →