← Latest papers
🤖 AI

TriMotion: Modality-Agnostic Camera Control for Video Generation

TriMotion is a modality-agnostic framework that unifies video, pose, and text inputs into a shared motion embedding space via a new dataset and latent consistency objective, enabling high-quality, flexible camera-controlled video generation across heterogeneous user inputs.

Original authors: Seunghyun Shin, Jifei Song, Wooseok Jeon, Hae-Gon Jeon, Jiankang Deng

Published 2026-06-23
📖 5 min read🧠 Deep dive

Original authors: Seunghyun Shin, Jifei Song, Wooseok Jeon, Hae-Gon Jeon, Jiankang Deng

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a movie director. You have a script (a text description), a storyboard (a reference video), and a technical blueprint (camera coordinates). In the past, if you wanted an AI to generate a video with a specific camera movement—like a smooth zoom-in or a spinning orbit—you had to speak the AI's specific language. If the AI only understood blueprints, you couldn't use your storyboard. If it only understood storyboards, you couldn't use your text script. You were stuck with one tool for one job.

TriMotion is a new "universal translator" for AI video generation. It allows you to control the camera using any of those three tools: text, a reference video, or technical camera data. The AI understands all of them as the same thing.

Here is how it works, broken down into simple concepts:

1. The Problem: The "Language Barrier"

Currently, most AI video tools are like specialists who only speak one dialect.

  • The Blueprint Specialist: Needs exact numbers (e.g., "move the camera 5 meters left"). This is hard for normal people to write.
  • The Storyboard Specialist: Needs a video clip to copy the movement from. This is inflexible if you want to change the speed or direction.
  • The Script Specialist: Needs a text description (e.g., "the camera pans left"). But AI often struggles to turn vague words into precise, smooth physical movements.

2. The Solution: The "Shared Dance Floor"

The researchers built a system called TriMotion. Think of the AI's internal brain as a giant dance floor.

  • Before, if you gave the AI a text description, it tried to dance to the text. If you gave it a video, it tried to dance to the video. They were different dances.
  • TriMotion teaches the AI to translate everything onto the same dance floor. Whether you give it a text script, a reference video, or a camera blueprint, the AI converts all of them into the exact same "dance steps" (mathematical numbers) inside its brain.
  • Because they are all on the same dance floor, the AI can switch between them instantly. You can start with a text description and then swap it for a video reference, and the camera movement stays perfectly consistent.

3. The Training: The "Motion Triplet"

To teach the AI this universal language, the researchers created a special training dataset called the Motion Triplet Dataset.

  • Imagine a set of training cards. Each card has three things on it:
    1. A video clip of a camera moving.
    2. The exact mathematical coordinates of that camera's path.
    3. A written description of what the camera did (e.g., "The camera slowly zooms in while turning right").
  • The AI studied millions of these cards. It learned that the math, the video, and the words all describe the exact same physical motion. This allowed it to build that "shared dance floor" where all three inputs mean the same thing.

4. The Secret Sauce: "Ghost Tracking"

One of the biggest challenges in AI video is that the AI might get the look right but the movement wrong (the camera might drift or jitter).

  • TriMotion uses a clever trick called Latent Motion Consistency.
  • Imagine the AI is drawing a picture. Usually, it draws the whole picture, then checks if the lines are straight. If they aren't, it erases and redraws. This is slow and can ruin the picture.
  • TriMotion has a "Ghost Tracker." As the AI is drawing the video in its "sketch phase" (before the final image is fully formed), this Ghost Tracker checks the movement immediately. It whispers to the AI: "Hey, the camera is supposed to be moving left, but your sketch is drifting right. Fix it now."
  • This happens inside the AI's "sketch" (latent space), so it doesn't have to waste time erasing and redrawing the final high-quality image. It keeps the movement smooth and accurate from the very first step.

5. The Results: What Can It Do?

The paper shows that TriMotion creates high-quality videos that follow the camera instructions perfectly, no matter which input you use.

  • Text to Video: You type "a slow zoom into a forest," and the camera moves smoothly.
  • Video to Video: You show a clip of a camera spinning, and the AI applies that exact spin to a new scene.
  • The "Mix-and-Match" Trick: Because everything is on the same dance floor, you can do cool things like Motion Composition. You can tell the AI to "start with a zoom-in" (from text) and then "switch to a spin" (from a video reference), and the AI will blend them seamlessly into one continuous shot without needing to be retrained.

Summary

TriMotion is like giving a movie director a universal remote control. Whether they speak in words, show a sample clip, or hand over a technical diagram, the AI understands the instruction as the exact same camera movement. It uses a special "training dataset" to learn this translation and a "ghost tracker" to ensure the camera moves smoothly and accurately, resulting in professional-looking videos that follow the director's vision perfectly.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →