← Latest papers
💻 computer science

CT-1: Vision-Language-Camera Models Transfer Spatial Reasoning Knowledge to Camera-Controllable Video Generation

The paper introduces CT-1, a novel Vision-Language-Camera model that leverages a specialized data pipeline and wavelet-based regularization to transfer spatial reasoning knowledge into camera-controllable video generation, achieving significantly improved trajectory accuracy and physically plausible motion synthesis.

Original authors: Haoyu Zhao, Zihao Zhang, Jiaxi Gu, Haoran Chen, Qingping Zheng, Pin Tang, Yeyin Jin, Yuang Zhang, Junqi Cheng, Zenghui Lu, Peng Shu, Zuxuan Wu, Yu-Gang Jiang

Published 2026-04-13
📖 5 min read🧠 Deep dive

Original authors: Haoyu Zhao, Zihao Zhang, Jiaxi Gu, Haoran Chen, Qingping Zheng, Pin Tang, Yeyin Jin, Yuang Zhang, Junqi Cheng, Zenghui Lu, Peng Shu, Zuxuan Wu, Yu-Gang Jiang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a director filming a movie, but instead of holding a physical camera, you are talking to a robot that generates the video for you. You want to say, "The camera should slowly zoom in on the cat's eyes," or "Fly the camera backward while turning right to show the whole room."

In the past, telling a computer this was like trying to give directions to a blindfolded person who only speaks a different language. You either had to be incredibly vague (which led to the robot guessing wrong) or you had to manually type out complex math coordinates (which is boring and hard to do).

Enter CT-1, a new "super-intelligent camera operator" created by researchers. Here is how it works, explained simply:

1. The Problem: The "Blind" Camera

Existing video generators are like talented painters who can draw amazing scenes, but they don't understand how to move the "camera" around the scene.

  • If you say "Zoom in," they might just make the picture bigger (like a digital zoom) rather than actually moving the viewpoint forward.
  • If you say "Pan right," they might just slide the image sideways, breaking the 3D illusion.
  • The old way: You had to manually calculate the exact math for every second of the video. It was like trying to drive a car by manually calculating the rotation of every wheel.

2. The Solution: CT-1 (The "Brain" of the Operation)

The researchers built a new model called CT-1 (Camera Transformer 1). Think of CT-1 as a smart translator that sits between your brain and the video generator.

  • It has two eyes and a brain: It looks at a reference image (the scene) and reads your text instructions (the director's notes).
  • It does the math for you: Instead of you typing coordinates, CT-1 figures out exactly how the camera needs to move through 3D space to match your words. It predicts a smooth, realistic path, like a drone pilot flying a route.
  • It understands context: If you say "Zoom in on the cat," CT-1 knows the cat is in the center of the image, so it moves the camera forward toward the cat, not just randomly.

3. The Secret Sauce: The "Wavelet" Filter

One of the coolest parts of this paper is how they taught the model to move smoothly.
Imagine you are drawing a line on a piece of paper.

  • The "Jitter" Problem: Sometimes, when computers try to draw a line, they shake a little bit (like a shaky hand). This makes the video look glitchy.
  • The Wavelet Trick: The researchers used a mathematical tool called a Wavelet Transform. Think of this like a noise-canceling headphone for movement.
    • It separates the "big picture" movement (the smooth path you want) from the "tiny shakes" (the jitter).
    • It tells the model: "Keep the smooth, big movements, but throw away the tiny, jittery shakes."
    • This ensures the camera glides like a professional drone operator, not a nervous tourist.

4. The Training Data: The "Camera Gym"

To teach CT-1 how to be this good, the researchers couldn't just use existing videos because they didn't have the "camera movement notes" attached.

  • They built a massive new dataset called CT-200K.
  • Imagine a giant gym where millions of videos are being watched by AI. These AIs are trained to look at a video and write a script describing exactly how the camera moved, frame by frame.
  • They created 47 million frames of training data. It's like giving the model a library of every possible camera move in the world so it can learn the rules of physics and perspective.

5. The Result: A Magic Director

When you put CT-1 together with a video generator:

  1. You upload a picture and type: "The camera flies over the city, then dives down into the street."
  2. CT-1 instantly calculates the perfect flight path.
  3. It hands this path to the video generator.
  4. The result is a video where the camera actually flies and dives in a physically realistic way, matching your description perfectly.

Why This Matters

Before this, making a video with specific camera moves was like trying to sculpt a statue with a sledgehammer—imprecise and messy.
CT-1 is like giving you a laser-guided chisel. It bridges the gap between your imagination (what you want to see) and the computer's execution (how the video is made). It makes video generation feel less like magic and more like a conversation with a skilled cameraman who knows exactly what you mean.

In short: CT-1 is the smart assistant that translates your "Director's Vision" into a perfect, smooth, 3D camera flight path, so you can finally get the video you imagined without needing a degree in mathematics.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →