← Latest papers
🤖 machine learning

HorizonRelight: Relighting Long-horizon Videos Consistently via Diffusion Transformers

Original authors: Jing Yang, Mayoore Jaiswal, Zian Wang, Steven Zeng, Rochelle Pereira, Yajie Zhao, Jianyuan Min

Published 2026-06-30
📖 4 min read☕ Coffee break read

Original authors: Jing Yang, Mayoore Jaiswal, Zian Wang, Steven Zeng, Rochelle Pereira, Yajie Zhao, Jianyuan Min

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a home video of a family dinner, and you want to change the lighting from a warm, cozy sunset to a cool, moody blue moonlight. You want the whole video to look like it was filmed under that new light, without the characters flickering or the room suddenly changing shape.

This is the problem HorizonRelight solves.

Here is the simple breakdown of how it works, using everyday analogies:

The Problem: The "Jigsaw Puzzle" Glitch

Current video AI tools are like artists who can only paint on small, 5-second sticky notes. If you give them a 2-minute video, they chop it up into 5-second chunks, paint each one, and tape them back together.

  • The Issue: Because the artist paints each sticky note independently, the end of one note might look slightly different from the start of the next. When you tape them together, you see a "seam" or a glitch where the lighting jumps or the object flickers.
  • The Paper's Term: This is called "temporal discontinuity at chunk boundaries."

The Solution: The "Conveyor Belt" Approach

HorizonRelight changes the workflow. Instead of painting 5-second notes in isolation, it treats the video like a conveyor belt.

  1. The "Warm-Start" (The First Step):
    Before the conveyor belt starts moving, the system takes a single "starter image" (generated by a smart AI like ChatGPT or a specialized tool) that shows exactly what the scene should look like under the new light. Think of this as setting the first domino. It tells the AI, "Start here, and keep this look."

  2. The "Passing the Torch" (Chained Propagation):
    This is the magic trick. When the AI finishes painting the first 5-second chunk, it doesn't just stop. It takes the very last few frames of that chunk and hands them to the AI working on the next chunk.

    • Analogy: Imagine a relay race. The runner (the video chunk) doesn't just run their own race; they pass the baton (the visual state) to the next runner. The next runner starts exactly where the previous one left off, ensuring the lighting and shadows flow smoothly without any jumps.
  3. The "Training Gym" (Masked Self-Conditioning):
    To teach the AI how to do this relay race perfectly, the researchers trained it in a special way. During training, they would hide (mask) parts of the video and ask the AI to guess what comes next based only on the frames it just saw.

    • Analogy: It's like a teacher showing a student the first half of a sentence and asking them to finish it. By practicing this over and over, the AI learns to "continue" the story naturally rather than restarting it every time.

What This Achieves

  • No More Flickering: Because the AI is constantly "remembering" the end of the previous chunk, the lighting stays consistent. You won't see the gray ball in the video suddenly change color when the video cuts to the next 5-second segment.
  • Long Videos: This allows the system to handle long videos (like a whole movie scene or a long travel vlog) without the quality degrading or the lighting getting weird halfway through.
  • Style Transfer: The paper notes this same "starter image" trick can be used to change the style of the video (e.g., turning a real video into a cartoon) while keeping the lighting consistent, not just changing the light source.

The Catch (Limitations)

The paper is honest about one limitation: If you make the video extremely long (much longer than a typical movie scene, say over 300 frames), the fine details (like the texture of a shirt or the sharpness of a face) might start to get a little blurry. The lighting stays perfect, but the "crispness" fades slightly the further you go down the conveyor belt.

Summary

HorizonRelight is a new way to change the lighting in long videos. Instead of treating the video as a stack of separate, disconnected clips, it treats it as one continuous stream where the end of one clip helps paint the beginning of the next. This ensures the new lighting looks natural and steady from start to finish.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →