← Latest papers
💻 computer science

HDR Video Generation via Latent Alignment with Logarithmic Encoding

This paper demonstrates that high-quality HDR video generation can be achieved by leveraging pretrained generative models through logarithmic encoding to align HDR imagery with existing latent spaces, combined with a camera-mimicking degradation strategy to infer missing details via lightweight fine-tuning.

Original authors: Naomi Ken Korem, Mohamed Oumoumad, Harel Cain, Matan Ben Yosef, Urska Jelercic, Ofir Bibi, Yaron Inger, Or Patashnik, Daniel Cohen-Or

Published 2026-04-15
📖 5 min read🧠 Deep dive

Original authors: Naomi Ken Korem, Mohamed Oumoumad, Harel Cain, Matan Ben Yosef, Urska Jelercic, Ofir Bibi, Yaron Inger, Or Patashnik, Daniel Cohen-Or

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have an old, black-and-white movie (SDR). It's clear, but it lacks the depth, the blinding brightness of the sun, and the deep, velvety darkness of a shadow. You want to turn it into a modern, high-definition, "High Dynamic Range" (HDR) movie that looks like it was shot with a $50,000 camera.

Usually, to do this, you'd think you need to build a brand-new, super-complex robot from scratch to understand what HDR looks like. But the authors of this paper, LumiVid, say: "Wait a minute. We already have a robot that knows how to paint beautiful pictures. We just need to teach it how to speak HDR."

Here is the simple breakdown of how they did it, using some everyday analogies:

1. The Problem: Speaking Different Languages

Think of the AI model they are using as a master chef who has spent years cooking only with standard, pre-packaged ingredients (SDR images). This chef knows exactly how to make a perfect burger or pizza.

Now, you want this chef to cook a gourmet, 5-star meal using raw, unprocessed, wild ingredients (HDR data).

  • The Issue: The wild ingredients are too intense. They are too bright (like a blinding sun) or too dark (like a cave). If you just hand them to the chef, they get confused. The chef's kitchen (the AI's "brain") isn't built to handle ingredients that are "off the charts."
  • The Old Way: Most researchers tried to build a whole new kitchen or train a new chef from scratch to handle these wild ingredients. This takes years and huge amounts of money.

2. The Solution: The "Translator" (Logarithmic Encoding)

The authors realized they didn't need a new chef. They just needed a translator.

They found a specific way to "translate" the wild HDR ingredients into a format the chef already understands. They used a Logarithmic Encoding (specifically called LogC3).

  • The Analogy: Imagine the HDR data is a volume knob turned all the way up to 100, which would blow out the speakers. The LogC3 translator is like a smart volume compressor. It takes that massive, overwhelming sound and compresses it down to a range the speaker (the AI) can handle, without losing the nuance of the music.
  • The Result: The AI thinks it's still cooking with its familiar ingredients, but it's actually processing the high-end HDR data. Because the data is now "familiar" to the AI, it can use all its existing knowledge to make the picture look amazing.

3. The Secret Sauce: The "Amnesia" Training

Here is the tricky part. In the real world, when you take a photo of a bright sun, the camera often "blows out" the details (it turns white). When you look at a dark shadow, it often turns into a black blob. The HDR data should have those details, but the input video (SDR) doesn't.

If you just show the AI the SDR video, it will just copy what it sees. It won't invent the missing details.

  • The Trick: The authors intentionally ruined the input video during training. They took the SDR video and deliberately added "noise," blurred the bright spots, and crushed the dark shadows.
  • The Analogy: Imagine you are teaching a student to draw a sunset. Instead of showing them a perfect photo, you show them a photo where the sun is a blank white circle and the sky is a muddy gray.
    • You tell them: "I know the sun is actually glowing and the sky has orange clouds. You have to use your imagination to fill in the blanks."
    • Because the "perfect" details were erased, the AI was forced to hallucinate (or rather, infer) the missing details based on what it already knows about how light works. It learned to "dream" the missing highlights and shadows back into existence.

4. The Result: A Magic Video Generator

By combining the Translator (LogC3) and the Amnesia Training (intentional blurring), they created LumiVid.

  • How it works: You feed it a normal, low-quality video.
  • What happens: The AI uses its "translator" to understand the data, then uses its "imagination" (trained by the blurring) to invent the missing bright suns and deep shadows.
  • The Output: It spits out a video that looks like it was shot in high-end HDR, with realistic lighting, glowing reflections, and deep shadows, all while keeping the movement smooth (no flickering).

Why is this a big deal?

  • It's Cheap: They didn't build a new AI. They just tweaked an existing one with a tiny bit of extra code (less than 1% of the model's size).
  • It's Fast: It took them only about 8 hours on a single computer to train, compared to the weeks or months usually required.
  • It Works: It creates videos that look professional, fixing the "blown out" skies and "crushed" shadows that plague normal videos.

In short: They didn't build a new brain; they just taught an old brain a new dialect and then challenged it to solve puzzles it couldn't easily copy. The result is a video generator that can turn a flat, dull video into a cinematic masterpiece.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →