HALLELUAI: A Hallucination-Aware AI System for Ultra-Realistic Image-to-Video Generation at Scale
HALLELUAI is an end-to-end system that ensures ultra-realistic, production-grade image-to-video generation at scale by integrating a hallucination-aware moderation module with an agentic regeneration framework to iteratively refine outputs and guarantee brand safety and visual fidelity.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a world where you can snap a single photo of a beach and, with a magic command, turn it into a moving movie where the waves crash and the seagulls fly. This is the promise of "Image-to-Video" generation, a branch of artificial intelligence that is rapidly changing how we make ads, tell stories, and show off products. But just like a child learning to draw, these AI artists sometimes get carried away. They might blur the picture, make the waves move in a weird, jittery way, or—here's the tricky part—suddenly invent a palm tree that wasn't in the original photo. In industries like travel or real estate, where showing the exact truth matters, these "hallucinations" (the AI making things up) are a big problem. We need a way to check the AI's work before we show it to the world, ensuring the video looks real, moves smoothly, and doesn't add any fake elements.
Enter HALLELUAI, a new system designed to be the ultimate "quality control inspector" for these AI-generated movies. Think of HALLELUAI not just as a judge, but as a tireless, super-smart editor who works in a loop. When an AI tries to turn a photo into a video, HALLELUAI watches the result closely. If the video is blurry, if the camera shakes too much, or if the AI accidentally adds a fake building to a hotel photo, HALLELUAI doesn't just say "No." Instead, it acts like a coach, telling the AI exactly what went wrong and then giving it a second (or third) chance to fix it. It tweaks the instructions, changes the camera angle, or even swaps out the AI model trying to do the job, repeating this process until the video is perfect. The researchers found that by using this "try, check, fix, try again" approach, they could produce ultra-realistic videos that creative experts loved, with the system agreeing with human experts on whether a video was good or bad about 87% of the time.
The Problem: When AI Gets Too Creative
Imagine you ask a friend to describe a photo of your dog. If they say, "Your dog is running," that's fine. But if they suddenly say, "Your dog is running, but now he has wings and is wearing a hat," they've started hallucinating. In the world of AI video, this is a disaster. If a travel company uses AI to show a hotel room, they can't have the AI invent a swimming pool that doesn't exist or make the furniture float.
Current tools for checking AI videos are like using a ruler to measure the taste of soup. They can tell you if a whole bunch of videos look "real" on average, but they can't look at a single video and say, "Hey, this specific one has a fake tree in it," or "The camera is shaking too much." They miss the small, annoying mistakes that ruin the experience for a viewer.
The Solution: The HALLELUAI Loop
The authors built HALLELUAI to solve this by creating a closed loop of generation and correction. Here is how it works, step-by-step:
1. The First Draft
The system takes a starting photo and a creative idea (like "make the camera zoom in slowly") and asks an AI to generate a video.
2. The Strict Inspector (Moderation Module)
Before the video is allowed to go out, HALLELUAI's "Moderation Module" puts it under a microscope. It checks three main things:
- Frame Quality: Is the picture blurry? Is it too dark or too bright? Is there weird static noise? It compares every single frame to the original photo to make sure the lighting and sharpness haven't drifted.
- Motion Smoothness: Is the camera moving like a professional filmmaker, or is it jittering like a shaky hand? It checks if the movement matches the instructions (e.g., if you asked for a "pan left," did it actually pan left?).
- The Hallucination Hunt: This is the most critical part. The system uses a powerful AI brain to look for things that shouldn't be there. Did an object disappear? Did two objects merge into a weird blob? Did a new object suddenly pop into the scene? It specifically looks for "New Structure Hallucinations" (fake objects appearing) and "Object Hallucinations" (real objects changing shape or vanishing).
3. The Coach (Agentic Regeneration Module)
If the video fails the inspection, HALLELUAI doesn't just throw it in the trash. It acts like a smart coach. It looks at the specific mistakes and decides on a fix.
- If the motion was too shaky, it tells the AI to "stabilize the camera."
- If the lighting was wrong, it says "adjust the exposure."
- If the AI kept inventing fake trees, it might switch to a different AI model or change the prompt to be more strict about keeping the original image's details.
- If the object was blurry, it might ask the AI to try again with a different "seed" (a random starting point for the generation).
The system then generates a new version of the video and checks it again. It keeps doing this—Generate, Check, Fix, Generate—until the video passes all the tests or until it hits a limit on how many times it can try.
What They Found
The team tested HALLELUAI with real human experts to see if it could do the job.
- The Agreement: When they compared HALLELUAI's decisions to those of human creative experts, the system agreed with the humans 86.9% of the time. This is a strong signal that the machine is learning to think like a human editor.
- The Precision: When the system said a video was "Good" (a PASS), it was right 88% of the time in the initial tests.
- The Real-World Test: In a "Pseudo Production" test, where the system was allowed to fix the videos before showing them to humans, the quality skyrocketed. Out of 1,158 videos the system approved, 1,126 were also approved by human experts. That's a 97% success rate.
The researchers also compared their system to other existing tools. They found that standard video quality checkers (like DOVER and COVER) couldn't tell the difference between a good AI video and a bad one because they weren't designed to spot these specific AI mistakes. Even powerful AI chatbots, when asked to just "look at the video" without special instructions, failed to catch the subtle errors. HALLELUAI worked because it was built specifically to look for these exact problems and had a plan to fix them.
Why This Matters
This isn't just about making pretty videos; it's about trust. In fields like real estate or travel, if an AI shows a house with a pool that doesn't exist, it's not just a funny mistake—it's a legal and reputational risk. HALLELUAI shows that we can scale up AI video production without losing control. By combining a strict inspector with a helpful coach, the system ensures that the final product is not only ultra-realistic but also faithful to the original image. It turns the chaotic, "guess-and-check" nature of AI generation into a reliable, industrial process that can handle thousands of videos while keeping the quality high. The authors suggest this framework is a major step toward making AI-generated video safe and usable for big companies, bridging the gap between cool technology and real-world reliability.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.