AHOY! Animatable Humans under Occlusion from YouTube Videos with Gaussian Splatting and Video Diffusion Priors
AHOY is a novel method that reconstructs complete, animatable 3D Gaussian avatars from heavily occluded, in-the-wild monocular videos by leveraging identity-finetuned diffusion models to hallucinate unobserved body regions and employing a specialized architecture to resolve multi-view inconsistencies.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you want to create a digital twin of a person for a video game or a movie. Usually, you need a special studio with dozens of cameras, bright lights, and the person standing perfectly still so nothing blocks the view.
AHOY is a new method that says, "No problem! We can do this with just a regular YouTube video, even if the person is hiding behind a table, a tree, or another person."
Here is how it works, broken down into simple steps with some creative analogies:
The Big Problem: The "Blind Spot"
Most 3D reconstruction tools are like a photographer who only takes pictures when the subject is standing in an empty room. If the person is sitting on a couch, the photographer can't see their legs. If they are holding a coffee cup, the photographer can't see their hand.
In the real world (like YouTube videos), people are constantly blocked by things. Traditional tools get confused and leave big holes in the 3D model, or they just give up.
The Solution: AHOY's "Magic Imagination" Pipeline
AHOY solves this by using a mix of 3D scanning and AI imagination (specifically, a Video Diffusion model). Think of it as a four-step cooking recipe:
Step 1: The Rough Draft (The "Coarse Avatar")
First, AHOY looks at the video and builds a very rough, blocky 3D model of the person using only the parts it can see.
- Analogy: Imagine a sculptor trying to make a statue of a person sitting on a chair. They can only carve the visible parts (the head, the shoulders). The legs and the back are just empty space because the chair is in the way. This is the "Coarse Avatar."
Step 2: The "Hallucination" (The AI Artist)
This is the secret sauce. AHOY takes that rough, blocky statue and asks a super-smart AI video generator (trained on the specific person's face and clothes) to "imagine" what the missing parts look like.
- Analogy: The sculptor hands the rough statue to a magical painter. The painter looks at the empty space where the legs should be and says, "Based on the style of the shirt and the person's identity, I'm going to paint the legs, the back, and the hidden hands."
- The Trick: The AI doesn't just guess randomly; it uses "Diffusion Inversion." It takes the rough 3D model, turns it into a video, and asks the AI to "re-imagine" that video but fill in the missing details with high-quality, realistic textures. It's like asking an artist to redraw a sketch, but this time, they fill in all the blank spots with perfect detail.
Step 3: The "Two-Pose" Dance (Fixing the Glitches)
Here is the tricky part. The AI's "imagination" isn't perfect. If the AI generates a video of the person turning around, the left side might look slightly different from the right side because the AI generated them separately. It's like a team of painters working on different parts of a mural without talking to each other—the seams might not match.
AHOY solves this with a clever "Two-Pose" system:
- The Map Pose: This is the "blueprint." It stays the same for the whole movement. It tells the system, "This is what the person looks like when their arms are up."
- The LBS Pose: This is the "movement." It changes every single frame to match the video.
- Analogy: Think of a puppet show. The Map Pose is the puppet's body shape (which stays consistent). The LBS Pose is the puppeteer's hands moving the strings. If the AI's video looks a little wobbly or inconsistent, the "puppeteer" (the LBS pose) adjusts the strings frame-by-frame to smooth out the glitches, while the "body shape" (the Map) stays true to the person's identity.
Step 4: The Face Guard (Keeping Identity)
AI video generators are great at bodies but terrible at faces—they often change the person's identity (making them look like a different person).
- Analogy: AHOY puts a "Face Guard" on the process. It separates the head from the body. It uses a special, stricter AI just for the face to make sure the digital twin looks exactly like the real person, while letting the "imagination" AI handle the body.
The Result
At the end of the process, you have a complete, 3D, animatable human.
- You can take this digital person and make them dance, sit, or jump.
- You can put them into a 3D scene (like a living room) that was also filmed with a phone.
- Even though the original video had the person hiding behind a sofa, the final 3D model has the sofa's shape filled in perfectly, and the person can be moved anywhere.
Why This Matters
Before AHOY, if you wanted a 3D model of a person, you needed a Hollywood studio. Now, you can take a random video from YouTube where someone is cooking, dancing, or playing with their dog (and getting blocked by furniture), and turn it into a high-quality, moving 3D character. It unlocks the entire internet as a source for 3D characters!
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.