ReGenHuman: Re-Generating Human Appearances for Realistic Full-Body Video Anonymization
The paper introduces ReGenHuman, a novel full-body video anonymization pipeline that synthesizes realistic and temporally consistent human appearances from identity-free structural cues using a "regenerate, don't edit" paradigm, thereby achieving an optimal trade-off between privacy, visual quality, and downstream task utility.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a home video of a busy hospital hallway or a crowded street. You want to share this video to help train AI systems (like self-driving cars or medical robots), but you can't because the people in the video are recognizable. Their faces, their clothes, the way they walk, and even the text on their shirts give them away.
Currently, the way people handle this is like taking a permanent marker and scribbling over the people's faces or blurring the whole screen. It protects their identity, but it ruins the video. The AI can't learn anything from a blurry mess, and the video looks terrible.
Enter "ReGenHuman."
Think of ReGenHuman not as an editor with a marker, but as a magical, super-fast painter who watches the video and then re-paints the entire scene from scratch, replacing the real people with brand-new, fake ones.
Here is how it works, broken down into simple steps:
1. The "Ghost Blueprint" (The Input)
Instead of showing the painter the real people, the system first strips the video down to a "ghost blueprint." It extracts three things that describe where people are and what they are doing, but contain zero information about who they are:
- The Skeleton: A stick-figure drawing of the person's pose (arms up, walking, sitting).
- The Outline: A silhouette showing exactly where the person's body is in the room.
- The Depth Map: A 3D map showing how far away things are, giving the scene shape without showing textures or colors.
- The Description: A text caption describing the scene (e.g., "Three people are standing on a stage giving an award").
Crucially, the painter never sees the original faces, clothes, or skin. It only sees these anonymous blueprints.
2. The "Re-Painting" (The Generation)
The system uses a powerful AI (a video diffusion model) to look at these blueprints and generate a completely new video.
- It paints new people who fit perfectly into the stick-figure skeletons.
- It paints new clothes and hair that look realistic.
- It keeps the background and the movement exactly the same as the original video.
Because the painter never saw the original person, it is mathematically impossible for the original person to appear in the new video. It's like if you gave an artist a photo of a tree and asked them to paint a new tree based only on a wireframe drawing; they could never accidentally paint the exact same tree from the photo.
3. The Result: A "Safe" Video
The final video looks incredibly realistic. The new people look like real humans, they move smoothly (no jittery flickering between frames), and they interact with the environment naturally.
- Privacy: The original people are gone. You can't recognize them by their face, their walk, or their clothes.
- Quality: The video looks high-definition and natural, not blurry or pixelated.
- Utility: Because the actions and movements are preserved, AI systems can still "watch" this video and learn how people move, how doctors interact with patients, or how crowds behave.
Why is this a big deal?
The paper compares this to old methods:
- Old Way (Blurring/Redacting): Like putting a black bar over a face. It's safe, but the video is useless for learning.
- Old Generative Way (Inpainting): Like trying to paint over a face on a photo. It often looks weird, the person might look different in every frame (flickering), and sometimes the original face still leaks through.
- ReGenHuman: It's the first method that is safe by design (because it never sees the original) and high quality (because it generates a whole new, realistic video).
The authors tested this on thousands of videos, including medical footage, and found that their method is the best at balancing privacy (keeping people anonymous) with utility (keeping the video useful for AI). They even showed that AI models can still answer questions about the video (like "What is the doctor doing?") just as well as they could with the original, un-anonymized video.
In short: ReGenHuman is a tool that takes a video of real people, erases them completely, and paints in new, fake people who act exactly the same way, allowing us to share and study video data without ever risking anyone's privacy.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.