CogPortrait: Fine-Grained Eye-Region Control in Portrait Animation via Hierarchical Agent Planning
CogPortrait is a two-stage framework that leverages hierarchical Multimodal Large Language Model agents to translate high-level labels into precise facial keypoints, enabling fine-grained control of eye-region dynamics and beyond-emotion states in portrait animation while maintaining high visual quality and identity consistency.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you want to make a static photo of a person come to life, talking and moving. Most current methods are like giving a director a very vague script: "Make him look angry" or "Make her look happy." The result is okay, but the eyes often look dead or the movements are too smooth and robotic. On the other hand, some methods let you control every tiny muscle movement, but that's like asking the director to manually move every single eyelid and eyebrow pixel by pixel—it's incredibly hard work and requires a lot of technical skill.
CogPortrait is a new system that tries to find the perfect middle ground. It lets you give a simple, high-level instruction (like "he's thinking hard" or "she's getting sleepy") and automatically figures out exactly how the eyes and head should move to make it look real, without you having to do the heavy lifting.
Here is how it works, broken down into two main acts:
Act 1: The "Brain" (The Agent Team)
Before the video is made, the system uses a team of three AI "agents" (think of them as a specialized production crew) to translate your simple idea into a detailed movement plan.
- The Planner (The Director): This agent takes your simple label (e.g., "Cognitive Effort") and breaks it down into a timeline. It decides, "First, the person looks up for a second, then they squint slightly, then they look to the left." It creates a storyboard of events.
- The Composer (The Librarian): This agent goes into a library of real human behavior recordings. It finds real examples of people "squinting while thinking" or "blinking slowly when tired." It stitches these real-life snippets together to create a realistic movement sequence, ensuring the eyes don't move in a weird, robotic way.
- The Critic (The Quality Control Inspector): This agent double-checks the plan. It asks, "Does this actually look like someone thinking? Is it physically possible for a human to blink that fast?" If the plan is weird, it sends it back to the Composer or Planner to fix it.
The result of this team is a precise set of coordinates (keypoints) that tell the computer exactly where the eyelids, eyebrows, and pupils should be at every single moment.
Act 2: The "Body" (The Video Generator)
Once the movement plan is ready, the second stage takes over. This is the part that actually paints the video.
- The Painter: It takes your original photo, the audio (so the lips move with the voice), and the detailed movement plan from Act 1.
- The Smart Brush (Dynamic Guidance): Usually, when AI paints, it might get the colors of the face and the background mixed up, or it might make the eyes look blurry. CogPortrait uses a "smart brush" that focuses extra attention on the eyes. It says, "Make the background stable and natural, but apply extra precision to the eyes and eyebrows so the blinking and looking feel real."
- The Safety Net (KTO Refinement): Sometimes, the AI struggles with tricky situations, like a person turning their head almost all the way to the side or raising just one eyebrow. To fix this, the system was trained on a special "hard mode" dataset. It learned from its own mistakes on these difficult cases to ensure the person's face doesn't distort and the identity stays consistent, even in weird angles.
Why is this a big deal?
The paper introduces a new test called EMH (Emotions and Beyond-Emotions) to see how well these systems work. It tests not just basic emotions like "happy" or "sad," but also subtle states like:
- Cognitive Effort: That look of deep concentration.
- Drowsiness: Heavy eyelids and slow nods.
- Evasive Response: Looking away quickly.
The results show that CogPortrait is much better at controlling the eyes than previous methods. It can make a character look "tired" or "thinking" with high accuracy, while keeping the person looking like themselves and keeping the video quality high. It bridges the gap between "easy to use" (just type a label) and "highly realistic" (precise eye movements).
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.