Controlla: Learning Controllability via Graph-Constrained Latent Geometry
This paper introduces Controlla, a modular framework that models controllability as structured latent geometry by aligning learned identity and attribute factors with graph priors via graph-constrained optimal transport, thereby improving identity preservation and cross-modal consistency in multimodal generation, which is validated on a new benchmark called AffectHuman-43K.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Problem: The "Drifting Identity"
Imagine you are an artist trying to paint a portrait of a specific friend. You want to change their expression from "neutral" to "excited" and "laughing," based on a text description and a recording of their laughter.
With current AI tools, you might get a result that looks like your friend, but with a weird twist: their nose might look slightly different, their hair might change color, or the "excitement" might look like a completely different person. The AI is good at making images, but it struggles to change one thing (the emotion) without accidentally changing everything else (the identity). It's like trying to change the flavor of a soup without the spoon accidentally swapping out the bowl.
The Solution: Controlla (The "Structured Map")
The authors propose a new system called Controlla. Instead of just telling the AI "make it excited" and hoping for the best, Controlla teaches the AI to understand the structure of emotions and identity before it starts drawing.
Think of the AI's internal "brain" (its latent space) as a giant, messy warehouse.
- Old Way: You just shout "Excitement!" into the warehouse. The AI grabs whatever "excitement" items it finds nearby, but it might accidentally grab a "stranger's face" or "blue hair" along with it.
- Controlla's Way: Controlla builds a structured map (a graph) inside that warehouse.
- The Identity Zone: It creates a safe, locked room for your friend's face. Once your friend's identity is placed there, the map ensures it stays exactly where it is, no matter what happens outside.
- The Emotion Pathway: It builds a specific, paved road for emotions. This road connects "Neutral" to "Happy" to "Excited" in a logical, smooth line.
How It Works: The "Train on Tracks" Analogy
The paper uses a concept called Graph-Constrained Optimal Transport. Let's break that down:
Imagine the AI's generation process is a train.
- The Tracks (The Graph): Before the train moves, Controlla lays down tracks that represent the rules of the world. The tracks for "Emotion" are curved and connected, showing that you can't jump instantly from "Sad" to "Manic" without passing through "Anxious." The tracks for "Identity" are a solid, unmovable platform.
- The Train (The Generation): When you ask for a change, the train (the AI) is forced to stay on the tracks. It cannot wander off into the "stranger's face" zone because the tracks don't go there. It can only move smoothly along the emotion pathway.
- The Result: The train arrives at "Excitement," but because it stayed on the tracks, the "Identity Platform" underneath it never moved. Your friend looks exactly like your friend, just with a new expression.
The New Playground: AffectHuman-43K
To prove this works, the authors built a new testing ground called AffectHuman-43K.
- The Challenge: Most old tests mixed up the person's face with their voice or text, making it hard to tell if the AI was actually preserving the face or just guessing.
- The Fix: This new dataset is like a strict exam. You give the AI a photo of a person (the "Reference") and separate instructions (a text saying "laugh" and an audio clip of laughter). The AI must keep the photo's face exactly the same while changing the expression to match the audio/text.
- The "Leakage" Guard: The exam is designed so that the AI can't cheat by memorizing the answers. The "students" (AI models) are tested on people they have never seen before.
The Results: Smoother Rides
When they tested Controlla against other top AI models:
- Better Identity: The faces stayed recognizable. No more "nose drift."
- Smoother Transitions: If you asked the AI to slowly change an expression from neutral to happy, Controlla did it in a smooth, natural flow. Other models often made jerky, unnatural jumps.
- Following the Map: The AI followed the "emotion road" much better, meaning the changes felt logical and consistent, not random.
Summary
In short, Controlla stops the AI from guessing how to change an image. Instead, it gives the AI a map and tracks. It says, "Here is the person's face (stay here), and here is the road to the new emotion (follow this path)." By forcing the AI to follow this structured geometry, it creates changes that are controlled, smooth, and true to the original person.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.