Omni-Persona: Systematic Benchmarking and Improving Omnimodal Personalization
This paper introduces Omni-Persona, the first comprehensive benchmark for omnimodal personalization that unifies text, image, and audio evaluation with a novel Calibrated Accuracy metric, revealing critical gaps in current models' grounding abilities and the limitations of existing post-training methods like SFT and RLVR.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a super-smart digital assistant that can see photos, hear voices, and read text all at once. This is the world of "omnimodal" AI, where a single brain tries to understand the world through eyes, ears, and words simultaneously. But there's a tricky part: making this assistant truly personal. You want it to recognize your face in a crowd, know your voice from a stranger's, and remember your favorite stories. However, real life is messy. Sometimes you ask, "Who is that?" and the person isn't even there. A truly smart assistant shouldn't just guess; it should know when to say, "I don't know."
For a long time, researchers have tested these assistants mostly on pictures and words, ignoring the voice part. They also mostly tested them in perfect, clean rooms where the answer is guaranteed to be present. This paper introduces a new, tougher test called Omni-Persona. Think of it as a "stress test" for personal AI. Instead of just asking, "Did it get the right answer?" (which is like checking if a student memorized the textbook), this test asks, "Did it know when not to answer?" It checks if the AI can handle the noise of the real world, where the person you're asking about might be missing from the memory bank entirely.
The researchers built a massive playground with about 750 different scenarios to see how well current AI models handle this. They created a system called the Persona Modality Graph, which is like a giant web of connections. Each person in the AI's memory is a node with three strands: a photo, a voice clip, and a text bio. When you ask a question, the AI has to find the right strand in the web and pull the correct thread. The test includes four main types of challenges: matching a face to a face, a voice to a voice, a text clue to a text bio, and the hardest one—using a text clue to find a person based on their voice or photo. Crucially, about half of these questions are "trick questions" where the person you are asking about isn't in the memory at all. The goal is to see if the AI can confidently say, "I can't decide," instead of making up a story.
The team put several famous AI models through this gauntlet, including some from big tech companies and some open-source ones. They found some surprising things. First, the AI models are much better at recognizing faces than voices. It's like they have sharp eyes but fuzzy ears; they can spot a face easily but often get confused by a voice, even when the voice is the only clue. Second, simply making the AI bigger (adding more "brain power" or parameters) doesn't always make it smarter at being personal. In fact, some of the largest models were the worst at knowing when to stay silent, often hallucinating (making things up) when the answer wasn't there.
The paper also tested two different ways to train these models to get better. The first method, called SFT (Supervised Fine-Tuning), is like giving the AI a massive stack of flashcards with the right answers written on them. The researchers tried giving it 1,000 cards and then 10,000 cards, but the AI didn't get much better at handling the tricky "missing person" questions. It seems that just memorizing more examples doesn't teach the AI how to handle uncertainty.
The second method, RLVR (Reinforcement Learning with Verifiable Rewards), is more like a video game. The AI tries to answer, and if it gets it right, it gets points. If it guesses wrong, it loses points. The researchers found that this method made the AI much better at finding the right answers when the person was present. However, it had a side effect: the AI became too eager to play the game. It started guessing even when it didn't know the answer, just to get points. This made it worse at the "calibrated accuracy" score—a fancy way of saying "being right and knowing when to stop." The paper suggests that while this game-like training helps the AI find things, it needs a better rulebook to teach it when to say, "I don't know."
In the end, the authors conclude that we are still a long way from a perfect personal AI. Current models are great at recognizing things when they are right in front of them, but they struggle with the messy reality of missing information. They are too confident when they should be humble. The paper doesn't offer a magic fix, but it provides a very clear map of where the AI is failing, showing that the next step isn't just making bigger models, but teaching them the art of knowing what they don't know.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.