← Latest papers
🤖 AI

VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactions

This paper introduces VitaBench 2.0, a new benchmark designed to evaluate the ability of large language model agents to perform personalized and proactive decision-making in long-term, fragmented user interactions, revealing significant gaps between current model capabilities and real-world requirements.

Original authors: Yuxin Chen, Yi Zhang, Zhengzhou Cai, Yaorui Shi, Zhiyuan Yao, Chenhang Cui, Jingnan Zheng, Yaqi Huo, Xi Su, Qi Gu, Xunliang Cai, Xiang Wang, An Zhang, Tat-Seng Chua

Published 2026-05-27
📖 5 min read🧠 Deep dive

Original authors: Yuxin Chen, Yi Zhang, Zhengzhou Cai, Yaorui Shi, Zhiyuan Yao, Chenhang Cui, Jingnan Zheng, Yaqi Huo, Xi Su, Qi Gu, Xunliang Cai, Xiang Wang, An Zhang, Tat-Seng Chua

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a digital assistant that is incredibly smart at solving puzzles and using tools, but it has a terrible memory of who you are. It knows how to book a flight, but it forgets that you hate flying at night, or that you prefer window seats, or that you suddenly decided to stop eating spicy food last month.

This paper, VitaBench 2.0, is like a rigorous "final exam" designed to test exactly that problem: Can AI assistants actually get to know you over time and act like a thoughtful friend rather than a robotic stranger?

Here is a breakdown of the paper's key ideas using simple analogies:

1. The Problem: The "Amnesiac Butler"

Current AI agents are like butlers who are great at following specific instructions ("Bring me coffee") but terrible at reading the room. If you say, "I'm feeling down," a smart butler might know to bring a comfort item. But if you've been chatting with them for weeks, and you suddenly say, "I'm feeling down," a truly personalized butler would remember that you usually feel down on rainy Tuesdays and suggest an indoor hobby you love, rather than just asking, "What's wrong?"

Existing tests for AI mostly check if the robot can follow a recipe. VitaBench 2.0 checks if the robot can remember your taste in food, your changing moods, and your hidden habits over a long period.

2. The Exam: VitaBench 2.0

The researchers built a simulation of a real life. Imagine a video game where you play a user, and the AI is your assistant.

  • The Setup: They created 56 different "users" with complex lives (jobs, families, allergies, hobbies).
  • The Twist: The AI doesn't get a cheat sheet with the user's preferences. Instead, it has to figure out what the user likes by listening to their fragmented conversations and watching their behavior over time (like seeing them buy a salad today, but a burger yesterday).
  • The Drift: Just like real people, these simulated users change their minds. Maybe they decide to go on a diet, or they suddenly love a new type of music. The AI has to notice these changes and update its "mental model" of the user.

3. The Three Skills Tested

To pass this exam, the AI needs to master three specific skills:

  • Detective Work (Extraction): Finding the clues. If a user says, "I'm trying to cut back on sugar," the AI must realize this is a new rule for future orders, even if the user doesn't explicitly say "Remember this."
  • The Librarian (Memory): Keeping the clues organized. The AI needs a "memory bank" to store these preferences. The paper tested two types of librarians:
    • The Active Manager: An AI that decides what to keep, what to throw away, and how to summarize the user's life.
    • The Search Engine: An AI that just dumps all past chats into a database and searches for keywords when needed.
  • The Proactive Friend: Sometimes the user gives a vague order like, "Get me a coffee." A good assistant knows when it's missing information. If the user usually drinks strong coffee in the morning but weak coffee in the afternoon, the AI should ask, "What time is your meeting?" before ordering.

4. The Results: The "Reality Gap"

The researchers tested many of the world's smartest AI models (from companies like OpenAI, Google, DeepSeek, etc.) on this exam. The results were sobering:

  • The "Thinking" Trap: They found that just because an AI can "think" harder (reasoning step-by-step) doesn't mean it gets better at remembering you. In fact, some of the smartest reasoning models still failed to remember simple preferences like "I don't like spicy food."
  • Memory is Hard: Even with a memory system, most AIs struggled. They often forgot old preferences, got confused by new ones, or couldn't connect the dots between a past conversation and a current request.
  • The Proactive Gap: The AIs were very bad at realizing when they were missing information. Instead of asking, "Do you want this for lunch or dinner?" they often just guessed and got it wrong.

5. The Verdict

The paper concludes that while AI is getting amazing at solving math problems or writing code, it is still struggling to be a personalized companion.

Think of it like this: We have built a car with a powerful engine (the reasoning AI), but we haven't figured out how to install the GPS that knows your favorite routes and stops (the personalization). VitaBench 2.0 is the map that shows us exactly where the GPS is broken, so engineers can fix it.

In short: The paper says, "Our AI is smart, but it's not yours yet. It forgets your name, your likes, and your changes in mood. We built a test to prove this, and the results show we have a long way to go before AI can truly act like a personalized assistant."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →