← Latest papers
🤖 AI

Setoka: A Benchmark for Hierarchical User Understanding in Personalized Agents over Heterogeneous Data

This paper introduces Setoka, a novel benchmark grounded in cognitive and personality psychology that evaluates personalized agents' ability to perform hierarchical user understanding—from semantic and episodic memory to behavior patterns and personality traits—across heterogeneous data, revealing that current systems struggle significantly with tasks requiring the integration and abstraction of fragmented long-term information beyond simple fact retrieval.

Original authors: Lingyang Zeng, Guangze Chen, Kaichen Yu, Zhicheng Pan, Siyang Weng, Zirui Hu, Xiangyun Du, Hailin He, Rong Zhang, Chengcheng Yang, Kai Huang, Xuan Zhou

Published 2026-07-30
📖 6 min read🧠 Deep dive

Original authors: Lingyang Zeng, Guangze Chen, Kaichen Yu, Zhicheng Pan, Siyang Weng, Zirui Hu, Xiangyun Du, Hailin He, Rong Zhang, Chengcheng Yang, Kai Huang, Xuan Zhou

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to get to know a new friend. You could just ask them, "What is your phone number?" or "What time is your meeting?" and write down the answer. That's easy; it's just grabbing a fact. But real friendship is deeper. It's noticing that your friend always laughs at the same type of jokes, or that they seem to get grumpy when it rains, or that they are the kind of person who plans their whole week in advance. These aren't facts you can just look up; they are patterns you have to piece together from many different conversations, texts, and moments over time.

In the world of artificial intelligence, we are building "agents" that act like personal assistants. Right now, these AI helpers are getting pretty good at the easy stuff: remembering your phone number or finding a meeting time in your calendar. But scientists are worried that these agents are terrible at the "deep" stuff. They struggle to understand who you really are—your habits, your personality, and your long-term behavior. To fix this, researchers need a way to test if an AI can actually "get" a human, not just read a file. This is where a new study comes in, trying to build a better test to see if AI can move from being a simple note-taker to becoming a true understanding companion.


The Problem: AI Has Amnesia for "Who You Are"

Think of your digital life as a giant, messy attic filled with different kinds of boxes. Some boxes are labeled "Messages," some are "Calendar Events," and others are "Social Connections." Right now, most AI assistants are like a librarian who can only pull out one specific box when you ask for it. If you ask, "What is Alice's phone number?" the AI finds the "Contacts" box and reads the number. Easy peasy.

But what if you ask, "Am I an extroverted person?" or "How often do I usually hang out with friends on weekends?" The AI can't just open one box. It has to look at your calendar to see group events, check your messages to see how active you are in chats, and scan your social connections to see how many friends you have. It has to connect the dots across different types of data to figure out a pattern. The problem is that current AI systems are really bad at this. They are great at finding facts but terrible at understanding the story behind the facts.

Enter Setoka: The Ultimate "Get to Know You" Test

To solve this, a team of researchers created a new benchmark called Setoka. You can think of Setoka as a giant, super-organized simulation lab. Instead of using real people's private data (which would be a privacy nightmare), the researchers built a "psychology factory."

Here is how their factory works:

  1. The Blueprint: First, they use real psychological science to create 10 fake people. They don't just guess what these people are like; they use a special math model to make sure the personalities make sense. For example, if a fake person is very "outgoing," the math ensures they aren't also secretly "shy" in a way that real humans rarely are.
  2. The Behavior: Next, they translate these personalities into daily habits. If a person is "outgoing," the system generates a realistic schedule where they talk to friends every other day, play multiplayer games daily, and watch sitcoms a few times a week.
  3. The Messy Data: Finally, they turn those habits into a massive pile of digital clutter. They create thousands of fake text messages, calendar entries, and social network logs that span two months. This is the "heterogeneous data"—a mix of different file types that an AI would have to sift through.

Once the factory is running, they ask the AI 1,426 questions about these fake people. The questions are divided into four levels of difficulty, like a video game with increasing boss battles:

  • Level 1 (Semantic Memory): "What is Alice's phone number?" (Just find the fact).
  • Level 2 (Episodic Memory): "What was the user doing on the evening of April 20?" (Combine a calendar entry, a text message, and a location log to reconstruct one specific event).
  • Level 3 (Behavior Pattern): "How many times per week does the user socialize?" (Look at weeks of data and count the total).
  • Level 4 (Personality Trait): "How extroverted is the user?" (Look at all the social data, the work habits, and the leisure activities to guess a deep personality trait).

What They Found: The "Fact-Finder" vs. The "Mind-Reader"

The researchers tested 3 different AI brains (ranging from a powerful cloud model to a smaller one that could run on a phone) paired with 5 different memory systems (some that organize data like a graph, some like a simple list).

The results were a bit of a wake-up call for the AI world:

  • The Easy Stuff is Solved: When the question was just about finding a fact (Level 1), the AI did great. In fact, sometimes the AI didn't even need a special memory system; it could just look up the raw data like a database.
  • The Middle Ground is Hard: As soon as the AI had to combine a few pieces of information to remember a specific event (Level 2), the scores dropped.
  • The Deep Stuff is Broken: When the AI had to find patterns or guess personality traits (Levels 3 and 4), the performance crashed. The best systems only got about 24% of the personality questions right.

The study suggests that simply making AI "remember" more facts isn't enough. The current systems are like students who can memorize a textbook perfectly but fail the essay question because they can't connect the ideas. The researchers found that to understand a user, an AI needs to do more than just retrieve data; it needs to link different sources, aggregate them over time, and generalize to find hidden traits.

The Takeaway

Setoka shows us that we are still far from having AI that truly "knows" us. While our digital assistants are getting better at remembering our phone numbers and meeting times, they are still struggling to understand our habits and personalities. The study suggests that the next big breakthrough won't come from just giving AI more memory, but from teaching it how to weave together the messy, scattered pieces of our digital lives into a coherent picture of who we really are. Until then, our AI friends might know what we did yesterday, but they still don't really know us.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →