Towards Real-world Human Behavior Simulation: Benchmarking Large Language Models on Long-horizon, Cross-scenario, Heterogeneous Behavior Traces
This paper introduces OmniBehavior, the first benchmark constructed entirely from real-world data to evaluate Large Language Models on long-horizon, cross-scenario human behavior, revealing that current models suffer from tunnel vision and a fundamental structural bias toward homogenized, overly positive personas that fail to capture authentic individual differences.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to build a perfect digital twin of a human being. You want an AI that can sit in a room, look at a video, browse a store, or chat with customer service, and act exactly like a real person would. This is what researchers call a "User Simulator."
For a long time, scientists thought Large Language Models (LLMs) were ready for this job. But a new paper called OmniBehavior says, "Not so fast."
Here is the story of the paper, explained simply with some analogies.
1. The Problem: The "Zoo Exhibit" vs. The "Real Jungle"
Imagine you are trying to learn how a lion behaves.
- Old Benchmarks (The Zoo): Previous tests only looked at lions in a cage. They watched the lion eat a piece of meat (one action) and then sleep (another action). They thought, "Okay, we know how lions work!"
- The Reality (The Jungle): In the real world, a lion's life is a messy, long story. It might hunt in the morning, get chased by a rival in the afternoon, sleep in a different spot at night, and interact with the whole pride. Its actions are connected across days and different places.
The paper argues that previous AI tests were like the Zoo. They only tested AI on isolated, short tasks (like "click this button" or "watch this video"). They missed the big picture: real human behavior is a long, messy, cross-scenario story.
2. The Solution: OmniBehavior (The "Real Jungle" Map)
The researchers created a new, massive dataset called OmniBehavior.
- Where did it come from? They didn't make up fake data. They took 3 months of real, anonymous logs from 200 real users on Kuaishou (a giant Chinese video app, similar to TikTok).
- What's in it? It tracks a user's entire digital life: watching videos, shopping, chatting with customer service, searching for things, and watching live streams.
- The Scale: It's like watching a movie that is 3 months long, where the character jumps between 5 different "rooms" (scenarios) and does 22 different types of things.
3. The Big Reveal: The AI is a "Too-Polite, Average Person"
The researchers tested the world's smartest AIs (like GPT-5, Claude, and others) using this new "Real Jungle" map. The results were shocking.
The AI failed to act like a real human in three weird ways:
A. The "Hyper-Active" Bias (The Over-Eager Fan)
- Real Humans: We are lazy sometimes. We scroll past 99 videos and only like one. We ignore most ads. We are often bored or indifferent.
- The AI: The AI acts like a super-fan. It thinks every video is amazing and wants to "Like," "Share," and "Buy" everything. It is too active. It can't simulate the feeling of "meh" or "I don't care."
B. The "Utopian" Bias (The Polite Robot)
- Real Humans: When things go wrong (like a package is late), real people get angry. They use harsh words, complain, and get frustrated.
- The AI: The AI is stuck in "Customer Service Mode." Even when the user is angry, the AI responds politely. It says, "Oh dear, that's unfortunate," instead of "Where is my package? I'm furious!" It's like a robot trying to be a diplomat in a bar fight. It's too nice.
C. The "Average Person" Bias (The Clone)
- Real Humans: We are all different. One person loves spicy food; another hates it. One is a night owl; another is an early bird.
- The AI: The AI tries to be everyone at once. It creates a "perfect average" human. It forgets the weird, specific, "long-tail" quirks that make you, you. If you ask the AI to simulate 100 different people, they all end up sounding like the same polite, enthusiastic middle-aged person.
4. Why Does This Matter?
You might ask, "So what if the AI is too nice? Isn't that good?"
No. If you are building a system to predict who will stop using an app (churn), or who is angry and might leave a bad review, you need an AI that can simulate anger, boredom, and indifference.
- If your AI is always happy and active, it will think everyone loves your product.
- In the real world, that's a dangerous lie.
The Takeaway
This paper is a wake-up call. It says: "Stop testing AI in a vacuum."
To build AI that truly understands humans, we need to stop looking at isolated actions and start looking at the whole, messy, long-term story of a person's life. And right now, our AI is still too polite, too active, and too "average" to be a real human twin. It's a mirror that only reflects our best ideals, not our messy reality.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.