HORIZON: A Benchmark for In-the-wild User Behaviour Modeling
The paper introduces HORIZON, a new large-scale, cross-domain benchmark built from 54 million users and 35 million items that reformulates user behavior modeling through novel tasks and evaluation metrics to address the limitations of existing narrow benchmarks and better reflect real-world deployment scenarios.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to predict what a person will buy next.
The Old Way (Current Benchmarks):
Most existing tests for recommendation systems are like playing a game of "Guess the Next Card" using a single deck of cards. If a user has been buying only cooking books, the test asks the AI: "Okay, they just bought a cookbook, what will they buy next?" The AI guesses another cookbook. It gets a high score because it's good at predicting the next item in a very narrow, short-term context.
But in real life, people aren't just one-dimensional. A person might buy a cookbook, then a tent, then a guitar, then a baby stroller. Their interests shift over years, not just minutes. The old tests fail to see the whole person; they only see the last few seconds of their life.
The New Solution: HORIZON
This paper introduces HORIZON, a new, massive "training ground" for AI that changes the rules of the game. Instead of a single deck of cards, HORIZON gives the AI a library containing 54 million people's entire shopping histories across 35 million different products (from toothpaste to telescopes).
Think of HORIZON as a time machine combined with a chameleon suit. It forces the AI to do three difficult things that old tests ignored:
- The "Time Travel" Test: Instead of guessing the next item immediately, the AI must predict what a user will buy years from now, based on what they bought in the past. It's like asking a weather forecaster to predict next winter's snow based on data from last summer, while the user's habits have completely changed.
- The "Stranger" Test: The AI is tested on people it has never seen before. In the old tests, the AI studied a user's history and then guessed their next move. In HORIZON, the AI meets a stranger, looks at their short history, and has to guess what they'll buy next, without ever having met them in training.
- The "Jigsaw" Test: The AI must connect dots across different worlds. If a user buys a camera, then a hiking boot, then a protein shake, the AI needs to understand that this person is an "outdoor photographer," not just a "camera buyer."
How They Tested It
The researchers didn't just look at how well the AI guessed the next item. They set up four different "exam rooms" to see how the AI handled stress:
- Room 1 (The Comfort Zone): Predicting the next item for a known user, right after the training data. (Easy mode).
- Room 2 (The Time Jump): Predicting for a known user, but years into the future. (Hard mode: Did their taste change?)
- Room 3 (The Stranger): Predicting for a new user, but in the same time period. (Hard mode: Can it generalize?)
- Room 4 (The Ultimate Boss): Predicting for a new user, years into the future. (Impossible mode: Can it handle both new people and new times?)
What They Found
The results were a bit of a wake-up call for the tech world:
- The "Smart" AI isn't that smart: The most advanced AI models (like BERT4Rec) were great at Room 1 (the Comfort Zone). But as soon as they entered Room 4 (The Ultimate Boss), their performance crashed. They were like a student who memorized the textbook but failed when asked to solve a real-world problem they hadn't seen before.
- LLMs (Large Language Models) are okay, but not magic: The researchers tried using giant AI chatbots (like Llama or Qwen) to act as "psychologists" for these users. They asked the AI to write down what the user might search for next. While these chatbots were decent at understanding the vibe of a user, they weren't great at actually picking the right product from a catalog of 35 million items. They were good at the "why," but bad at the "what."
- The "Long Tail" Problem: Most items in the real world are rare. The AI struggled to predict purchases for obscure items (like a specific type of telescope) because it had never seen them before. It relied too much on popular items.
The Big Takeaway
The paper argues that we are currently building AI that is too specialized and too short-sighted. It's like training a dog to fetch a ball in a living room, then expecting it to hunt a deer in a forest.
HORIZON is the forest. It's a new standard that forces researchers to build AI that can:
- Understand a person's life story, not just their last purchase.
- Adapt to new people and new times.
- Handle the messy, changing reality of the real world.
In short, HORIZON is saying: "Stop playing with toy datasets. If you want to build a recommendation system that works in the real world, you need to test it in the real world."
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.