← Latest papers
💬 NLP

PersonaTrail: Benchmarking Personalized Web Agents through Browsing Trails

This paper introduces PersonaTrail, a benchmark for evaluating personalized web agents' ability to infer user preferences from raw browsing histories, alongside PACMem, a framework that structures this history into factual and preference memories to significantly improve navigation performance.

Original authors: Seungbin Yang, Chaewoon Ki, Dohyun Lee, Jaegul Choo, ChaeHun Park

Published 2026-08-05
📖 4 min read☕ Coffee break read

Original authors: Seungbin Yang, Chaewoon Ki, Dohyun Lee, Jaegul Choo, ChaeHun Park

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a super-smart robot assistant that can browse the internet for you. You ask it to "find a movie," and it does. But in the real world, you rarely give perfect instructions. You might say, "Find me a movie like the ones I usually watch," or "Show me that recipe I looked at last Tuesday." To do this, the robot needs to remember your past habits and specific moments from your digital life. This is the world of web agents: software that acts like a human on websites, clicking buttons and reading pages to get things done. For a while, scientists tested these robots using very strict, perfect instructions, like a math problem with all the numbers filled in. But that doesn't feel like real life. Real life is messy, full of half-finished thoughts and memories of things we did days ago. The big question is: Can these robots actually learn you from your messy history, or do they just follow orders blindly?

This paper introduces a new way to test these robots, called PERSONATRAIL. The researchers built a giant playground where they created fake but realistic "users" with specific hobbies, jobs, and tastes. They then watched how these users browsed the web, creating a long trail of clicks, searches, and page visits. They gave the robots two tricky challenges. First, Preference Inference: The robot has to guess what you like based on a pattern. For example, if you always click on "Mystery" movies and "One-pot" recipes, the robot needs to figure out you love mysteries and easy cooking, even if you don't say it out loud. Second, Episodic Grounding: The robot has to find a specific memory. If you say, "Show me the curry recipe I searched for last Tuesday," the robot has to dig through its memory bank to find that exact moment and go back there.

The researchers found that most existing robots were terrible at this. They were like students who memorized a textbook but couldn't apply it to a new situation. When the robots tried to remember your past, they often threw away the important details, like what you clicked on, and only kept the general idea of how you clicked. To fix this, the team built a new memory system called PACMEM. Think of PACMEM as a super-organized librarian for the robot. Instead of just dumping a pile of papers on the desk, this librarian sorts them into two special bins. One bin, the Factual Memory, is a diary of exactly what happened on specific days (e.g., "On May 24, Sarah watched Bird Box"). The other bin, the Preference Memory, is a summary of your habits (e.g., "Sarah always picks Mystery movies and ratings above 6.0").

When the robot gets a question, PACMEM doesn't just dump all the history on it. It quickly checks the diary bin if the question is about a specific day, or the habit bin if the question is about what you like. The results showed that robots using this two-bin system were much better at solving these personalized puzzles than the old ones. They didn't just guess; they actually used the right clues to find the right movie or recipe. The paper suggests that for robots to truly help us in the real world, they need to stop treating our browsing history like a messy pile of trash and start organizing it into clear stories and habits. While this is a big step forward, the researchers note that real life is even more complicated than their test, with habits that change and questions that are impossible to answer, but this new system gives us a solid foundation to build smarter, more personal helpers.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →