← Latest papers
💬 NLP

Language Models Don't Know What You Want: Evaluating Personalization in Deep Research Needs Real Users

This paper introduces MyScholarQA, a personalized Deep Research tool that outperforms baselines on synthetic benchmarks but reveals critical, undetectable limitations through real-user interviews, arguing that genuine progress in personalization requires evaluation with actual users rather than relying solely on LLM judges.

Original authors: Nishant Balepur, Malachi Hamada, Varsha Kishore, Sergey Feldman, Amanpreet Singh, Pao Siangliulue, Joseph Chee Chang, Eunsol Choi, Jordan Lee Boyd-Graber, Aakanksha Naik

Published 2026-03-18
📖 4 min read☕ Coffee break read

Original authors: Nishant Balepur, Malachi Hamada, Varsha Kishore, Sergey Feldman, Amanpreet Singh, Pao Siangliulue, Joseph Chee Chang, Eunsol Choi, Jordan Lee Boyd-Graber, Aakanksha Naik

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: The "Smart" Librarian Who Doesn't Know You

Imagine you walk into a massive, high-tech library (the internet's entire collection of scientific papers) and ask a super-smart librarian (an AI) to find you the best books on a specific topic.

In the past, the librarian would just grab the most popular books or the ones with the most stars. But you, a researcher, have a specific style. Maybe you hate long introductions, maybe you only care about experiments, or maybe you prefer visual diagrams over text. You want the librarian to know you so they can curate the perfect list just for your brain.

This paper introduces a new system called MyScholarQA. It's a "Deep Research" tool that tries to be that personal librarian. It does three things:

  1. Reads your past work to build a "profile" of who you are.
  2. Asks you what you want by suggesting specific ways to tweak the search (like "skip the basics" or "focus on math").
  3. Writes a custom report based on your profile and your tweaks.

The Problem: The "Fake" Test vs. The Real Test

The researchers wanted to see if their new "Personal Librarian" was actually good. So, they tried two different ways to test it.

1. The Simulation Test (The "Robot Judge")

First, they did what most computer scientists do: they created a fake test.

  • The Setup: They made up fake researchers and fake papers. They used other AIs (called "LLM Judges") to grade the reports.
  • The Result: The AI Judges gave the new system a gold star! They said, "Great job! It cited the right papers and followed the rules."
  • The Analogy: Imagine a cooking contest where the judges are robots. The robot tastes a soup and says, "Perfect! It has salt and pepper." But the robot can't tell if the soup tastes good to a human, or if it's too spicy for a specific person's palate. The robot is just checking boxes, not tasting the food.

2. The Real User Test (The "Taste Test")

Then, the researchers did something rare: they asked 21 real human researchers to use the system.

  • The Setup: Real people used the tool, edited the profiles, and rated the reports.
  • The Result: The humans were much pickier. They found nine specific ways the system failed that the robot judges completely missed.
  • The Analogy: When a real human tasted the soup, they said, "It's too salty for me," or "I don't like the texture," or "You forgot to mention the main ingredient I asked for." The robot never noticed these things because it was just looking for "salt" and "pepper."

The Nine "Hidden" Flaws

The paper found that while the AI thought it was doing great, it was actually making subtle mistakes that only humans would catch. Here are a few examples:

  • The "Over-Confident" Profile: The AI would say, "You are an expert in Quantum Physics," when the user was actually just a beginner. It was flattering but wrong.
  • The "Off-Topic" Suggestion: The AI suggested actions that were technically correct but totally irrelevant to what the user actually wanted to know.
  • The "Vague" Report: The AI wrote a report that was too high-level. The user wanted deep details, but the AI gave a summary.

The Big Lesson: The "Robot Judges" (other AIs) were terrible at predicting what real humans wanted. They got the score wrong 100% of the time compared to real human feedback.

The Takeaway: Stop Simulating, Start Talking

The authors argue that the field of Artificial Intelligence is relying too much on simulations. We are building AI tools and testing them with other AIs, thinking, "If the AI likes it, humans will too."

This paper says: No, that doesn't work.

  • The Metaphor: You wouldn't design a new car and only test it in a video game. You need to drive it on real roads with real drivers to see if the seats are comfortable or if the brakes feel right.
  • The Conclusion: To make AI that truly helps people, we need to stop just using "Robot Judges" and start talking to Real Users. We need to watch how they use the tool, listen to their complaints, and watch them get frustrated.

Summary in One Sentence

Just because an AI thinks it's doing a great job at personalizing a report doesn't mean it actually knows what you want; you can't test a personal assistant with a robot, you have to test it with a real person.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →