Beyond expert users: agents should help users construct preferences, not just elicit them
This paper argues that AI agents should actively help users construct preferences by providing necessary domain knowledge, rather than merely eliciting them, a challenge highlighted by the proposed CoPref framework and CoShop benchmark which reveal current frontier models struggle to expand user understanding despite multiple interaction turns.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are walking into a high-end shoe store to buy hiking boots. You tell the salesperson, "I need boots that are comfortable, not too expensive, and look good."
In the world of current AI shopping assistants, the system assumes you are an expert. It thinks, "Great! You know exactly what you want. I just need to ask a few more questions to get the details: 'What size? What color? What's your budget?'"
The paper argues that this assumption is wrong. Most of us aren't experts. We don't know that "lug depth" (the tread on the bottom) matters for rocky trails, or that "waterproof" is different from "water-resistant." If the AI just asks, "Do you want waterproof boots?" and you say "I don't know," the AI hits a dead end. It can't find the right boot because it hasn't taught you why that feature matters yet.
The Core Idea: Building Preferences, Not Just Finding Them
The authors propose that AI shouldn't just be a question-asker; it needs to be a teacher.
They introduce a model called COPREF (Co-constructed Preferences). Think of your preferences not as a hidden treasure chest that the AI is trying to unlock with a key (a question), but as a house that is being built brick by brick during your conversation.
- Search Features: These are easy bricks. You know you want "red" or "size 10." The AI just asks, and you answer.
- Experience Features: These are bricks you can't see until you hold them. You don't know if you like "suede" until the AI shows you a picture of a suede boot and says, "See this texture?" You then realize, "Oh, I hate that."
- Credence Features: These are the blueprints. You don't even know the wall exists until the AI explains, "By the way, some boots have a special lining that stops blisters." Once you learn this, you can decide if you want it.
The Experiment: COSHOP
To test this, the researchers built a playground called COSHOP. It's like a video game where an AI agent tries to help a simulated user buy things (clothes, movies, or books).
The "users" in this game are programmed to be real people: they start with very few ideas. They don't know what they want until the AI helps them learn. The goal isn't just for the AI to find the right item; it's for the AI to help the user understand what they want so they can pick the right item themselves.
What They Found: The AI is Too Shy to Teach
The researchers tested five of the smartest AI models available (like the latest versions of GPT and Claude). Here is what happened:
- The "Expert" Trap: When the AI was given a user who already knew everything (an expert), the AI was amazing. It found the right item 94% of the time.
- The "Real User" Crash: When the AI tried to help a "real" user (who needed to learn), the performance plummeted. No model got above 56% accuracy.
- The Reason for Failure: The AI wasn't bad at searching the store. It was bad at teaching.
- Too many questions, not enough examples: The AI kept asking, "Do you like this?" or "Is this too expensive?" without ever showing the user what "too expensive" or "this style" actually looked like.
- Premature endings: The AI often got tired of asking questions and said, "Okay, I think I have enough info," and went to search. But the user still didn't know enough to pick the right item from the list.
- Hallucinations: Sometimes the AI made up features that didn't exist, confusing the user further.
The Analogy of the "Bad Report"
Imagine the AI is a researcher who finds the perfect hiking boot. It writes a report for you.
- The Problem: The report says, "Here is a great boot." But it forgets to mention that the boot is made of heavy leather (which you hate) or that it has a high ankle (which you need).
- The Result: Even though the AI found the perfect boot, you can't pick it because the report didn't give you the information you needed to make a choice. You pick a different, worse boot because you don't know the first one was actually the right one.
The Takeaway
The paper concludes that for AI to be truly helpful, it needs to stop acting like a retrieval system (just fetching answers) and start acting like a collaborator. It needs to realize that sometimes the user doesn't know what they want, and the AI's job is to show them examples and explain technical details so the user can construct their own preference.
Currently, even the smartest AIs are failing at this. They are great at finding things, but they are terrible at helping us figure out what we actually need.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.