Mind the Gap: How Elicitation Protocols Shape the Stated-Revealed Preference Gap in Language Models
This paper demonstrates that the stated-revealed preference gap in language models is highly dependent on elicitation protocols, showing that while allowing neutrality in stated preferences improves correlation with forced-choice revealed preferences, introducing abstention in revealed preferences or using system prompt steering fails to reliably resolve the mismatch.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to understand the personality of a new friend. You have two ways to do this:
- Ask them directly: "Do you value honesty more than kindness?" (This is their Stated Preference).
- Watch them act: Put them in a tough situation where they have to choose between being honest or being kind, and see what they actually do (This is their Revealed Preference).
In the world of AI, researchers have noticed a problem: sometimes an AI says it loves "Honesty," but when put in a tricky story, it chooses "Kindness" instead. This mismatch is called the Stated-Revealed (SvR) Gap.
This paper, titled "Mind the Gap," investigates why this gap exists and, more importantly, how the way we ask the questions changes the answer. Think of the "elicitation protocol" as the rules of the game we play with the AI.
Here is what the researchers found, using simple analogies:
1. The "Forced Choice" Trap (The Old Way)
Previously, researchers played a game where the AI was forced to pick a side. Imagine a referee shouting, "You must choose: A or B! No talking, no hesitation!"
- The Problem: If the AI is actually unsure, or thinks "It depends on the situation," it still has to pick A or B. It's like forcing a person to choose between "Pizza" and "Salad" even if they are full and don't want either.
- The Result: The AI picks randomly or based on the order of the words, creating "fake" preferences. When researchers compared these fake choices to what the AI said it valued, the connection was weak and messy.
2. The "Opt-Out" Button (The New Way for Stated Preferences)
The researchers tried a new rule for the "Ask them directly" part. They said, "You can pick A, B, or you can say 'I don't care' or 'It depends'."
- The Analogy: Imagine asking a friend, "Do you prefer coffee or tea?" but allowing them to say, "I'm not a drinker," or "It depends on the time of day."
- The Result: This acted like a filter. It removed the "weak signals" (the AI's guesses or random choices). When the researchers only looked at the times the AI strongly picked a side, their stated preferences matched their actual behavior much better. The "gap" shrank significantly.
3. The "Too Many Opt-Outs" Problem (The New Way for Revealed Preferences)
Then, they tried letting the AI say "It depends" during the tough scenarios (the revealed preferences).
- The Analogy: Now, imagine putting your friend in a real-life crisis and saying, "You can choose Action A, Action B, or just say 'I don't know'."
- The Result: The AI started saying "It depends" or "Equal" almost all the time! It was like a student who, when given a hard test, just writes "I'm not sure" on every answer.
- The Consequence: Because the AI refused to make a firm choice so often, the researchers couldn't build a clear ranking of what the AI actually valued. The connection between what they said and what they did vanished (dropped to near zero or even negative).
4. The "Instruction Manual" Experiment (Prompt Steering)
Finally, the researchers tried to "fix" the AI's behavior by giving it a cheat sheet. Before the tough scenarios, they told the AI: "Here is your own list of values you said you liked. Please follow this list strictly."
- The Analogy: It's like telling a driver, "Remember, you said you love safety, so please drive safely!" right before a race.
- The Result: It didn't work reliably. For some AIs, it helped a little. For others (like the "Claude" family), it actually made things worse. The AI seemed to ignore the cheat sheet or get confused by it. The researchers found that this "instruction manual" trick works for small lists of values but fails miserably when the list is long (16 different values).
The Big Takeaway
The main lesson of this paper is that how you ask the question changes the answer.
- If you force an AI to choose when it's unsure, you get noisy, fake data.
- If you let an AI say "I don't know" too much, you get no data at all.
- Simply telling an AI to "follow its own rules" doesn't reliably fix the problem.
The authors conclude that to truly understand AI values, we need methods that acknowledge that sometimes, the AI (just like a human) genuinely doesn't have a clear preference, and we need to design our tests to handle that uncertainty rather than forcing a choice or ignoring it.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.