Surprisal Minimisation over Goal-directed Alternatives Predicts Production Choice in Dialogue
This paper demonstrates that modeling utterance production as a probabilistic cost-sensitive choice over language model-generated, goal-directed alternatives reveals that surprisal minimisation relative to these alternatives is the strongest predictor of speaker choices in dialogue, outperforming uniform information density and length-based costs.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are at a dinner party, and you need to finish a sentence your friend started. You have a specific point you want to make (your goal), but you also know your friend is listening and trying to understand you.
This paper is like a detective story investigating how we choose the exact words to finish that sentence. The authors wanted to solve a mystery: When we speak, are we trying to make our own job easier, or are we trying to make our listener's job easier?
To solve this, they used a clever trick involving two different "lists of options" (alternatives) and a digital brain (a large language model).
The Two Lists of Options
The researchers realized that to understand why we pick certain words, we have to know what other words we could have picked. They created two types of lists:
The "Same Goal" List (Goal-Directed): Imagine you want to say, "The meeting was a disaster."
- Option A: "The meeting was a disaster."
- Option B: "The meeting was a total catastrophe."
- Option C: "The meeting went terribly."
- All these options mean the exact same thing. They are just different ways to say the same thing. This list represents the choices a speaker has when they are just trying to get their specific message across.
The "Anything Goes" List (Goal-Agnostic): Now, imagine the listener doesn't know what you're going to say yet. They are just guessing based on the conversation so far.
- Option A: "The meeting was a disaster."
- Option B: "The meeting was rescheduled."
- Option C: "The meeting was canceled."
- Option D: "The meeting was fun."
- This list represents everything a listener might expect to hear next, regardless of what you actually intended to say.
The Four "Cost" Measures
The paper tests four different ideas about what makes a sentence "expensive" or "hard" to produce. Think of these as different ways to measure effort:
- Surprisal (The "Surprise" Factor): How unexpected is this word? If you say "The meeting was a..." and then say "disaster," it's not very surprising. If you say "The meeting was a... pineapple," that's high surprise. High surprise is "expensive" for the brain to process.
- Uniform Information Density (The "Smoothness" Factor): Do the words flow evenly, or is there a sudden spike of hard words followed by easy ones? Speakers generally prefer a smooth, steady rhythm.
- Length (The "Brevity" Factor): Shorter is usually easier than longer.
- Global Uniformity: Similar to smoothness, but looking at the whole sentence at once.
The Experiment: Who is the Boss?
The researchers used a super-smart AI (GPT-4) to generate thousands of these "Same Goal" and "Anything Goes" lists for real human conversations. Then, they compared what humans actually said against the AI's lists.
They asked: Did humans pick the option that was cheapest according to these rules?
Here is what they found:
1. The "Speaker" Rule Wins (Surprisal):
When looking at the "Same Goal" list, humans overwhelmingly picked the option with the lowest surprise.
- Analogy: Imagine you are driving to work (your goal). You have three routes that all get you there. You always pick the one with the least traffic (lowest surprise).
- Conclusion: When we know exactly what we want to say, we pick the words that are easiest for us to produce. We stick to the most common, predictable paths. This suggests Surprisal is a speaker-side cost.
2. The "Listener" Rule is Weak:
When looking at the "Anything Goes" list, the results were mixed.
- Humans didn't consistently pick the "smoothest" or "shortest" options to help the listener.
- In fact, for some measures, humans sometimes picked more surprising or longer options than the AI thought they should.
- Conclusion: While we care about the listener, our drive to pick the most natural, low-surprise words for our own goal is a much stronger force than our drive to smooth out the information for the listener.
The Big Takeaway
The paper argues that you can't just say "speakers minimize cost." You have to ask: "Minimizing cost for whom, and against what options?"
- If you look at the options that all say the same thing (Goal-Directed), speakers act like efficient drivers, picking the smoothest, most predictable route (Low Surprisal).
- If you look at the options that change the meaning entirely (Goal-Agnostic), the rules get messy.
In simple terms: When we speak, we are mostly trying to make our own job easy by picking the most familiar words to express our specific idea. We aren't constantly rearranging our sentences to make them perfectly smooth for the listener, at least not as much as previous theories suggested. We are the drivers, and we pick the path of least resistance to our destination.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.