Diverse, not Short: A Length-Controlled Data Selection Strategy for Improving Response Diversity of Language Models
The paper introduces Diverse-NS, a length-controlled data selection strategy that mitigates the systematic bias toward shorter outputs in language models, thereby significantly enhancing response diversity and expressiveness across various creative tasks while maintaining quality.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Problem: The "Short and Boring" Trap
Imagine you have a very talented writer (a Large Language Model) who has been trained to be helpful, safe, and polite. You ask them to write a story.
The problem is that in trying to be "perfect" and safe, the writer has started acting like a nervous student taking a test. They are afraid to write too much because they think shorter answers are better answers.
Why? Because the tools used to grade their work (diversity metrics) are broken. These tools are like a judge who thinks a 3-word sentence is more "creative" than a 30-word paragraph simply because it's shorter. The writer notices this and starts giving you short, safe, repetitive answers to get a good grade.
The authors call this the "Brevity Bias." It's like a chef who stops adding spices because the taste-tester only likes plain food. The result is that the AI becomes less creative, less expressive, and eventually, all AI outputs start sounding exactly the same (a phenomenon called "model collapse").
The Solution: "Diverse, not Short" (Diverse-NS)
The authors created a new strategy called Diverse, not Short (Diverse-NS). Think of this as a new training program for the writer that fixes the broken grading system.
Here is how it works, step-by-step:
The "Two-Draft" Game:
The researchers ask the AI to write a story twice.- Draft A: The AI writes naturally (its default, often boring style).
- Draft B: The AI is asked to rewrite the story with a specific rule: "Make it more interesting and varied, but keep it the exact same length as Draft A."
The "Fair Judge" Filter:
Usually, if you ask for more creativity, the text gets longer. But the researchers have a strict rule: If Draft B is longer than Draft A, throw it away.
They only keep the pairs where the "more creative" version is roughly the same length as the "boring" version. This teaches the AI: "You can be creative without being wordy."The "Small Teacher" Trick:
The researchers found something surprising. A smaller, cheaper AI model (like a 7-billion parameter model) can act as a "diversity teacher" for a much larger, more expensive model (like a 13-billion parameter model).- Analogy: Imagine a small, scrappy local artist teaching a famous, big-budget studio how to paint more uniquely. The small artist knows how to break the rules creatively, and the big studio learns from them. This saves a lot of money and computing power.
The Results: More Flavor, Same Size
The researchers tested this on four different creative tasks:
- Divergent Associations: Listing words that are very different from each other.
- Persona Generation: Creating unique character profiles (names, jobs, cities).
- Alternate Uses: Thinking of weird ways to use a common object (like a brick).
- Creative Writing: Writing short stories using three random words.
What happened?
- More Variety: The models trained with "Diverse, not Short" produced much more varied and interesting outputs. They used more unique words and ideas.
- No Loss in Quality: Usually, when you make things more diverse, they get worse quality (like a messy, chaotic story). But here, the quality stayed high. In some cases, it even got better.
- Efficiency: They only needed 3,000 examples to teach the model this new skill. That's like teaching a student for a few hours instead of years.
The New Tool: The "Diversity Decile"
The authors also realized that the old tools for measuring creativity were unfair to long answers. So, they invented a new ruler called Diversity Decile.
- The Analogy: Imagine a race. The old ruler just measured how fast you ran, but it ignored that some runners had heavy backpacks (long text) and some didn't (short text).
- The New Ruler: The "Diversity Decile" looks at how fast you ran relative to other runners with the same size backpack. It asks: "Is this long story more creative than other long stories?" This ensures that long, expressive answers get the credit they deserve.
Summary
The paper argues that AI models have become too short and repetitive because the way we measure "creativity" accidentally rewards brevity.
By using a strategy called Diverse, not Short, the authors taught models to be creative without making their answers longer. They did this by:
- Comparing a boring draft to a creative draft of the same length.
- Using a smaller, cheaper AI to teach a bigger AI how to be diverse.
- Creating a new, fairer ruler to measure creativity that doesn't penalize long answers.
The result is AI that is more expressive, more creative, and still high-quality, all while being cheaper and faster to train.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.