More Is Not More: What Matters for Diversity in LLM Opinions?
This paper employs a factorial experiment to demonstrate that enhancing LLM opinion diversity relies less on scaling single factors like persona detail or temperature, and more on strategically combining distinct interaction architectures and recognizing the diminishing returns of excessive persona elaboration.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to host a massive, virtual dinner party where every guest is a different person with their own unique thoughts, memories, and opinions. You want to hear a wild mix of ideas, from spicy takes on politics to weird theories about the future. But here's the catch: you aren't inviting real humans. Instead, you are asking a super-smart computer program, called a Large Language Model (or LLM), to pretend to be these guests. These models are like brilliant actors who can read the entire internet and mimic human speech perfectly. However, when you ask them to play a crowd, they often all start sounding the same. They agree too much, use the same phrases, and end up with a boring, one-note conversation. This is a problem because if you want to understand what real people think, you need a crowd that actually sounds like a crowd, not a choir singing in perfect unison.
Scientists have been trying to fix this "boring party" problem. They've tried giving the computer more details about who it's pretending to be (like saying, "You are a 45-year-old teacher named Sarah") or changing how the guests talk to each other (like making them argue in a group instead of just answering one by one). They've also tried turning up the "creativity dial" on the computer to see if that makes it more random and diverse. But until now, no one had a clear map of which trick actually works best. It was like trying to bake a cake by mixing every ingredient at once and hoping for the best, without knowing if the flour, the eggs, or the oven temperature was actually what made it taste good.
This paper, titled "More Is Not More," is like a giant, scientific taste-test that finally separates the ingredients to see what really matters. The researchers set up a massive experiment with 100 different questions and 7 different computer models. They tested two main ways to make the opinions more diverse: Input Conditioning (how much detail you give the computer about who it is pretending to be) and Interaction Architecture (how the computer guests talk to each other). They also tested some "cheap tricks" people often try, like just telling the computer to "be diverse" or turning up the randomness knob.
Here is what they discovered, and it turns out the answer is a bit surprising:
1. The "More Details" Trap
Many people thought that the more backstory you give the computer, the more unique its opinions would be. You might think, "If I tell the computer it's a 34-year-old Black woman named Sarah who lives in Atlanta, goes to church, and makes $78,000 a year, it will act super unique!" But the study found that this isn't true. The biggest jump in diversity happened with just a tiny bit of info—like simply saying, "You are a school teacher." Once you added that one sentence, the computer started acting more like a person. Adding all the extra details (age, race, income, hobbies) didn't help much more. In fact, for some models, giving too many details actually made the opinions less diverse, as if the computer got confused or started sticking to stereotypes. It's like trying to make a character interesting by giving them a 10-page biography; sometimes, just knowing their job is enough to spark a unique voice, and the rest just clutters the script.
2. The Power of the Group Chat
The researchers also looked at how the computer guests interact. They compared three styles:
- The Solo Artist: The computer answers a question once and leaves.
- The Deep Interview: The computer answers, then gets asked a follow-up question, then another, building a long conversation with itself.
- The Focus Group: Five different computer "people" talk to each other in a group chat.
The results showed that the "Deep Interview" and the "Focus Group" both created much more diverse opinions than the "Solo Artist." But here's the kicker: they didn't just make the same opinions better; they explored completely different worlds of thought. The "Deep Interview" style dug deep into personal, introspective angles, while the "Focus Group" style found different, more comparative angles. The study suggests that if you want the widest range of opinions, you shouldn't just pick the "best" method. Instead, you should run both methods and combine their answers. It's like having a solo poet and a debate team write about the same topic; their answers won't overlap much, so you get a much richer collection of ideas by reading both.
3. The "Cheap Tricks" Don't Work
Finally, the paper tested the low-effort hacks that many people try first. They tried raising the "temperature" (a setting that makes the computer more random and less predictable) and adding instructions like, "Please give us diverse perspectives." The result? These tricks barely made a dent. Turning up the randomness didn't create new ideas; it just made the same old ideas sound a little more chaotic. Telling the computer to "be diverse" was like telling a shy person to "be funny"—it didn't actually change their personality. The study found that a simple, one-sentence job description (like "You are a teacher") was 2.5 times more effective at creating diverse opinions than the best "cheap trick."
The Bottom Line
The main message of this paper is that diversity isn't about doing more of the same thing. It's not about writing longer bios or turning up the randomness dial. Instead, it's about the structure of how you ask the questions. To get a truly diverse set of opinions from a computer, you need to give it a simple identity to start with, and then let it explore the topic in different ways—like having it talk to itself in a long interview and having it chat with a group of friends. If you want a real simulation of human opinion, you need a mix of strategies, not just a bigger pile of instructions.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.