← Latest papers
💬 NLP

Sixteen models, fewer than two voices: measuring ensemble dispersion where no answer is uniquely correct

This paper introduces a per-model dissent contribution metric to analyze how sixteen language models generate diverse interpretations of psychotherapeutic cases, revealing that while model identity significantly structures this dissent, the resulting dispersion is driven more by clinical content and specific ensemble composition than by standard model categories or the intended interpretive openness of the cases.

Original authors: Mario Vega-Barbas, Lidia Mora-Valenciano, Iván Pau, Fernando Seoane, Farhad Abtahi

Published 2026-08-04
📖 7 min read🧠 Deep dive

Original authors: Mario Vega-Barbas, Lidia Mora-Valenciano, Iván Pau, Fernando Seoane, Farhad Abtahi

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to solve a tricky puzzle, but instead of asking just one friend, you ask a whole room full of them. You hope that because they are all different people, they will give you a wide variety of creative solutions. This is the basic idea behind using "ensembles" of Artificial Intelligence (AI) models. In the world of computer science, an ensemble is simply a group of different AI programs working together to solve a problem. The big hope is that by gathering many different "voices," you get a richer, more diverse set of answers than you would from just one.

But here is the catch: just because you have a room full of people doesn't mean they are all saying different things. If they all went to the same school, read the same books, and were taught the same rules, they might all end up giving the exact same answer, just with slightly different words. This paper dives into a specific corner of AI research called "diversity measurement." It asks a simple but profound question: When we ask a group of AIs to solve a problem that doesn't have one single right answer (like interpreting a complex human emotion), do they actually think differently? Or are they all just echoing each other? The researchers care about this because if we think we are getting a wide range of perspectives, but we are actually getting a crowd of clones, we might make bad decisions based on a false sense of security.

The Great AI Echo Chamber

In this study, the researchers set up a massive experiment to see how much "noise" a group of AI models actually makes. They gathered a panel of 16 different AI models from 10 different families (think of these as different brands or lineages of AI). They asked these models to act like experienced therapists and write a case study for a made-up patient. The patient's story was vague enough that there wasn't just one "correct" way to interpret it; a good therapist could look at the same story and see many different angles.

The team wanted to measure the "dissent" or disagreement among the models. They used a special mathematical tool called the Vendi Score. Imagine you have a group of people shouting out answers. If everyone shouts the exact same word, the Vendi Score is low (like 1). If everyone shouts a completely different word, the score is high (like 16). The researchers wanted to see if their 16 models would act like 16 unique individuals or if they would collapse into a single, boring voice.

The Big Surprise: Fewer Than Two Voices

The results were a bit of a shock. Even though they had 16 different models from 10 different families, the group behaved, on average, as if it were only 1.69 distinct voices.

To put that in perspective, the researchers also tested a single AI model by asking it the same question 16 times on its own. Even just one model, by itself, managed to produce 1.43 different versions of the answer just by changing its mood slightly (what scientists call "stochastic variation"). So, the entire team of 16 models only added about a quarter of a "voice" more diversity than a single model could generate by itself.

It's like hiring a choir of 16 singers from 10 different music schools, expecting a symphony of unique styles, only to find out they all sound like one person humming the same tune, slightly off-key. The paper explicitly rules out the idea that simply adding more models to the group automatically creates more diversity. In fact, the study shows that adding more models to a panel is a poor way to get more perspectives.

Who is the Rebel? (And Why It Depends on the Crowd)

The researchers also wanted to know: "Is there one specific model that is always the rebel, the one that disagrees with everyone else?" They tracked which model was the "most divergent" in every single round of the experiment.

Here is the twist: The identity of the rebel changed depending on who else was in the room.

In a smaller experiment with just three models, one specific model (GPT-4o) was the rebel most of the time. But when the team expanded the group to 16 models, that same model dropped to third place. New "rebels" emerged, like Ministral 14B and Llama 3.3 70B, who took the lead in disagreeing with the group.

This proves that being the "outlier" isn't a permanent personality trait of a specific AI. It's a relationship. A model might seem very different only because the other models in the group happen to be very similar to each other. If you swap out the group, the "rebel" might suddenly look like the "conformist." This means that if a system shows you a "divergent voice" to help you make a decision, that voice is telling you more about the specific group of AIs present right now than it is about the specific AI model itself.

Does Size or Brand Matter?

The team also tested two common assumptions people have about AI:

  1. Does bigger mean different? They compared "small" models to "large" models within the same family. The results were messy. Sometimes the smaller model was the rebel; sometimes the larger one was. There was no consistent rule.
  2. Does the family name matter? They looked at models from the same "family" (like different versions of the same brand). While models from the same family did tend to be a bit more similar to each other than to strangers, the evidence was weak because they only had a few pairs to compare.

The paper concludes that you cannot predict how "divergent" an AI will be just by looking at its size or its brand name. You have to actually measure it in the specific context you are using it.

What Does the AI Actually Disagree About?

Finally, the researchers checked if the models disagreed more on the cases that were designed to be open to many interpretations. They had sorted the patient stories into three types: those with one clear reading, those with a shared but ineffective reading, and those with genuinely distinct readings.

They expected the models to disagree the most on the "open" stories. They didn't. The models disagreed just as much (or even slightly less) on the "open" stories as they did on the others. Instead, the amount of disagreement was driven by the clinical content of the story. For example, stories about "depression" sparked more disagreement than stories about "trauma."

This suggests that the AI models aren't sensitive to the "openness" of a problem in the way humans might be. They are reacting to the specific details of the story (like the type of illness mentioned) rather than the philosophical complexity of the situation.

The Takeaway

This study doesn't say AI is useless; it just says we need to stop assuming that a big group of AIs is automatically a smart, diverse group. If you put 16 models in a room, you might only get the wisdom of two. The "diversity" you get isn't a fixed number you can buy by adding more models; it's a fragile thing that depends on the specific mix of models and the specific problem they are solving.

The most important lesson is that if you are using a group of AIs to help you make a tough decision, don't just count the number of models you have. You need to measure how different they actually are, because the "outlier" voice you see might just be a temporary effect of who is sitting at the table, not a deep insight from a unique machine.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →