← Latest papers
💻 computer science

Diversity is the Strength of the AI Crowd

This paper demonstrates that maximizing the accuracy of AI forecasting ensembles requires strategically combining diverse models with complementary errors rather than simply aggregating multiple forecasts from similar high-performing LLMs.

Original authors: Matthew Aitchison, Scott Jeen, Toby Shevlane, Ben Day

Published 2026-06-30
📖 4 min read☕ Coffee break read

Original authors: Matthew Aitchison, Scott Jeen, Toby Shevlane, Ben Day

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to predict the outcome of a series of coin flips, but instead of flipping a coin, you are asking a group of very smart, but slightly different, computers (AI models) to guess whether a future event will happen or not.

This paper asks a simple question: If you have a limited number of "guesses" you can make, should you ask your single smartest computer to guess five times, or should you ask five different computers to guess once each?

Here is the breakdown of what the researchers found, using everyday analogies.

The Old Way: "The Super-Expert"

For a long time, the standard approach was to find the single smartest AI model available and ask it the same question over and over again. The logic was: "If this model is the best, more guesses from it must be better."

The researchers call this "indiscriminate sampling." It's like asking your one best friend to write five different essays on the same topic. Even if they try to write differently, they are using the same brain, the same memories, and the same style. Their answers will be very similar to each other.

The New Discovery: "The Diverse Team"

The paper argues that this approach is wrong. Instead, the best results come from building a diverse team of different models.

Think of it like a sports team. If you have five players who are all excellent at shooting basketballs, but they all stand in the exact same spot and shoot the same way, you aren't covering the whole court. You need a mix: one player who is great at shooting, one who is great at defense, and one who is great at passing. Even if the "passer" isn't the best shooter, their unique skill makes the whole team stronger.

The Key Findings

1. Smart models think alike
The researchers tested several top-tier AI models (like GPT-5, Gemini 3 Pro, and others). They found that these "super-smart" models are actually very similar in how they think. When asked the same question, they tend to give very similar answers.

  • The Analogy: If you ask five identical twins the same riddle, they will likely give you the same answer. Asking them five times doesn't give you new information; it just gives you the same answer five times.

2. The "Outlier" is the MVP
One model in their test group, called Grok 4, was interesting. It wasn't the absolute #1 most accurate model on its own (it was actually #3). However, it thought very differently from the others. Its answers were less correlated with the rest of the group.

  • The Analogy: Imagine a puzzle. You have four pieces that fit together perfectly but leave a gap. You find a fifth piece that doesn't look like the others, but it fits perfectly into that gap. Even though the fifth piece looks "weird" compared to the others, it is the most valuable piece for finishing the puzzle.

3. Diversity beats raw accuracy
When the researchers combined the models, the winning team wasn't made of the single best model repeated five times. The winning team was a mix:

  • A custom-tuned model (the "specialist").
  • Two of the top-tier models (the "generalists").
  • Grok 4 (the "diverse thinker").

Even though Grok 4 wasn't the most accurate on its own, it was the hardest to replace. If you took Grok 4 out of the team, the group's performance dropped significantly. If you took out one of the "top" models, the group barely noticed.

The Big Lesson

The paper concludes that the strength of an "AI Crowd" doesn't come from asking the same smart person to guess a million times. It comes from asking different smart people who make different kinds of mistakes.

  • The Mistake: Asking the same model to guess 5 times. (Like asking one person to write 5 different stories; they will all sound the same).
  • The Solution: Asking 5 different models to guess once. (Like asking a poet, a scientist, a comedian, and a historian to tell a story; their different perspectives create a much richer and more accurate picture).

In short: To get the best predictions, don't just stack the "best" model on top of itself. Mix in models that think differently, even if they aren't the absolute #1 on the leaderboard. Diversity is the secret sauce.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →