A Consensus-Based Framework for Relative Preference Evaluation of Large Language Models
This paper proposes a scalable, consensus-based framework that evaluates Large Language Models by aggregating their mutual rankings of anonymized responses to generate a Relative Intelligence Index, offering a model-driven alternative to traditional static benchmarks for assessing response quality in scenarios where multiple valid answers exist.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Great AI Taste-Test
Imagine a world where computers have learned to talk, write, and solve problems so well that they can pass as human. This is the realm of Large Language Models (LLMs), the super-smart digital brains behind many of the chatbots and tools we use today. For a long time, scientists tested these brains using strict "answer keys," like a math test where there is only one right number. But life isn't always a math test. Sometimes, a question has many good answers, and the "best" one depends on how clear, helpful, or clever it feels. This is where things get tricky: how do you grade an essay when there are five different ways to write a perfect one?
To solve this, researchers have started using AI to grade AI. It's like asking a panel of judges to decide which of their friends told the funniest joke. But here's the catch: if the judges are all friends who went to the same school, they might all like the same style of joke, even if it's not actually the funniest. This paper dives into that exact problem. It asks: If we let a group of different AI models judge each other's answers blindly, without knowing who wrote what, can we find a pattern of what they prefer? It doesn't claim to find the "truth," but rather a map of how these digital minds agree on what sounds good.
The Blind Taste-Test of the Digital World
Imagine you are at a massive, high-tech potluck dinner. Instead of bringing food, five different super-intelligent robots (let's call them the "Chefs") bring their best recipes for the same dish. The goal isn't to see who cooked the "correct" meal—because maybe there are ten ways to make a perfect stew—but to see which meal the other robots think looks and tastes the best.
This is exactly what Mohtashim Khan did in this paper. Instead of asking a human to taste-test the answers, he set up a "blind taste-test" where the Chefs judge each other.
The Setup: A Secret Identity Game
The experiment involved five top-tier AI models: Claude, ChatGPT, Gemini, Grok, and Mistral. The researcher gave them 25 different challenges, ranging from solving tricky math puzzles and writing code to answering safety questions and general knowledge trivia.
Here is the magic trick: When the robots were judging, they didn't know who wrote the answer. The answers were stripped of names and shuffled around, labeled only as "Answer 1," "Answer 2," and so on. Each robot had to rank all the answers from best to worst based on how clear and helpful they seemed. They did this completely independently, without talking to each other or seeing what the others thought.
The Score: The "Relative Intelligence Index"
After the judging was done, the researcher tallied up the votes. He created a new score called the Relative Intelligence Index (RII). Think of this not as a grade for being "smart," but as a popularity contest score. A high RII means a robot's answers were frequently picked as the "best" by its peers. A low RII means its answers were often ranked lower.
What Did They Find?
The results were a bit like a sports league where different teams win on different days.
- The Consistent Winners: Claude and Grok often came out on top. Claude was particularly strong in math and coding, while Grok shined in logical puzzles and creative thinking.
- The Steady Performers: Gemini and Mistral didn't always win the "best answer" prize, but they were incredibly consistent. Their scores didn't swing wildly from one test to the next.
- The Speed vs. Quality Trade-off: The study also timed how long each robot took to cook up an answer. Interestingly, the robots that took the longest to think (like Gemini and Claude) were often the ones the others liked the most. The fastest robots (like Mistral) were quick but didn't always get the highest votes.
The Big "But"
The paper is very careful to say what this score isn't. It is not a measure of who is actually "right" or who is most helpful to a human. It's just a measure of how much the robots like each other's style. It's possible that all the robots were trained on similar books and websites, so they might just be agreeing because they have the same "taste buds," not because one is objectively better.
The Takeaway
This framework suggests that when there isn't one single right answer, we can use a "consensus" method to see which AI outputs feel the most polished and coherent to other AIs. While it doesn't replace human judgment, it offers a scalable way to compare models when the answers are subjective. The study found that while some models are generally preferred by their peers, the results can shift depending on the topic and the specific run of the test. It's a new way to look at AI performance, focusing on "what do the models think of each other?" rather than just "did they get the answer right?"
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.