Who Defines "Best"? Towards Interactive, User-Defined Evaluation of LLM Leaderboards
This paper critiques the limitations of static LLM leaderboards by analyzing biases in the LMArena dataset and proposing an interactive visualization tool that enables users to define custom evaluation priorities, thereby enhancing transparency and supporting context-specific model assessment.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Problem: The "One-Size-Fits-All" Scorecard
Imagine you are trying to buy a car. You go to a website that ranks cars based on a single number: "The Best Car Score."
This score is calculated by a small group of car enthusiasts who love racing on tracks. They weigh things like "top speed" and "cornering" very heavily. Because of this, a Formula 1 race car gets a score of 100/100, and a family minivan gets a score of 40/100.
The problem: You aren't a race car driver. You are a parent who needs to haul three kids, a dog, and groceries to the grocery store. You care about "safety," "cargo space," and "fuel efficiency."
If you buy the "Best Car" (the race car) because it has the highest score, you will be miserable. The score didn't tell you the truth; it only told you what the designers of the score cared about.
This is exactly what is happening with AI.
Currently, we use "Leaderboards" (like LMArena) to rank Large Language Models (LLMs). These leaderboards give every AI a single score based on how well it performs on a specific set of questions. But the people who designed the questions are often developers or tech enthusiasts. They ask a lot of coding questions and logic puzzles.
If you are a teacher, a doctor, or a customer service manager, that single score might be completely useless to you. It hides the fact that an AI might be terrible at your specific job, even if it's "number one" on the leaderboard.
The Investigation: Peeking Under the Hood
The authors of this paper decided to investigate the most popular AI leaderboard (LMArena) to see what was really going on. They treated the dataset like a giant bag of marbles and sorted them by color (topic).
What they found:
- The Bag is Skewed: The bag is full of "coding" and "AI" questions (about 30% of the data). It's like a car review site that only reviews sports cars.
- The Rankings Change: When they looked at specific slices of the data, the rankings flipped.
- Analogy: Imagine a "Best Athlete" contest. If the contest is 100% swimming, the swimmer wins. If the contest is 100% weightlifting, the weightlifter wins. But the leaderboard only shows the swimmer as #1. If you need a weightlifter, the leaderboard is lying to you.
- The "Tie" Problem: Sometimes, two AIs give the exact same correct answer to a math problem. But humans still vote for one over the other because one answer was written more nicely or looked prettier. The leaderboard counts this as a "win" for style, not for truth.
- The "Politics" Trap: When people ask about sensitive political topics, there is no single "correct" answer. Some people want a neutral answer; others want a strong opinion. The leaderboard just averages these conflicting opinions into one score, which doesn't make sense for anyone.
The Solution: The "Build-Your-Own-Report" Tool
Instead of giving you a fixed score, the authors built a prototype tool (a visualization interface) that lets you be the judge.
How it works (The Analogy):
Imagine a customizable buffet instead of a pre-plated meal.
- The Old Way: The chef gives you a plate with 50% broccoli and 50% steak. You have to eat it all and rate the meal.
- The New Way: The chef gives you a menu of ingredients (topics like "Math," "History," "Coding," "Creative Writing"). You get a plate and a set of tongs.
- You say, "I only care about Math and History."
- You say, "I want double the portion of History."
- You say, "I don't want any Coding."
The tool then instantly recalculates the rankings based only on your choices.
What the tool shows you:
- The "Slice" View: You can see that Model A is amazing at coding but terrible at history. Model B is the opposite.
- The "Why" View: You can click on a specific result to see the actual questions and answers. You can see why Model A won (e.g., "It gave a very structured list," or "It was concise").
- The "Trade-off" View: You can see that if you prioritize "Speed," you might lose "Accuracy." The tool helps you decide what trade-off you are willing to make.
The Experiment: Does It Work?
The researchers tested this tool with 10 real-world professionals (engineers, teachers, data analysts).
What happened:
- Surprise: People were often shocked. One person thought "Model X" was the best because it was famous. But when they filtered the data to their specific needs (e.g., "customer service"), a smaller, lesser-known model actually performed better.
- Confidence: Instead of blindly trusting a number, people felt confident in their choices because they could see the evidence. They could say, "I chose Model Y because it is great at explaining math to kids, which is what I need."
- Critical Thinking: They stopped asking "Who is the best?" and started asking "Who is the best for me?"
The Big Takeaway
The paper argues that there is no single "Best" AI. "Best" is a question that depends entirely on who is asking and what they need.
- Old Mindset: "The leaderboard says Model A is #1, so we must use Model A."
- New Mindset: "The leaderboard is a map. But I need to draw my own route based on my destination. I will use this tool to find the AI that fits my specific journey."
In short: Stop letting a small group of people decide what "good" means for everyone. Give users the tools to define "good" for themselves.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.