← Latest papers
💬 NLP

Mediocrity is the key for LLM as a Judge Anchor Selection

This paper demonstrates that the selection of an anchor model is critical for the reliability of "LLM-as-a-judge" evaluations, revealing that extreme anchors (best or worst models) undermine ranking accuracy and advocating for the use of "mediocre" anchors alongside larger benchmark sizes to ensure robust results.

Original authors: Shachar Don-Yehiya, Asaf Yehudai, Leshem Choshen, Omri Abend

Published 2026-03-18
📖 5 min read🧠 Deep dive

Original authors: Shachar Don-Yehiya, Asaf Yehudai, Leshem Choshen, Omri Abend

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: Why "Average" is the New "Best"

Imagine you are a talent scout trying to rank 22 singers in a competition. You can't listen to every single pair of singers against each other (that would take forever and cost a fortune). So, you decide to pick one singer to be the "Anchor." You will only compare every other singer against this one Anchor.

The common wisdom in the AI world has been: "Pick the best singer as your Anchor!" or "Pick the worst singer as your Anchor!" The idea was that if you compare everyone to the absolute best, you'll see who is close to the top. If you compare everyone to the worst, you'll see who is better than the bottom.

This paper says: You are wrong. Both of those choices are terrible.

Instead, the paper argues that the "mediocre" singer—the one who is just okay, right in the middle of the pack—is actually the perfect Anchor.


The Problem: The "Superstar" and the "Struggling" Anchor

Let's look at why picking the extremes fails, using two analogies:

1. The Superstar Anchor (The "Olympian")

Imagine your Anchor is a world-famous Olympic gold medalist.

  • The Scenario: You ask the judge to compare a local high school singer against the Olympian.
  • The Result: The judge immediately says, "The Olympian wins!" 99 times out of 100.
  • The Problem: This gives you zero useful information. You already knew the Olympian was better. You wasted your time and money on a comparison that didn't tell you anything about the other singers. You can't tell if Singer A is slightly better than Singer B because both of them lost to the Olympian in a landslide.

2. The Struggling Anchor (The "Amateur")

Now, imagine your Anchor is someone who can barely carry a tune.

  • The Scenario: You compare a professional singer against this amateur.
  • The Result: The judge immediately says, "The pro wins!" 99 times out of 100.
  • The Problem: Again, zero useful information. Everyone beats the amateur. You still don't know who is the best among the professionals.

The Conclusion: When you pick an Anchor that is too good or too bad, the results are "saturated." It's like trying to measure the height of a skyscraper and a garden gnome using a ruler that only goes up to 10 feet. You can't distinguish between the tall things, and you can't distinguish between the short things.


The Solution: The "Goldilocks" Anchor

The paper found that the best Anchor is the mediocre one—the singer who is average.

  • Why it works: When you compare a good singer to an average Anchor, the judge has to think hard. "Hmm, this one is slightly better."
  • The Result: When you compare a bad singer to the same average Anchor, the judge also has to think. "Hmm, this one is slightly worse."
  • The Magic: Because the Anchor is in the middle, the results are balanced. Some people win, some people lose, and the margins are close. This creates a rich dataset where you can actually see the subtle differences between the singers.

The Analogy: Think of the Anchor as a thermostat.

  • If the thermostat is set to 100°F (too hot), everyone feels cold. You can't tell who is shivering a little and who is shivering a lot.
  • If the thermostat is set to 32°F (too cold), everyone feels hot. Same problem.
  • If the thermostat is set to a comfortable 70°F, you can clearly feel who is slightly too warm and who is slightly too cool. That "Goldilocks" temperature gives you the most data.

The Surprising Findings

The researchers did a massive experiment (comparing over 850,000 pairs of AI models) and found three shocking things:

  1. The "Inverted U" Shape: If you plot the "quality" of the Anchor against how well it ranks the other models, you get a U-shape. The best Anchors are in the middle (the bottom of the U), and the worst Anchors are at the very top and very bottom.
  2. It Matters More Than the Judge: People spend a lot of time arguing about which AI model should do the judging (e.g., "Is GPT-4 better than Claude?"). The paper found that choosing the right Anchor is just as important as choosing the right Judge. A bad Anchor can ruin your results even if you have the smartest Judge in the world.
  3. We Are Wasting Money: Current benchmarks (like Arena-Hard) are often too small. Because "extreme" Anchors waste so much data (by producing too many obvious wins/losses), we need much larger datasets to get reliable results. If you use a bad Anchor, you might need double the data to get the same accuracy.

What Should You Do? (The Cheat Sheet)

If you are running an AI evaluation, here is the advice from the paper:

  • Don't pick the "Best" model as your Anchor. It's too easy to beat.
  • Don't pick the "Worst" model. It's too easy to lose to.
  • Pick the "Middle" model. Find a model that is roughly as smart as the ones you are testing. This will give you the most accurate ranking.
  • Check your data first. Before running a huge, expensive test, run a tiny test to see if your Anchor is "informative" (does it produce a mix of wins and losses?). If it's too one-sided, pick a different Anchor.

The Takeaway

In the world of AI evaluation, mediocrity is a superpower. By choosing a "good enough" Anchor instead of a "perfect" or "terrible" one, you get a clearer, fairer, and more accurate picture of how AI models actually compare to each other. It's the difference between a blurry photo and a high-definition picture.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →