CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks
CoEval is an open-source framework that generates contamination-free, attribute-controlled benchmarks and employs a diverse ensemble of judge models to reliably rank language models for custom tasks without requiring labeled data or trusting public benchmarks.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a hiring manager trying to pick the best employee for a very specific job, like "organizing a chaotic warehouse" or "writing funny jokes about gardening." You have a problem: you don't have any past performance reviews (labeled data), and the standard tests everyone else uses (public benchmarks) are useless because the candidates have already memorized the answers from studying them beforehand.
This is the exact problem CoEval solves. It's a new, open-source tool that helps you rank AI models for your specific needs without needing human experts to grade them or worrying about cheating.
Here is how it works, broken down into simple analogies:
1. The "Fresh Test" Generator (No Cheating Allowed)
Usually, AI models take the same old tests (like the SATs or GRE) that have been around for years. Because these tests are so popular, the AI models have likely seen the questions during their training and just memorized the answers. It's like a student memorizing the answer key instead of learning the math.
CoEval changes the game. Instead of using old questions, it uses a "Teacher AI" to invent brand-new questions on the spot, based only on a description of your task.
- The Analogy: Imagine a teacher who creates a math test while the student is in the room. The student couldn't have memorized the answers because the questions didn't exist until that second.
- The Result: The paper proves these new questions are 100% fresh. They share no 13-word phrases with any major public test banks, meaning the AI can't cheat by recalling old data.
2. The "Jury" of Judges (Diversity Over Numbers)
Once the AI models answer these fresh questions, someone has to grade them. You could ask one super-smart AI to grade them, but that's risky. That single judge might have weird biases, like preferring long answers over short, good ones, or liking answers that sound like its own family of models.
CoEval uses a "Jury" approach. It gathers a small group of judges from different companies (e.g., one from OpenAI, one from Anthropic, one from Google).
- The Analogy: Think of a courtroom. If you have one judge who is known to be grumpy, the verdict might be unfair. But if you have a jury of three people from different backgrounds, their individual quirks cancel each other out. If one judge likes long answers and another hates them, the average score becomes fair.
- The Big Discovery: The paper found that who is on the jury matters more than how many people are there. Adding more judges from the same company just amplifies their shared bias. But a small, diverse group of judges is incredibly reliable. In fact, a single judge can sometimes be "anti-correlated" (giving the wrong answer), but the diverse jury almost never is.
3. The "No-Human" Factory
Usually, to build a good test, you need humans to write questions and grade answers. This is slow and expensive.
- The Analogy: CoEval is like a fully automated factory. You give it a blueprint (a description of your task), and it automatically:
- Designs the test questions.
- Has the AI candidates take the test.
- Assembles the diverse jury to grade the answers.
- Spits out a final ranking.
- The Cost: The authors ran a massive study with nearly 8,000 evaluations for the price of a cup of coffee ($5.89). It's so cheap you could run a new test every time a new AI model comes out.
4. Why This Matters
The paper tested CoEval in three real-world scenarios where no "correct" answers existed:
- Drug Interactions: Figuring out if two medicines clash.
- Clinical Reasoning: Diagnosing patient symptoms.
- Legal Analysis: Interpreting laws.
In these areas, there are no standard "gold standard" tests. CoEval successfully ranked the models, showing which ones were best for these specific jobs. It even caught that a smaller, cheaper model was actually worse than a larger one, something a single biased judge might have missed.
The Bottom Line
CoEval is a tool that says: "Don't trust the old, memorized tests. Don't rely on a single judge who might be biased. Instead, generate a fresh test for your specific job and have a diverse, small team of AI judges grade it."
It turns the messy, expensive process of testing AI into a cheap, repeatable, and fair process that any team can use to find the right AI for their specific needs.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.