Reviewer Scores Are Not Comparable Across Research Areas in ML Peer Review
This position paper argues that ML conference acceptance outcomes are driven by measurement design failures rather than individual bias, demonstrating through an analysis of over 50,000 papers that reviewer scores are fundamentally incomparable across research areas, leading to acceptance probabilities that vary by up to 8x for the same score depending on the topic.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Great Scorecard Mystery
Imagine a massive, global science fair where thousands of brilliant students submit their projects every year. To decide who wins a spot in the final showcase, a team of judges gives each project a simple number score, like a rating from 1 to 10. The rule of the fair is supposed to be simple and fair: a score of 7 means "great work" no matter what the project is about. Whether you built a robot that dances, a program that predicts the weather, or a new way to teach computers to read, a 7 should mean the same thing.
But what if the judges in the "Robotics" section are super strict and only give 7s to perfect robots, while the judges in the "Weather" section are super excited and give 7s to almost anything that looks like a cloud? If you don't know which section a project belongs to, that single number becomes confusing. A 7 in one area might be a golden ticket, while a 7 in another might be a polite rejection. This is the puzzle that researchers in the field of Machine Learning are trying to solve. They are looking at how computers learn to recognize patterns, make decisions, and solve problems, and they are asking a scary question: Is the "score" we use to judge these papers actually fair, or does it depend entirely on which club you belong to?
The Paper's Big Discovery: The Scorecard is Broken
This paper, written by a team of researchers, dives into the data from a major machine learning conference called ICLR. They looked at over 50,000 papers submitted between 2021 and 2026 to see if the "score" really meant the same thing for everyone. Their conclusion is startling: The scores are not comparable.
Think of it like a video game where players in different levels are using different controllers. In the "Trending" levels (like the hot new topics of AI), the controllers are super sensitive. A tiny tap gives you a high score. In the "Old School" levels (like older, established topics), the controllers are stiff and hard to press. If you get a score of 5.0 in the "Trending" level, you might be a superstar. But if you get a 5.0 in the "Old School" level, you might be struggling. The paper found that for the exact same reviewer score, a paper's chance of getting accepted can be up to 8 times higher depending on its topic.
For example, a paper about "thinking" or "reasoning" with a score of 5.0 had a 52.6% chance of being accepted. But a paper about "image recognition" or "adversarial attacks" with that exact same score of 5.0 only had a 6.6% chance. That is an eight-fold difference! The researchers call this a "measurement design failure." It's not that the reviewers are mean or the area chairs are biased; it's that the system itself is broken because it tries to use one single ruler to measure things that are fundamentally different.
What It's NOT (The Red Herrings)
The authors were very careful to rule out other reasons for this gap. They asked: "Is it because the 'Trending' topics are just better quality?" The answer is no. If that were true, the accepted papers in those hot topics would have higher scores. Instead, the data shows the opposite: the hot topics are accepting papers with lower scores than the cold topics.
They also checked if it was because the "Old School" topics have super-expert reviewers who are just harder to please. The data says no again. The reviewers in the hard-to-please areas weren't actually more confident or more expert; they were just as enthusiastic as everyone else.
Finally, they wondered if maybe the "Trending" topics were just letting in lower-quality work because there were too many submissions (a "quality dilution"). The data showed no link between how fast a topic was growing and how much the quality dropped. In fact, some fast-growing topics kept their acceptance rates steady while others crashed.
The Real Culprit: The "Hype Cycle" and the "Patch"
So, why is this happening? The paper argues it's a structural problem. When a topic is "hot" (like a new, exciting trend), it attracts a huge crowd of new, junior reviewers who are excited and maybe a bit lenient. When a topic is "mature" (like an old, established field), it attracts senior experts who are more critical. Because the system doesn't know who is giving the score, it treats a "6" from an excited junior the same as a "6" from a grumpy senior.
To fix this, the Area Chairs (the people who make the final yes/no decision) have to guess. They use their "community priors"—basically, their gut feeling about how hard it is to get a paper accepted in that specific topic. They are essentially patching a broken system with their own intuition. This means that for a paper sitting right on the edge (a "borderline" paper), its fate isn't decided by how good the work is, but by whether the topic is currently "in fashion."
What Should We Do?
The paper doesn't just point out the problem; it offers a three-step plan to fix it, without needing to completely rebuild the whole system:
- Be Transparent: The conference organizers should publish a report showing the acceptance rates for every topic, broken down by score. If a paper gets a 5.0, everyone should be able to see: "Oh, a 5.0 in Topic A gets accepted 50% of the time, but a 5.0 in Topic B only gets accepted 6% of the time." This makes the hidden bias visible.
- Create Custom Rulebooks: Instead of using one generic "1 to 10" scale for everything, each research community should write its own specific checklist (rubric) for what makes a good paper. A "good" theory paper looks different from a "good" engineering paper, and the rules should reflect that.
- Use Better Signals: The paper suggests using "calibrated signals." For example, instead of just looking at the raw number, the system could adjust the score based on the reviewer's history (did they usually give high scores or low scores?). It also suggests letting authors rank their own papers if they submit multiple ones, which helps the system understand the relative quality of work better than a raw number ever could.
The Bottom Line
The paper concludes that the current system is like a race where runners in different lanes are running on different surfaces, but everyone is judged by the same stopwatch. The runners in the "hot" lanes are getting a head start not because they are faster, but because the track is easier. The authors argue that until we fix the track and make the scoring fair across all lanes, the best papers might still get rejected simply because they are working on a topic that isn't currently "cool." The solution isn't to blame the judges, but to admit that the ruler we are using is the wrong size for the job.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.