← Latest papers
📊 statistics

Rejoinder: The ICML 2023 Ranking Experiment: Examining Author Self-Assessment in ML/AI Peer Review

This rejoinder addresses feedback on the ICML 2023 Ranking Experiment by structuring its response around four key themes: framing peer review as statistical estimation, mitigating equity and strategic issues in the Isotonic Mechanism, integrating complementary signals like reviewer rankings, and proposing a human-centered framework for the era of generative AI.

Original authors: Buxin Su, Jiayao Zhang, Natalie Collina, Yuling Yan, Didong Li, Kyunghyun Cho, Jianqing Fan, Aaron Roth, Weijie Su

Published 2026-05-26
📖 5 min read🧠 Deep dive

Original authors: Buxin Su, Jiayao Zhang, Natalie Collina, Yuling Yan, Didong Li, Kyunghyun Cho, Jianqing Fan, Aaron Roth, Weijie Su

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the world of top-tier computer science research (specifically Artificial Intelligence) as a massive, chaotic music festival. Every year, thousands of bands (researchers) want to play on the main stage (get their papers published). But there are only a few hundred judges (reviewers) to listen to them, and the festival is growing so fast that the judges are drowning in applications.

The authors of this paper, a team of statisticians and computer scientists, argue that we need to stop treating this festival like a simple popularity contest and start treating it like a statistical puzzle. Here is a breakdown of their ideas using everyday analogies:

1. The Problem: Too Many Tickets, Not Enough Judges

The festival is exploding. In 2026, they received over 26,000 songs to judge. The judges are tired, and they can't listen to every song perfectly. Sometimes a judge is having a bad day, sometimes they just don't "get" a specific genre, and sometimes the music is just too good or too bad for them to judge fairly.

The authors say: "We can't just let the crowd decide (like counting 'likes' on social media) because the crowd isn't expert enough. We need a better way to figure out which songs are actually the best."

2. The Solution: The "Author's Playlist" (The Isotonic Mechanism)

The authors propose a clever trick called the Isotonic Mechanism.

The Analogy: Imagine you are a judge at a cooking contest. You have to rate 100 dishes. You might be a bit harsh on spicy food or a bit too soft on desserts. Your scores are "noisy" (a bit messy).
Now, imagine the chef (the author) who made the dish is standing right there. They know their own food better than anyone. They know exactly how their 10 dishes compare to each other.

  • The Old Way: The chef just says, "My dish is a 9 out of 10." (This is easy to lie about; everyone wants a 10).
  • The New Way: The chef is asked to rank their own 10 dishes from best to worst. "Dish A is better than Dish B, which is better than Dish C."

Why this works: It is very hard to lie about a ranking of your own work without getting caught. If you say your "Dish C" is the best, but your "Dish A" is actually terrible, the math can spot the lie. By asking authors to rank their own submissions, the system can "clean up" the messy scores from the judges. It's like using the chef's own knowledge to fix the judge's bad hearing.

3. Addressing the "Unfairness" Worry

Some people worried: "What if a famous chef with 50 dishes gets a huge advantage over a new chef with only 2 dishes? The ranking system works better for the person with more items to rank."

The Authors' Fix: They realized that if you just let one person rank everything, it's unfair. So, they suggest a "grouping" strategy.

  • The Analogy: Instead of asking the famous chef to rank all 50 dishes in one giant line, ask them to group them into small "tasting menus" (e.g., "The Soup Menu," "The Dessert Menu"). Rank the soups against each other, and the desserts against each other.
  • The Result: This levels the playing field. The math shows that even with this grouping, the system still makes the scores much more accurate, but it stops the famous chefs from dominating the results just because they have more entries.

4. The "Strategy" Problem: Can Chefs Cheat?

The authors admit that if the ranking actually decides who gets to play on the main stage, chefs might try to game the system.

  • The Risk: A chef might rank a terrible dish as "great" just to get it accepted, or rank their best dish lower to "save" it for next year.
  • The Reality Check: The authors say, "We can't stop all cheating, but we can design the rules to make cheating harder." They suggest starting small (using rankings just to flag which papers need a second look, rather than immediately accepting/rejecting them) so we can learn how people behave before making it a high-stakes game.

5. The Future: Humans + AI, Not AI vs. Humans

The paper ends with a big picture view. Artificial Intelligence is changing how research is written and reviewed. AI can write code, check math, and even draft reviews.

The Authors' Stance:

  • Don't replace the humans. AI is a tool, like a super-powered calculator. It can organize the data and find patterns, but it shouldn't be the final judge.
  • The "Human-in-the-Loop": Think of AI as the stagehand who sets up the lights and organizes the instruments. The human judge is the one who decides if the music is actually good.
  • Why? Because science isn't just about data; it's about creativity, ethics, and "taste." AI might miss the subtle soul of a paper or hallucinate (make things up). Humans need to stay in charge of the final decision.

Summary

The paper argues that to fix the overwhelmed system of judging AI research, we should treat it like a math problem. By asking authors to honestly rank their own work, we can "de-noise" the judges' scores and find the best papers more accurately. However, we must be careful not to let the system be gamed, and we must ensure that while AI helps us organize the chaos, humans remain the ones holding the gavel.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →