Justice in Judgment: Unveiling (Hidden) Bias in LLM-assisted Peer Reviews
This paper investigates bias in LLM-assisted peer reviews by revealing that models exhibit significant affiliation, seniority, and publication history biases favoring prestigious authors, with these underlying preferences becoming even more pronounced when analyzed at the token level despite potential alignment masking.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a massive, high-stakes talent show where thousands of scientists submit their best work to be judged. Traditionally, human experts review these submissions anonymously to decide who gets a prize (publication). But recently, the organizers started hiring a new kind of judge: a super-smart computer robot (a Large Language Model, or LLM) to help write the reviews.
This paper is like a "mystery shopper" investigation. The researchers wanted to see if these robot judges are truly fair, or if they secretly have favorites based on who the author is, rather than just the quality of the work.
Here is what they found, explained simply:
1. The "Name-Tag" Test (Affiliation Bias)
The researchers took the exact same scientific paper and gave it to the robot judges multiple times. The only thing they changed was the name of the university the author supposedly came from.
- Scenario A: The paper said it came from a world-famous, top-tier university (like MIT or Cambridge).
- Scenario B: The paper said it came from a smaller, less famous university.
The Result: The robots consistently gave higher scores to the papers from the famous universities. It's as if the robot judge looked at the name tag, saw "Harvard," and immediately thought, "This must be good!" without actually reading the work as carefully. If the name tag said "Small State College," the robot was much harsher.
The Hidden Twist: The researchers noticed something even sneakier. When the robots gave their final "Yes/No" answer, they sometimes tried to act neutral. But when the researchers looked at the robot's internal thought process (its "soft ratings"), the bias was even stronger. It's like a person saying, "I'm being fair," while their internal monologue is screaming, "I definitely prefer the rich kid!"
2. The "Resume" Effect (Seniority & History)
The researchers also tested if the robots cared about the author's experience.
- Scenario A: The author was a famous, older professor with hundreds of past publications.
- Scenario B: The author was a young student with no past publications.
The Result: The robots loved the famous professor. They gave higher scores to papers from "Senior Principals" and authors with long lists of past work. They were much more skeptical of the student, even though the actual paper content was identical. The robots seemed to think, "If they've done it before, they must be doing it right again."
3. The "Gender" Guess (Gender Bias)
The researchers also swapped male and female names on the papers.
The Result: This was a bit messier. Some robots favored men, some favored women, and some were neutral. It wasn't as consistent as the university bias, but it was still there. It's like a coin flip where the coin is slightly weighted depending on which robot you ask.
4. The "Borderline" Danger
The most worrying finding was about papers that were "on the fence"—the ones that were neither clearly amazing nor clearly terrible.
- When a "famous university" name was attached to a borderline paper, the robot often changed its mind from "Reject" to "Accept."
- When a "lesser-known university" name was attached to the same borderline paper, the robot often changed its mind from "Accept" to "Reject."
This means the robot's bias wasn't just a small preference; it was strong enough to decide a scientist's career and whether their work gets published.
The Big Picture
The paper concludes that these AI judges are not the neutral, objective referees we hoped they would be. They have absorbed the real-world biases of the internet they were trained on. They seem to believe that "Prestige = Quality" and "Experience = Goodness," even when the actual work doesn't support that conclusion.
The authors warn that if we let these robots run the show without checking their work, we might end up only publishing papers from famous people at famous schools, while ignoring great ideas from talented people at smaller places. The robots are "hallucinating" fairness while actually being quite biased.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.