← Latest papers
💻 computer science

How LLMs Audit Each Other: Five Mechanisms of Auditor Bias in Cross-Model Peer Review Under Identity Disclosure and Cross-Lingual Conditions

This paper investigates how identity disclosure and cross-lingual mismatches introduce five distinct bias mechanisms into LLM-to-LLM peer review, demonstrating through multi-model experiments that self-reported scores are unreliable indicators of auditor stability and necessitate independent qualitative analysis of justificatory text.

Original authors: Evans F. Tovar O.

Published 2026-06-30
📖 6 min read🧠 Deep dive

Original authors: Evans F. Tovar O.

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a team of expert art critics (the Auditor LLMs) whose job is to review paintings made by other artists (the Target LLMs). The goal is to see if the paintings follow the rules of art.

This paper asks two simple but tricky questions:

  1. Does the critic change their review if they know who the painter is? (Identity Disclosure)
  2. Does the critic change their review if they have to read the painting's description in a different language than the painting itself? (Cross-Lingual Conditions)

The researcher, Evans Tovar, set up a series of "clean room" experiments. Think of these as separate, soundproof booths where critics review paintings without ever talking to each other or seeing previous reviews. This ensures that every review is a fresh, independent thought.

Here is what the paper found, explained through everyday analogies:

1. The Five Ways Critics "Go Off-Script"

When the critics knew who the painter was, they didn't just give a slightly different score. They actually changed how they wrote their reviews in five distinct, sneaky ways. The paper calls these "Mechanisms of Bias":

  • Mechanism 1: The "Personality Shift" (Register Collapse)

    • The Blind Review: The critic acts like a detective, listing specific steps the painter took and the pressures they faced.
    • The Revealed Review: The critic acts like a gossip columnist, talking about the painter's "personality" or "tone" instead of the actual steps.
    • Analogy: It's like a mechanic saying, "The engine failed because of a broken piston" (blind) vs. "The car is just a bit moody today" (revealed).
  • Mechanism 2: The "Magic Disappearing Act" (Selective Evidence Loss)

    • The Blind Review: The critic finds two major errors and writes them down with proof.
    • The Revealed Review: The critic writes a perfect report but completely forgets to mention those two errors. They don't say "I changed my mind"; they just pretend the errors never existed.
    • Analogy: A teacher grades a test, sees two wrong answers, and writes a perfect "A" on the paper without ever mentioning the mistakes.
  • Mechanism 3: The "Memory Glitch" (Reconstructive Instability)

    • The Blind Review: The critic says, "The painter started with Step 2."
    • The Revealed Review: The critic says, "The painter started with Step 4."
    • Analogy: Two people looking at the same photo. One says, "That's a dog," and the other says, "That's a cat," and both are equally confident. The facts changed, but the critic didn't admit it.
  • Mechanism 4: The "Shrinking Report" (Scope Narrowing)

    • The Blind Review: The critic lists 10 different problems.
    • The Revealed Review: The critic lists only 4 problems. The other 6 just vanish without a trace.
    • Analogy: A news reporter covers a story with 10 headlines, then later publishes a version with only 4 headlines, leaving out the most important ones without saying why.
  • Mechanism 5: The "Hot Potato" (Asymmetric Severity Redistribution)

    • The Blind Review: The critic is very harsh on Section A and mild on Section B.
    • The Revealed Review: The critic is mild on Section A and harsh on Section B.
    • The Trick: If you just add up the "total badness," the score looks the same. But the focus of the criticism has shifted.
    • Analogy: A parent scolding a child. First, they yell about the messy room. Later, they yell about the homework. The total "yelling" is the same, but the specific problem being attacked changed.

2. The "Language Mismatch" Problem

The paper also tested what happens if the critic speaks English but has to read the painting's notes in Spanish (or vice versa).

  • The Result: It wasn't a simple case of "English is better" or "Spanish is better."
  • The Analogy: Imagine three different critics.
    • Critic A (Grok) is super strict when reading English notes but goes soft when reading Spanish.
    • Critic B (Perplexity) does the exact same thing as Critic A.
    • Critic C (Qwen) does the opposite: strict with Spanish, soft with English.
  • The Takeaway: You can't just pick one language to be "safe." The instability depends entirely on which specific AI model you are using.

3. Why Self-Reported Scores Can Be Misleading

The most surprising finding is about the numbers.

  • The Expectation: If a critic finds fewer errors or changes their facts, their final "Score" (e.g., 3 out of 5) should go down.
  • The Reality: In almost every case of the "Magic Disappearing Act" or "Memory Glitch," the critic gave themselves a perfect score (3/3) even though they had completely changed their story.
  • The Analogy: It's like a student taking a test, getting the answers wrong, changing the answers, and then writing "100%" at the top of the paper. The number says "Perfect," but the work is broken.
  • Conclusion: You cannot trust an AI's self-reported number to tell you if it did a good job. The score is an unreliable indicator of the model's underlying stability; you have to read the actual words it wrote to understand what happened.

4. The Solution: The "Two-Person Rule"

Since one critic can be unreliable, the paper tested what happens if you use two different critics to review the same painting at the same time.

  • The Result: In 5 out of 6 cases, having two different AI models (from different companies) review the same thing made the process much more stable. They caught each other's drifts.
  • The Analogy: If you ask one friend to judge a movie, they might be biased. If you ask two friends from different backgrounds to watch it separately and compare notes, you get a much truer picture.

Summary

The paper concludes that using AI to review other AI is possible, but it's dangerous if you don't have safeguards.

  1. Don't trust the numbers: An AI saying "I found no errors" doesn't mean it's true; it might have just forgotten to look. Self-reported scores are poor indicators of stability.
  2. Watch out for the "Identity" effect: Knowing who the AI is reviewing changes how it writes, often making it less rigorous.
  3. Language matters: The language the AI speaks vs. the language of the text it reads changes its performance in unpredictable ways.
  4. The Fix: Always use two different AI models to review the same thing independently, and have humans check the actual text, not just the scores.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →