← Latest papers
💰 quantitative finance

AI evaluation may bias perceptions: The importance of context in interpreting academic writing

This paper demonstrates that using pooled benchmarks to detect AI-generated text in academic writing introduces significant biases across different countries and fields, and argues that context-specific benchmarks are essential for accurate and equitable evaluation.

Original authors: Shang Wu, Randol Yao

Published 2026-05-27
📖 4 min read☕ Coffee break read

Original authors: Shang Wu, Randol Yao

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a judge at a talent show, and your job is to spot which singers are using a "magic voice box" (AI) to sound perfect. The paper by Shang Wu and Randol Yao argues that the current way we try to spot these magic voice boxes is broken because it ignores the singers' natural backgrounds.

Here is the breakdown of their findings using simple analogies:

The Problem: One Size Does Not Fit All

The researchers found that most current methods use a single "Gold Standard" checklist to judge everyone. This checklist was built by looking at how AI changes text, assuming that all human writing starts from the same place.

The Analogy:
Imagine you are judging a cooking competition. You have a checklist of "suspicious ingredients" that AI chefs might use.

  • The Flaw: You apply this same checklist to a French chef, a Japanese chef, and a Mexican chef.
  • The Reality: The French chef naturally uses ingredients like "butter" and "shallots." The Japanese chef naturally uses "soy sauce" and "miso."
  • The Mistake: If your checklist says, "If you use butter, you are cheating," you will unfairly accuse the French chef of using AI, even though they have been using butter for centuries. Meanwhile, you might miss the Mexican chef who uses a different set of ingredients that your checklist doesn't flag.

In the paper, the "ingredients" are specific words (like "notably," "utilizing," or "fostering"). The study shows that some countries and academic fields naturally use these words all the time, long before AI existed.

The Experiment: The "Time Travel" Test

To prove this, the researchers looked at scientific papers from 2021 (before AI writing tools were common).

  • The Expectation: Since no one was using AI yet, the "AI detector" should have found zero AI usage.
  • The Result with the Old Method (The Pooled Benchmark): The detector went crazy. It claimed that researchers in the US and UK were using AI heavily, while researchers in Russia or Brazil were using almost none.
  • The Reality: This was a false alarm. The detector was just confused by the fact that American and British scientists naturally write in a style that looks like AI. It was mistaking a "British accent" for a "robot voice."

The Solution: Customized Checklists

The researchers created a new method: Country-and-Field Specific Benchmarks.
Instead of one giant checklist for everyone, they made 234 tiny, custom checklists.

  • One checklist for "Math in the US."
  • One checklist for "Economics in China."
  • One checklist for "Psychology in Germany."

The Analogy:
Now, instead of asking the French chef, "Did you use butter?" (which is normal for them), the judge asks, "Did you use extra butter compared to what you usually use?"

  • If the French chef uses their normal amount of butter, the judge says, "Clean."
  • If the French chef suddenly uses double the butter, the judge says, "Suspicious."

The Results: Who Gets Unfairly Blamed?

When they applied these custom checklists to papers from 2025 (when AI was actually being used), they found that the old method was creating a distorted map of the world:

  1. Overestimation (False Accusations): The old method made it look like researchers in English-speaking countries (like the US and UK) and fields like Economics and Humanities were using AI the most.
    • Why? Because their natural writing style already sounded "AI-like" to the generic detector.
  2. Underestimation (Missed Detection): The old method made it look like researchers in Russia, Brazil, Japan, and Eastern Europe, and fields like Math and Physical Sciences, were using very little AI.
    • Why? Because their natural writing style was so different from the "AI style" that the detector didn't recognize the AI even when it was there.

Why This Matters

The paper concludes that if we keep using the "One Size Fits All" method, we are going to:

  • Unfairly punish scientists from English-speaking countries or social sciences by making them look like they are cheating.
  • Let others off the hook who might actually be using AI, simply because their writing style is different.
  • Create inequality: It reinforces the idea that some groups are "less original" than others, not because they are, but because the measuring tape is broken.

The Bottom Line:
To fairly judge if a scientist is using AI, you can't just look at the words they use. You have to understand who they are and what they usually write. You need to know if they are naturally a "butter chef" before you accuse them of using a magic voice box.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →