ARB: A Matched Authorship-Rewriting Benchmark Dataset for AI-Text Detector Evaluation
This paper introduces the Authorship-Rewriting Benchmark (ARB) to demonstrate that standard AI-text detectors, while effective at identifying direct LLM generation, suffer a severe performance drop (60–78 percentage points) when detecting human-authored text that has been rewritten by an LLM, revealing a critical gap in current evaluation methodologies.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to catch a forger. In the world of writing, the "forger" is an Artificial Intelligence (AI) that can write stories, essays, and news articles that look and sound just like a human. For a while, the police (the scientists who build detectors) have been testing their tools by handing them two piles of papers: one pile written by real people and another pile written directly by the AI. If the detector can tell them apart, it gets a gold star. But here is the twist: in the real world, people rarely just copy-paste AI text. Instead, they often write a draft themselves and then ask the AI to "fix it up," "make it sound better," or "rewrite this paragraph." This is like a human artist sketching a picture, then asking a robot to add the final polish. The big question is: if the detector is great at spotting the robot's raw drawing, will it still work when the robot is just polishing a human's sketch?
This paper, titled "ARB: A Matched Authorship-Rewriting Benchmark Dataset for AI-Text Detector Evaluation," dives right into that messy, real-world scenario. The researchers, Gaetano Perrone and Simon Pietro Romano, built a new testing ground called ARB. Instead of just comparing "Human vs. Robot," they created four distinct scenarios to see how detectors handle different mixes of human and robot work. They wanted to know if the detectors that pass the simple tests are actually ready for the complex, polished texts that students, journalists, and writers are actually producing today.
The Four-Regime Game
To understand the experiment, imagine a kitchen where a chef (the human) and a robot (the AI) are making soup. The researchers set up four specific bowls to taste:
- The Human Bowl (HUMAN): The chef makes the soup from scratch. This is the baseline.
- The Robot Bowl (FREE-LLM): The robot makes the soup from scratch, using only its own ingredients. This is the standard "AI text" everyone tries to catch.
- The Chef-Robot Bowl (H2L): The chef starts the soup, but then the robot comes in and rewrites the recipe, changing the words and the style while keeping the same flavor. This is a human draft polished by AI.
- The Robot-Robot Bowl (LLM2L): The robot makes the soup, and then the same robot comes back and rewrites its own recipe. This is AI text that has been polished by AI.
The clever part of this study is that they used the exact same starting ingredients (the same human drafts or robot drafts) for the rewriting steps. This way, they could compare the "Chef-Robot" soup directly against the "Robot-Robot" soup to see if the origin of the soup mattered, or if it was just about how much the robot changed the words.
The Shocking Results
When the researchers tested five different "detective" tools on these soups, the results were a bit of a plot twist.
First, on the standard test (Human vs. Robot Bowl), the detectives were doing great. Two of the top detectors, FastDetectGPT and Binoculars, caught about 91% to 94% of the raw robot soup. They were like sharp-eyed guards spotting a robot from a mile away.
But then, they tried the Chef-Robot Bowl (the human draft polished by AI). Suddenly, the detectives went blind.
- FastDetectGPT dropped from catching 91.2% of the raw robot text to catching only 30.8% of the human-draft-polished-by-AI text.
- Binoculars fell even harder, from 93.5% down to just 15.1%.
This is a massive drop of 60 to 78 percentage points. It's as if the detective could spot a robot in a red hat, but the moment the robot put on a human's blue hat, the detective couldn't tell the difference at all.
However, when they tested the Robot-Robot Bowl (where the robot polished its own work), the detectives stayed mostly awake.
- FastDetectGPT only dropped a little, from 91.2% to 78.3%.
- Binoculars went from 93.5% to 83.0%.
This suggests that when a robot polishes its own work, it leaves behind a "robot fingerprint" that is still easy to spot. But when a robot polishes a human's work, that fingerprint gets wiped away, and the text becomes very hard to distinguish from a real human.
What This Means
The paper argues that the current way we test AI detectors is misleading. If a detector looks great on the simple "Human vs. Raw Robot" test, it doesn't mean it will work in the real world where humans use AI to help them write. The standard tests are overestimating how good these tools really are.
The researchers found that this isn't just because the robot changed the words more in the human-draft scenario (though it did change them more). The gap between the two scenarios is an operational one: the detectors perform differently based on whether the text started with a human or a machine, but this difference is also mixed with the fact that the robot changed the human text more aggressively than it changed its own text. The study cannot prove that the "origin" alone is the sole cause of the drop in performance, but it clearly shows that the standard tests fail to predict how detectors will behave on human drafts that have been polished by AI.
The Bottom Line
The study concludes that we need to stop relying on simple tests. If you are a teacher, a boss, or a journalist trying to figure out if a text was written by a human or an AI, you can't just trust a detector that says "95% accurate" based on old tests. If that text was a human draft that got a little AI help, the detector might be wrong most of the time.
The paper suggests that for detectors to be truly useful, they need to be tested on these "polished human drafts" specifically, not just on raw robot text. Until then, we have to be careful not to accuse innocent humans of being robots just because a tool got confused by a little bit of AI help. The authors are clear: this isn't a solved problem, and the current tools are not as robust as we thought they were.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.