← Latest papers
💻 computer science

Benchmark Evaluation of the WordBinary AI-Text Detection Model on 9,000 AI-Generated Texts and 2,000 Pre-2015 Human Academic-Paper Samples

This study reports that the WordBinary AI-text detection model achieved perfect classification accuracy (100%) on a specific benchmark of 11,000 samples comprising 9,000 AI-generated texts and 2,000 pre-2015 human academic papers, while acknowledging that these results require further independent validation and broader testing before generalizing to universal accuracy.

Original authors: Kiah Curan

Published 2026-08-04
📖 5 min read🧠 Deep dive

Original authors: Kiah Curan

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the internet as a giant, bustling library where everyone is writing stories, essays, and reports. For a long time, we knew who wrote what: a human sat down with a pen or a keyboard and created the words. But recently, a new kind of "ghost writer" has moved into the library. These are Artificial Intelligence (AI) systems, powerful computer programs that can read millions of books and then write their own stories that sound almost exactly like a human wrote them. This has created a bit of a mystery: if you pick up a new book, how do you know if a human wrote it or if a robot did?

This is the world of "AI-text detection." Think of it like a super-smart librarian trying to spot the difference between a handwritten letter and a printout from a machine. The goal is to catch the robot-written stories without accidentally accusing the real humans. If the librarian is too eager, they might shout "Fake!" at a real person's homework, which is unfair and causes trouble. If they are too slow, they might miss the robot entirely. Scientists are racing to build detectors that are sharp enough to spot the AI but gentle enough to protect the humans.

Now, let's meet the new detective in town: a system called WordBinary. A researcher named Kiah Curan decided to put this detective through a massive test to see how good it really is. They didn't just test it on a few sentences; they threw a mountain of text at it. The test involved 11,000 different pieces of writing. To make it a fair fight, they split the mountain into two piles.

The first pile was the "Robot Pile." It contained 9,000 texts that were definitely written by AI. But here's the tricky part: these weren't just plain, boring robot texts. Some were written by famous AI models like OpenAI, Claude, and Gemini. Even more challenging, some of these robot texts had been "humanised." Imagine a robot writing a story, and then a human (or another robot) going through it with a red pen, changing the words, fixing the grammar, and trying to make it sound more natural. The test included both the raw robot texts and these "humanised" versions to see if the detective could still spot them.

The second pile was the "Human Pile." It contained 2,000 samples of real academic papers written by humans. To make sure these were truly old-school human writing, the researcher only picked papers published before 2015. Why 2015? Because that was before the modern, super-powerful AI writing tools became popular. These papers came from five different fields: computer science, economics, statistics, physics, and mathematics.

So, what happened when WordBinary looked at this mountain of 11,000 texts? The results were nothing short of perfect.

The detective got every single one right. It correctly identified all 9,000 AI texts as being written by a machine. It didn't miss a single one, even the ones that had been "humanised" to sound more like a person. At the same time, it correctly identified all 2,000 human academic papers as being written by humans. It didn't accidentally accuse a single human writer.

In the world of statistics, this is a huge deal. The researcher calculated that the system was 100% accurate on this specific test. It caught every robot and spared every human. Even the "humanised" robot texts, which are usually the hardest to catch, were spotted 100% of the time. The system gave the robot texts very high "AI scores" (mostly between 98% and 100%), while the human texts got very low scores (mostly below 1%). The highest score a human paper ever got was just 1.07%, which is like a whisper of a doubt, not a shout of guilt.

However, the researcher is very careful not to say this is the end of the story. While WordBinary won this specific game perfectly, the researcher warns us not to think it can win every game forever. The test was like a practice match on a perfect field with known players. The researcher points out that we don't know if WordBinary was trained on these exact same texts before, which would be like a student memorizing the answers to a test. We also don't know how it would handle texts in other languages, or if it would get confused by a human who used AI just to fix their spelling.

The paper concludes that WordBinary is incredibly promising and performed flawlessly on this specific set of 11,000 samples. But just like a detective who solves one big case, it still needs to prove it can solve many different kinds of mysteries in the real world before we can trust it to make final decisions on its own. For now, it's a very strong tool, but it's not a magic wand that solves everything.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →