Distinguishing Artificial from Authentic: Evaluating LLMs for Detecting LLM-Generated Content
This paper evaluates the capability of large language models to detect their own generated content across educational tasks, finding that while detection is effective for programming and longer reflective writing, it remains unreliable for short-answer questions and is highly sensitive to prompt variations, suggesting caution in using LLMs as standalone tools for identifying AI-generated student work.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to spot a fake painting in a gallery. You have a special pair of glasses that claims to tell you instantly if a masterpiece was painted by a human hand or by a robot. This is the world of Large Language Models (LLMs). Think of these models as super-smart robots that have read almost everything ever written on the internet. They can write essays, solve math problems, and even write computer code that looks just like it came from a human student. But here is the tricky part: as these robots get better, it becomes harder to tell their work from real human work. This creates a big problem for schools and teachers who need to know if a student actually did their homework or just asked a robot to do it for them. Scientists are currently racing to find a way to catch these "robot essays," and one of the most interesting ideas is to ask the robots themselves: "Hey, did you write this?" It's like asking a forger to look at a fake painting and admit, "I made that."
This paper is a report from a team of researchers who decided to test this "ask the robot" idea in a real classroom setting. They wanted to see if an AI could act as a detective to spot its own work. They set up a digital experiment using a university programming course where students had to do four different types of tasks: writing a long reflection about learning a new language, answering a short one-sentence question about code, and solving two different coding puzzles. The researchers then asked an AI to generate fake answers for all these tasks, trying to make them look like they came from a student (some with typos, some with "beginner" mistakes). Finally, they asked a second AI to look at all these answers—both the real student ones and the fake AI ones—and guess which was which.
The results were a bit like a rollercoaster ride. The researchers found that the AI detective was very good at its job in some areas but completely confused in others. When the task was writing code, the AI was a sharp-eyed detective. It could easily tell the difference between a robot's code and a human's code, mostly because the robot's code tended to be longer and more perfectly structured, like a robot that never makes a typo. Similarly, when students wrote long, reflective essays (several paragraphs long), the AI detective could spot the fake ones because the robot's writing was too polished and wordy, like a speech that sounds too rehearsed.
However, the AI detective hit a wall with short answers. When the task was to write just one sentence about why a piece of code was easy to read, the AI got it backwards. It actually thought the real human students' short answers were more likely to be written by a robot than the robot's own answers! The researchers suggest this happened because the robot's short answers were often deliberately made to look "human" by adding typos or using informal language, which accidentally made them look less suspicious to the detector. Meanwhile, the real students' short answers were sometimes so clear and direct that the detector mistakenly thought, "This is too perfect to be human; it must be a robot!"
The study also discovered that the "costume" the robot wore mattered a lot. If the robot was told to act like a student and add typos or use slang, it became much harder to catch in writing tasks. But in coding tasks, even if the robot tried to act like a clumsy beginner by adding bugs, it was still pretty easy to spot. The researchers conclude that while asking an AI to catch other AIs might work for long essays or coding, it is not a magic bullet. It is too unreliable for short answers, and a teacher should never rely on it alone to accuse a student of cheating. It's a tool that works in some rooms of the house but fails miserably in others.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.