← Latest papers
💻 computer science

Convergent assessment cannot be policed: a residual-entropy bound on authorship attribution and its consequences for integrity in quantitative disciplines

This paper argues that authorship attribution is structurally impossible for convergent assessment formats like multiple-choice questions because correct responses contain zero residual entropy, rendering current AI-integrity policies based on essay detection unsound for quantitative disciplines and necessitating modality-specific controls.

Original authors: TANZIM ISLAM KHAN

Published 2026-08-04
📖 6 min read🧠 Deep dive

Original authors: TANZIM ISLAM KHAN

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a mystery: "Who wrote this?" In the world of schools and universities, this is the ultimate question of academic integrity. For years, the main tool detectives have used is a special kind of "fingerprint scanner" designed for long, flowing stories—essays. These scanners look for tiny, unique habits in how a person writes, like how they use commas or structure their sentences. But here is the twist: AI chatbots are getting so good at writing these stories that the scanners are struggling. Schools are trying to fix this by applying the same "essay scanner" logic to everything else students do, like math problems, multiple-choice quizzes, and coding tasks.

To understand why this might be a dead end, we need to look at two simple ideas. First, think of "information" as the amount of unique detail in a message. A long essay is like a messy, unique fingerprint; it has tons of tiny details that can tell you who made it. A math answer, like the number "42," is like a perfect, smooth marble. If two people solve the same problem correctly, they both produce the exact same "42." There is no messiness, no unique fingerprint, and no extra detail to inspect. Second, think of "authorship attribution" as the act of using those unique details to prove who did the work. If there are no unique details to begin with, no amount of fancy technology can magically invent them. This paper asks a bold question: What happens when we try to use a fingerprint scanner on a smooth marble?


The Paper's Big Discovery: The "Smooth Marble" Problem

This research article, titled "Convergent assessment cannot be policed," argues that schools are making a fundamental mistake. They are trying to use essay-based rules to police math, science, and multiple-choice tests, but the author says this is impossible. It's not that the AI detectors are bad or that the tools need to be improved; it's that the type of work itself doesn't contain enough information to identify the author.

The author introduces a concept called "residual entropy." Imagine you are baking cookies. If you ask a student to "write a story about a cookie," they might write a poem, a horror story, or a recipe. There are millions of ways to do it, and each version has a unique "fingerprint" of the writer's style. This is high entropy. But if you ask a student to "solve for X" and the answer is "5," there is only one correct way to write that answer. Whether a human genius, a struggling student, or a super-intelligent robot wrote it, the result is the exact same "5." The author calls this zero residual entropy. Once the answer is correct, there is zero variation left to analyze.

The paper proves mathematically that you cannot extract more information than what is actually there. If the answer is just a single letter or number, there is no "signal" to find. Trying to use AI detectors on these formats is like trying to find a fingerprint on a perfectly smooth, blank piece of glass. No matter how powerful your microscope is, you won't find a print because the glass doesn't have one.

The Three Zones of Risk

The author used data from 19 different studies across 13 subjects to map out where different school tasks fall. They discovered three distinct "zones" or regimes:

  1. The Exposed Zone (The Danger Zone): This includes multiple-choice questions, short numerical answers, and standard problem sets. In this zone, AI is incredibly good (often scoring over 96% on multiple-choice), but the answers have zero authorship signal. The paper argues that in this zone, you simply cannot tell if a human or a robot did the work. It's not a matter of "hard to tell"; it's that the question "Who did this?" has no answer.
  2. The Latent Zone (The Waiting Room): This includes things like short numeric problems where AI currently isn't quite good enough to pass easily. These are "safe" only because the AI isn't smart enough yet. But the author warns this is a ticking clock. Once AI gets better at these, the "safety" disappears instantly because there is still no authorship signal to fall back on.
  3. The Detectable Zone (The Safe Zone): This includes long essays, reflective writing, and complex coding projects. Here, the answers are so varied and unique that AI detectors can work. The paper notes that research into detecting AI in code is actually successful because code has enough "fingerprint" details, unlike a simple math answer.

The Shocking Numbers

When the author applied this logic to a typical science and engineering university program, the results were startling. They found that 38% of the final grades in these programs come from tasks in the "Exposed Zone" (like multiple-choice and problem sets). In contrast, only 5% of humanities programs fall into this trap.

If you raise the standard of proof to what is needed for a criminal court case (a very high bar), the number of exposed tasks in science and engineering jumps to 52%. This means that for more than half of a science student's final grade, the school has no way to prove who actually did the work, and no tool exists that can ever do it.

What This Means for Schools

The paper explicitly rules out the idea that we can just buy better software to fix this. The author states clearly that no improvement in detection technology can solve the problem for multiple-choice or math questions because the information simply isn't there to be found. Trying to accuse a student of using unauthorized AI on a math quiz based on "AI detection" is not just ineffective; the author argues it is mathematically impossible to prove.

Instead, the paper suggests schools need to change their strategy entirely:

  • Stop trying to police the unpoliceable: Don't use authorship detectors on multiple-choice or short-answer math. It's a waste of time and leads to unfair accusations.
  • Change the game: If you want to know who did the work, you need to change the task. Use supervised exams (where a teacher watches you), ask students to explain their thinking in person, or design questions that are specific to the local classroom so AI can't just look up the answer.
  • Accept the limits: For the "Exposed Zone," the only real security is preventing the AI from being used in the first place (like in a locked room), not trying to catch it after the fact.

In short, the paper concludes that the current "one-size-fits-all" approach to academic integrity is broken. You can't use a tool designed for essays to police math problems. For a huge chunk of science and engineering education, the question isn't "How do we catch the cheaters?" but rather "How do we design our tests so that the question of unauthorized AI use doesn't even make sense?"

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →