← Latest papers
🤖 AI

Watermarking Should Be Treated as a Monitoring Primitive

This paper argues that watermarking should be treated as a monitoring primitive rather than just a per-sample detection tool, demonstrating that internal and external monitoring capabilities inevitably emerge through signal aggregation and statistical patterns, thereby revealing a fundamental tension between attribution and privacy.

Original authors: Toluwani Aremu, Nils Lukas, Jie Zhang

Published 2026-05-14
📖 5 min read🧠 Deep dive

Original authors: Toluwani Aremu, Nils Lukas, Jie Zhang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: Watermarks Are Like Invisible Ink That Never Dries Up

Imagine you are a baker who puts a tiny, invisible speck of glitter in every loaf of bread you sell. You do this so you can prove later, "Yes, I baked this loaf." This is how AI watermarking currently works: it hides a tiny signal in AI-generated text or images so we can tell if a computer made it.

Right now, experts mostly worry about two things:

  1. Can someone wash the glitter off? (Can an attacker remove the watermark?)
  2. Can someone fake the glitter? (Can an attacker make a loaf look like it has glitter when it doesn't?)

This paper argues that we are looking at the wrong side of the coin. The authors say we should stop thinking of watermarks just as a "stamp of authenticity" and start thinking of them as a surveillance tool.

They argue that even if the glitter is invisible to the naked eye, if you collect enough loaves over time, you can start to see patterns. You can figure out who baked the bread, how often they bake, and even link two different loaves together to say, "These definitely came from the same person."

The Two Types of "Watchers"

The paper introduces two types of people who might be watching the bread (the AI outputs):

1. The Internal Watcher (The Baker's Manager)

  • The Setup: The baker gives every employee a unique, secret recipe for their glitter. The manager has a master list of all these recipes.
  • The Power: If the manager sees a loaf, they can instantly check the recipe list. They know exactly which employee made it.
  • The Paper's Claim: This is unavoidable. If you give every user a unique key (recipe) to watermark their content, the person holding the keys can always track who generated what. It doesn't matter if the watermark is "zero-bit" (just a simple "yes/no" signal) or complex. The unique key leaves a unique statistical fingerprint.

2. The External Watcher (The Street Detective)

  • The Setup: This person doesn't have the secret recipes. They are just standing on the street corner watching people eat bread. They can't see the glitter.
  • The Power: At first, they can't tell who made what. But, if they watch for a long time and collect thousands of loaves from different people, they start to notice subtle patterns. Maybe "Employee A" always puts a slightly larger crumb in the corner, or "Employee B" always uses a specific type of flour texture.
  • The Paper's Claim: Even without the secret keys, an outsider can learn to identify the source just by looking at the patterns left behind by the watermarks. The more bread they collect (the more data they see), the better they get at guessing who made it.

The Experiments: Proving the Point

The researchers tested this idea with real AI models (for text and images) using a "multi-key" setup (where different users have different keys).

  • The Internal Test: They gave a computer access to the keys. Result? It could identify the source almost perfectly, even when there were 16 different "bakers."
  • The External Test: They gave a different computer no keys, only the bread itself.
    • At first, the computer guessed randomly (like flipping a coin).
    • But as they fed it more and more examples (from 100 to 4,000 samples per person), the computer got smarter.
    • Eventually, it could correctly identify the source 70% to 90% of the time, just by learning the subtle statistical "fingerprint" left by the watermark.

Crucial Control: When they removed the watermarks entirely, or when everyone used the same key, the external computer went back to guessing randomly. This proved that the tracking ability came specifically from the unique watermarks, not just the content of the text or images.

The "Dual-Use" Problem

The paper highlights a fundamental tension:

  • Good Use: Watermarks help us prove who made something (provenance) and stop bad actors from faking it.
  • Bad Use (Unintended): The same technology allows for monitoring. It lets someone track a user's activity over time, link their different posts together, and build a profile of their behavior.

The authors call this a "monitoring primitive." Just like a hammer can build a house or break a window, a watermark can verify safety or enable surveillance.

The Takeaway

The paper concludes that we cannot simply say, "This watermark is safe because it's hard to remove." We must also ask: "If someone collects enough of these watermarked outputs, what can they learn about the user?"

If we want to protect privacy, we need to design watermarking systems that don't leave these persistent, trackable fingerprints, or we need to accept that using unique keys for every user will inevitably allow for tracking. The authors argue that current safety tests are missing this big picture and need to change to account for this "aggregation" risk.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →