← Latest papers
💬 NLP

Geometry-Aware Localized Watermarking for Copyright Protection in Embedding-as-a-Service

The paper proposes GeoMark, a geometry-aware localized watermarking framework for Embedding-as-a-Service that resolves the robustness-utility-verifiability trade-off by decoupling localized triggering from centralized attribution to ensure copyright protection against various attacks while preserving downstream utility.

Original authors: Zhimin Chen, Xiaojie Liang, Wenbo Xu, Yuxuan Liu, Wei Lu

Published 2026-08-05
📖 4 min read☕ Coffee break read

Original authors: Zhimin Chen, Xiaojie Liang, Wenbo Xu, Yuxuan Liu, Wei Lu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the internet as a giant, bustling library where the most valuable books aren't made of paper, but of "meaning." These are digital maps that turn words, images, and ideas into lists of numbers, allowing computers to understand that a "cat" is similar to a "kitten" but different from a "toaster." Companies build these super-smart maps, called embedding models, and offer them as a service, like renting a high-powered microscope. But here's the catch: just as a thief could sneak into a library, copy the books, and sell them as their own, a hacker can trick these digital microscopes into revealing their secrets. By asking the service millions of questions and recording the answers, a thief can build a fake, cheap copy of the original model. This is called "model stealing," and it's a huge problem because it steals the hard work and money of the people who built the original tools. To stop this, scientists have tried to hide invisible "watermarks" in the answers, like a secret stamp on a banknote. But so far, these stamps have been tricky: some wash away if you rewrite the question, others break if you change the size of the answer, and some accidentally stamp innocent people by mistake.

This paper introduces a new, clever way to stamp these digital maps called GeoMark. Think of the digital map not as a flat sheet of paper, but as a giant, bumpy 3D landscape where similar ideas are close together and different ideas are far apart. Previous methods tried to stamp the whole landscape or specific flat zones, which often failed when the landscape was squished or stretched. GeoMark takes a different approach. Instead of stamping everywhere, it picks a few special "anchor" spots on the landscape that are far away from a secret "target" spot. It only stamps the area immediately around these anchors. If a thief tries to steal the model, the stamp is hidden in these specific neighborhoods. Even if the thief tries to rewrite the questions (paraphrasing) or chop off parts of the answer (dimensional perturbation), the stamp stays because it's tied to the shape of the neighborhood, not the specific words or numbers.

The researchers tested this idea on four different types of data, ranging from news articles to email spam. They found that GeoMark is like a super-durable stamp. It survived attacks where the thief tried to rewrite the stolen data using advanced AI tools, and it even survived attacks where the thief tried to shrink or shift the data's dimensions. Most importantly, it didn't accidentally stamp innocent, clean data. In the experiments, the system correctly identified the stolen models with a very high degree of confidence (a statistical value called a p-value less than 0.05), while rarely making a mistake on models that hadn't been stolen. The paper suggests that by separating where the stamp is applied from what the stamp looks like, they created a system that is both tough to break and easy to verify.

The team also checked if this new method slowed down the service. They found that adding the stamp took almost no time at all—about 0.017 milliseconds per question—which is so fast it's practically invisible. They also tested how changing the number of "anchor" spots or the size of the stamped neighborhoods affected the results. They found that having a few anchors (around 5) and covering a small, specific percentage of the area (about 4%) worked best, keeping the service fast and the stamp strong.

In short, the authors propose that by using the natural shape of the data landscape to hide the watermark in specific, scattered neighborhoods, they can protect digital models from thieves much better than before. The paper suggests that this "geometry-aware" approach solves the old problem where watermarks were either too fragile or too prone to false alarms. While the results are based on computer simulations and tests on specific datasets, the evidence points to a method that keeps the service useful for everyone while making it very hard for thieves to steal the technology without getting caught.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →