Linguistic Holonomy and Statistical Watermarks: Inner Geometry of Meaning-Preserving Transformations
This paper demonstrates that statistical watermarks for language models are fundamentally vulnerable to meaning-preserving transformations because their detection relies on endpoint semantic similarity rather than the hidden "holonomy" of the transformation path, leading to a precise mathematical identity showing that signal survival depends critically on the specific location of edits rather than just the overall retention rate.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the modern landscape of artificial intelligence, large language models generate text that is increasingly indistinguishable from human writing. To manage the risks of this technology, such as the spread of misinformation, researchers have developed a method to invisibly mark machine-generated content. These marks, known as watermarks, do not change the meaning or quality of the text; instead, they subtly shift the statistical probability of which words are chosen, creating a hidden signal that can be detected later. The central challenge for these watermarks is robustness: they must survive when the text is edited, translated, or paraphrased. For years, the scientific community has measured the success of an attack on a watermark by looking only at the final result. If a rewritten sentence still means the same thing as the original, it was assumed that the watermark had survived just as well, regardless of the journey the text took to get there. This assumption treats the intermediate steps of rewriting as mere scaffolding, irrelevant to the final outcome.
A researcher at the Instituto Superior Técnico and the University of Algarve has challenged this long-held view, demonstrating that the path a text takes is just as important as its destination. By treating the sequence of word choices as a geometric journey, they discovered that two texts can end up with identical meanings while carrying completely different amounts of watermark signal. The researcher proved that the standard method of measuring "semantic similarity"—how close the final meaning is to the start—is blind to a hidden property of the transformation. They found that the watermark's survival depends entirely on the specific arrangement of edits made along the way, not just on how many words were changed or how similar the final text looks. In some cases, an attacker can remove the entire watermark signal while changing only a small fraction of the words, provided those changes are spaced in a specific, rhythmic pattern.
The study begins by re-examining the mathematics behind "linguistic loops," which are chains of transformations that alter the form of a sentence without changing its core meaning. The researcher showed that the previous mathematical tools used to analyze these loops were flawed because they collapsed into a simple count, missing the complex structure of the journey. They replaced this with a more precise geometric description, showing that the transformation of text is like a path traced on a sphere. The distance between the start and end points tells you how far the meaning drifted, but the shape of the path itself carries a hidden rotation, a property known as holonomy. This rotation is invisible to detectors that only look at the final meaning, yet it is the very thing that determines whether the watermark survives. The researcher proved that the endpoint and the path are independent; you can have a text that returns to its original meaning after a short, direct trip or a long, winding detour, and the watermark will behave differently in each case.
On the practical side of the research, the researcher derived a precise law governing how watermarks decay when text is edited. They found that the survival of the signal is not determined by the overall percentage of words that remain unchanged, but by whether the specific "seeding windows" of the watermark remain intact. A watermark often relies on a group of preceding words to decide what the next word should be. If an edit breaks up this group, the signal is lost. The researcher demonstrated that an attacker who knows the size of this group can strategically place edits to destroy the entire watermark while leaving the vast majority of the text untouched. For example, if the watermark relies on a context of one preceding word, an attacker could remove every second word and completely erase the signal, even though half the text remains. This finding explains why watermarks sometimes fail much faster than expected when text is edited in blocks or patterns, rather than randomly.
To verify these theoretical insights, the researcher conducted extensive experiments using real machine translation systems. They took thousands of sentences and sent them through chains of round-trip translations, moving them through languages like German, French, and Spanish and back to English. They measured the watermark signal at every step and compared it against the geometric properties of the path the text took. The results confirmed their theory: the amount of signal remaining was strongly linked to the specific shape of the path, not just the final similarity score. In one striking test, they held the final meaning of the text constant while varying the path length and complexity. They found that texts ending at the exact same meaning could retain anywhere from a full signal to no signal at all, depending solely on the route taken. This proves that the current method of judging watermark robustness based on final similarity is fundamentally incomplete.
The implications of this work are significant for the future of digital provenance. It suggests that the resilience of a watermark is a function of the entire history of the text, not just its final state. The researcher showed that the standard metrics used to evaluate these systems are insufficient because they ignore the geometric structure of language. They also highlighted a vulnerability in current designs: because the survival of the signal depends on the arrangement of edits, a knowledgeable adversary can exploit this to neutralize the watermark without making the text look significantly different. While the study was conducted with specific models and translation chains, the geometric principles they uncovered appear to be universal. The work concludes that to truly understand how watermarks survive, we must stop looking only at the destination and start mapping the journey.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.