When Writing Style Drifts: Benchmarking Authorship Verification under Distribution Shifts in Genre, Time and the AI-Era
This paper introduces AVShift, the first German benchmark for systematically evaluating authorship verification under distribution shifts in genre, time, and the AI era, revealing that fine-tuned LLMs generalize best across genres while temporal drift significantly degrades performance, though no measurable AI-era shift was detected.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Every person leaves a unique trail of habits in the way they write. Just as a fingerprint is a physical mark that identifies a body, a writer's "idiolect" is a collection of linguistic habits—how they use punctuation, how long their sentences are, and which words they favor—that identifies their mind. For decades, experts in literature and law have relied on these patterns to solve mysteries, from catching plagiarists to identifying the anonymous author of a threatening letter. The core idea is simple: while we might change our vocabulary depending on whether we are writing a love letter or a business report, the deep, underlying rhythm of our writing remains constant enough to be recognized.
However, this assumption faces a serious challenge in the modern world. Writing does not happen in a vacuum; it shifts depending on the genre, the time period, and the tools we use. A person might write a short, informal forum post in the morning and a long, polished story in the evening. They might write differently today than they did twenty years ago. Furthermore, the recent explosion of artificial intelligence tools that help people write raises a new question: does using a computer to help draft a sentence change the writer's unique fingerprint enough to hide their identity? To answer these questions, researchers needed a way to test if these digital "fingerprints" hold up when the context changes.
A team of researchers at the University of Technology Nuremberg has tackled this problem by creating a massive new test set called AVShift. Instead of looking at a single type of writing, they gathered over 150,000 pairs of texts from a German website dedicated to fan fiction, reviews, and forum discussions. This collection spans twenty-one years, from 2004 to 2025, capturing writing before and after the public release of modern AI writing assistants. The researchers used this data to see if computer programs could still correctly identify that two different texts were written by the same person, even when those texts were from different genres, written years apart, or created during the era of AI assistance.
The team tested three different types of computer methods to solve this puzzle. The first type relied on handcrafted rules, where experts manually selected specific writing features for the computer to count. The second type used mathematical models that learned to compress text into dense digital summaries. The third and most recent type used large language models, the same kind of powerful AI systems that can write essays and code, which were trained specifically to answer the question: "Did the same person write both of these?"
The results revealed that the type of writing matters immensely. When the computer tried to match texts from the same genre, such as two forum posts, it performed very well. However, when the task required matching a forum post to a story or a review, the performance of most methods dropped significantly. The most successful approach was the large language model trained on a mix of all three writing styles. By exposing the computer to a wide variety of writing voices during its training, it learned to ignore the superficial differences between genres and focus on the deeper, consistent habits of the author. This model achieved a high success rate even when comparing a story to a forum post, proving that a diverse training diet makes the system much more robust.
Time, however, proved to be a much stronger obstacle than genre. The researchers found that as the time gap between two texts grew, the ability to verify the author's identity steadily declined. The longer the wait between when a person wrote two pieces, the harder it became for the computer to link them. This decline was noticeable even after just one year, suggesting that our writing style is not a static fingerprint but a living thing that evolves continuously. The drop in accuracy was significant enough that for high-stakes situations, like legal investigations, comparing documents written many years apart without special adjustments would likely lead to errors.
Surprisingly, the arrival of artificial intelligence did not cause the disruption many feared. The researchers looked closely at texts written before and after the widespread adoption of AI writing tools to see if the "AI era" created a new, confusing distribution of writing styles. They found no evidence that the use of these tools created a systematic shift that made verification impossible. While performance varied between different time periods, the changes did not follow a pattern that could be blamed on AI assistance. The variation seemed to stem from other factors, such as the specific topic or the length of the text, rather than the mere presence of AI tools.
The study also examined the specific habits that make up a writer's style. They discovered that while some features, like the use of common function words, remained stable across different genres, others changed drastically depending on the type of writing. There was no single set of "magic" features that worked for every situation; instead, the most reliable indicators depended on which genres were being compared. This suggests that to build a truly reliable system, one must understand that style is flexible and that the rules for identifying an author change depending on the context.
Ultimately, this work shows that while writing styles are fluid and change with time and genre, they are not impossible to track. The key to success lies not in finding a single unchangeable rule, but in training systems to recognize the author's voice across a wide spectrum of conditions. By using diverse training data and acknowledging that time erodes these signals, researchers can build tools that are far more reliable for real-world applications. The findings offer a clear path forward: to verify authorship in a complex world, we must teach our tools to expect change, rather than assuming the writer's style will remain frozen in time.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.