Authorship Verification of Transcribed German-Language Videos
This paper addresses the underexplored challenge of authorship verification for spoken German by evaluating ten methods on video transcripts, finding that traditional character- and token n-gram approaches outperform modern transformer-based models with up to 88% accuracy.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to solve a mystery, but instead of looking for fingerprints or DNA, you are looking for the invisible "fingerprint" of a person's writing style. This field is called Authorship Verification. Think of it like a literary lie detector test: you have two pieces of text, and you have to decide, "Did the same person write both of these?" It's a bit like trying to tell if two paintings were made by the same artist just by looking at the brushstrokes, without knowing what the paintings are actually of. While this has been studied for written letters and essays for years, a big question remained: Does this work for spoken words? And does this work for languages other than English? This matters because in the real world, people talk more than they write. If we can figure out who is speaking just by how they talk (even if we can't hear their voice, only the words they chose), it could help catch fraudsters, verify online identities, or even spot academic cheaters. But here's the catch: when people speak, they often talk about specific topics (like fixing a sink or solving a math problem). If a computer just guesses "this person likes math," it's not really checking who they are, just what they are talking about. So, the real challenge is to strip away the topic and see if the unique "voice" of the speaker remains.
This paper dives into that exact challenge, but with a twist: it focuses on German-language videos. The researchers, Oren Halvani and Sophie Titze, wanted to see if the old-school tricks for spotting writers work on transcripts of people talking in videos. They gathered 300 videos from 150 different speakers, covering three very different worlds: financial advice, math tutoring, and do-it-yourself (DIY) home repair. To make sure the computers weren't just memorizing the topic (like "this person always talks about oil changes"), they used a clever tool called POSNoise. Imagine this as a magic eraser that removes all the "meaty" words (nouns, verbs, specific facts) and leaves behind only the "glue" words (like "the," "and," "but") and the punctuation. It's like taking a recipe and removing all the ingredients, leaving only the instructions on how to mix them. If the computer can still guess who wrote the recipe just by the mixing instructions, that's a true style match.
They tested ten different computer methods, ranging from simple, old-school counting tricks to fancy, modern AI models (called transformers) that usually read books and write essays. The results were a bit of a plot twist. The "fancy" AI models, which are usually the stars of the show, stumbled badly. They performed poorly, often getting confused when the topic was removed. It's as if the AI models were so used to reading about specific things that when you took the story away, they had nothing left to go on. In fact, none of the modern AI models managed to get an accuracy score higher than 75%.
On the other hand, the traditional methods—the ones that just count how often certain small groups of letters or words appear together—were the real heroes. The best performer, a method called COAV, got it right up to 88% of the time on the DIY videos, and another method, LambdaG, reached 90% in terms of a statistical score called AUC. The paper suggests that these older methods work better because they focus on the tiny, unconscious habits of speech: how a speaker uses commas, where they put their "and"s and "but"s, and their specific rhythm. These habits survive even when you erase the topic. The researchers found that when they tried to use the fancy AI models, they often failed because those models rely too much on the content of the speech, which was the very thing they tried to hide.
The authors are careful to say that their results are based on these specific experiments with German monologue videos and might not work exactly the same way for every language or every type of conversation (like a chaotic debate with many people talking at once). They also note that their dataset was relatively small, so while the traditional methods look very promising, the gap between them and the AI models is a strong suggestion rather than a final, unchangeable law of physics. However, the message is clear: in the world of verifying who is speaking from a transcript, sometimes the simplest tools, which pay attention to the small, boring details of grammar and word choice, are still the sharpest detectives in the room. The modern, complex AI models, despite their reputation, seem to need a little more training to handle the messy reality of spoken language without getting distracted by the topic at hand.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.