← Latest papers
🤖 AI

Measuring Embedding Sensitivity to Authorial Style in French: Comparing Literary Texts with Language Model Rewritings

This paper investigates how well French language model embeddings capture and retain authorial stylistic features after rewriting, demonstrating that these signals persist and exhibit model-specific patterns, thereby offering new avenues for detecting authorship imitation.

Original authors: Benjamin Icard, Lila Sainero, Alice Breton, Evangelia Zve, Jean-Gabriel Ganascia

Published 2026-05-12
📖 5 min read🧠 Deep dive

Original authors: Benjamin Icard, Lila Sainero, Alice Breton, Evangelia Zve, Jean-Gabriel Ganascia

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Question: Can AI "Fake" a Writer's Soul?

Imagine you have a group of famous chefs (like Proust, Céline, and Yourcenar). Each has a unique way of cooking: one uses long, complex recipes with fancy ingredients; another cooks fast, loud, and messy; the third is precise and historical.

Now, imagine a super-smart robot chef (an AI or Large Language Model) that tries to copy these human chefs. The robot is told: "Cook this specific dish (a bus ride in Paris), but make it taste exactly like Chef Proust."

The big question this paper asks is: If we take a bite of the robot's food, can we still taste the human chef's unique "flavor," or does the robot's own mechanical taste take over?

In the world of computers, "taste" is measured using embeddings. Think of embeddings as a digital fingerprint or a GPS coordinate for a piece of text. Every sentence gets a specific set of numbers that tells the computer where it lives in a giant map of language.

The Experiment: The "Bus Ride" Test

The researchers set up a controlled experiment using French literature:

  1. The Base Recipe (The Topic): They started with a book by Stéphane Tufféry that tells the exact same story (a bus ride in Paris) but written in many different styles. This ensures the story is the same, so any differences are purely about style.
  2. The Human Chefs: They took texts from three real French authors:
    • Proust: Known for long, flowing, complex sentences.
    • Céline: Known for short, punchy, spoken-sounding sentences.
    • Yourcenar: Known for balanced, historical, and precise writing.
  3. The Robot Chefs: They asked three different AI models (GPT, Mistral, and Gemini) to rewrite the "Bus Ride" story, trying to copy the style of Proust, Céline, and Yourcenar.

How They Measured the "Flavor"

The researchers didn't just ask humans to read the texts. Instead, they used a mathematical map (called UMAP) to plot where all these texts landed.

  • The Human Map: When they plotted the real human authors, their texts formed distinct, separate islands. Proust's texts were in one cluster, Céline's in another, and Yourcenar's in a third.
  • The Robot Map: When they plotted the AI-generated texts, they wanted to see if the robots managed to land on the same islands as the humans, or if they drifted off into their own "Robot Zone."

They measured dispersion, which is like measuring how spread out a group of people is in a room.

  • If the AI successfully copied the style, the AI texts should be clustered tightly with the human texts (low dispersion).
  • If the AI failed, the AI texts would scatter away from the humans (high dispersion).

What They Found

1. The Robots Can Mimic, But Not Perfectly
The study found that the AI models did capture some of the human authors' styles. The "Robot Proust" texts did land closer to "Real Proust" than to "Real Céline." However, the connection wasn't perfect. The AI texts were like a slightly blurry photocopy of the original; the "fingerprint" was there, but it was fainter.

2. Different Robots Have Different "Hands"
Just like how a human chef might be better at baking than frying, different AI models were better at copying different authors:

  • Mistral was the best at copying Proust's complex style.
  • GPT was surprisingly good at capturing Céline's messy, spoken style.
  • Gemini did well with Yourcenar's historical precision.

3. The "Taste" Changes Depending on Who You Ask
The researchers noticed something interesting: The way they measured the "flavor" mattered.

  • If they looked at surface features (like how many commas or short words were used), one robot looked like the best copy.
  • If they looked at deeper structural features (like the complexity of the sentence structure), a different robot looked like the best copy.

It's like judging a painting: If you only look at the colors, one artist wins. If you look at the brushstrokes, a different artist wins. The AI models preserved different parts of the human style depending on which part of the "fingerprint" you were checking.

The Conclusion

The paper concludes that AI embeddings are sensitive to human style, but the signal gets weaker and distorted after rewriting.

Think of it like a game of "Telephone."

  • Round 1: A human whispers a story to another human. The message stays mostly the same, but the voice changes slightly.
  • Round 2: A human whispers to a robot, and the robot whispers back. The robot gets the gist of the story, but the unique "voice" of the original human gets muddled.

The study shows that while AI can mimic the look of a famous author, the deep mathematical "soul" of that author's writing is only partially preserved. This means that in the future, we might be able to detect if a text was written by a human or an AI by looking at how "fuzzy" or "distorted" these style fingerprints are.

Important Note: The paper does not claim this technology is ready to catch AI cheaters in real-time or to be used in court. It is a preliminary study showing that the "style signal" exists but is fragile. The authors suggest this is just the beginning of understanding how AI handles human creativity.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →