← Latest papers
💬 NLP

Probing Stylistic Appropriation using Large Language Models: An Evaluation Framework for Copyright Infringement under EU Law

This paper introduces PSALM, a novel LLM-as-a-judge framework that operationalizes EU copyright doctrine to evaluate stylistic appropriation, revealing that current safeguards focusing solely on verbatim memorization fail to detect systematic infringement in narrative patterns and creative elements generated by fine-tuned models.

Original authors: Noah Scharrenberg, Chang Sun

Published 2026-07-01
📖 4 min read☕ Coffee break read

Original authors: Noah Scharrenberg, Chang Sun

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Problem: Copying the "Vibe" vs. Copying the Words

Imagine you are a chef who learns to cook by reading a famous cookbook.

  • The Old Way of Checking: Current technology acts like a strict librarian. It only checks if you copied a recipe word-for-word. If you changed a few words or rearranged the sentences, the librarian says, "You didn't copy this! You're safe."
  • The Legal Reality: European copyright law is more like a food critic. The critic doesn't just care about the exact words; they care about the style, the flavor, the plating, and the unique way the chef tells the story of the dish. If you take a famous chef's unique "voice" and "method" and serve it as your own, even if you changed the ingredients slightly, the law says you might still be infringing on their copyright.

The Gap: There is a mismatch. The technology says, "No exact words copied, so it's fine." The law says, "But the style and story are stolen."

The Solution: PSALM (The Digital Food Critic)

The authors created a new tool called PSALM (Probing Stylistic Appropriation by Language Models). Think of PSALM as a team of AI food critics designed to act like a judge in a copyright court.

Instead of just counting how many words match, PSALM looks at 10 different "flavors" of a story:

  1. Writing Style: Does the sentence rhythm and vocabulary feel like the original author?
  2. Narrative Voice: Is the story told from the same perspective (e.g., a grumpy old man vs. a curious child)?
  3. Characters: Are the personalities, motivations, and relationships identical?
  4. Plot: Is the sequence of events and the twists the same?
  5. World Building: Are the rules of the world (magic systems, geography, culture) copied?
  6. Defenses: PSALM also checks if the new story is a parody (making fun of the original), a pastiche (a loving tribute), or a quotation (citing the source).

PSALM gives a score from 0 to 10. A high score means the AI has "appropriated" the style, even if it didn't copy the words.

The Experiment: Teaching an AI to Be a Specific Author

The researchers tested this on Llama 3.2, a popular AI model. They taught it (fine-tuned it) using a dataset of questions and answers based on old Dutch literature.

The Results:

  1. Before Training: The AI was like a generalist chef. It had a generic style.
  2. After Training: The AI became a mimic. When asked to write a story, it didn't just remember specific sentences; it completely adopted the personality, voice, and plot structure of the old Dutch authors.
    • Analogy: It wasn't just reciting the cookbook; it started cooking exactly like the famous chef, using the same unique techniques and presentation, even if the ingredients were slightly different.
    • Key Finding: The AI passed the "word-for-word" test (it didn't copy exact sentences), but PSALM gave it a 10/10 on stealing the author's style and story structure.

The "Unlearning" Attempt: Can You Forget a Style?

The researchers then tried to make the AI "unlearn" these specific books using a technique called Negative Preference Optimization (NPO). Think of this as trying to scrub the chef's brain so they forget the specific recipes.

The Results:

  • The Good News: The AI successfully stopped copying the exact words. If you asked it to recite a passage, it couldn't do it anymore. The "verbatim" score dropped to zero.
  • The Bad News: The AI did not forget the style.
    • Analogy: You scrubbed the chef's brain of the specific recipes, but they still cook with the exact same unique rhythm, plating style, and flavor profile as the original chef.
    • PSALM still gave the "unlearned" AI high scores for stealing the author's voice and story structure. The AI was still "too similar" to the original, even though it wasn't copying words.

The Conclusion

The paper concludes that current safety tools are not enough.

  • Current Tools: Only check if you copied the words (like a librarian checking for exact text).
  • The Law: Cares if you stole the style and story (like a critic checking the vibe).
  • The Reality: Even if you "unlearn" the exact words, the AI can still hold onto the style and structure of the copyrighted work.

The Takeaway: To truly respect copyright in the age of AI, we need to measure and stop the theft of style and structure, not just the theft of words. The PSALM tool is a first step toward building a system that can catch these "vibe thieves."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →