← Latest papers
🤖 machine learning

Generative AI Training and Copyright Law

This paper argues that current legal justifications for generative AI training, such as "fair use" in the USA and "Text and Data Mining" exceptions in Europe, are fundamentally flawed because AI training differs from TDM and involves independent copyright risks through data memorization, while also proposing the ISMIR's role in establishing fair practices for all stakeholders.

Original authors: Sebastian Stober, Tim W. Dornis

Published 2026-03-18
📖 5 min read🧠 Deep dive

Original authors: Sebastian Stober, Tim W. Dornis

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Picture: The Great AI Feast

Imagine the world of music, art, and writing as a massive, endless buffet. For decades, chefs (human creators) have been cooking up delicious dishes.

Now, a new kind of chef has arrived: Generative AI. This robot chef wants to learn how to cook. To do this, it doesn't just taste a few bites; it swallows the entire buffet whole. It eats millions of songs, books, and paintings to learn the "flavor" of human creativity.

The paper asks a very serious question: Is it legal for the robot to eat the whole buffet without asking the original chefs for permission?

The authors, Sebastian Stober and Tim W. Dornis, say: "Probably not." They argue that the robot isn't just "studying" the food; it's trying to become a copycat chef who can serve up dishes that taste exactly like the originals.


1. The "Study Hall" vs. The "Copycat Kitchen"

For a long time, companies building AI have argued: "We aren't stealing; we are just studying. We are doing 'Text and Data Mining' (TDM)."

  • The Old Analogy (TDM): Imagine a student in a library. They read a book to find out how many times the word "love" appears, or to analyze the sentence structure. They take notes, but they don't write a new book that looks like the original. This is usually allowed by law.
  • The New Reality (GenAI): The AI isn't just taking notes. It's memorizing the entire book, the entire melody, and the entire style, so it can write a new song that sounds exactly like the original artist.

The Authors' Point: The law has exceptions for "studying" (TDM) and "fair use" (like parody or criticism). But the authors argue that training an AI to generate new content is different from studying. It's like the difference between a student taking a math test (analysis) and a forger trying to replicate a Picasso painting so well that it fools the gallery (generation). The law might not cover this new kind of "forgery."

2. The "Photographic Memory" Problem (Memorization)

Here is the scariest part. Even if the AI is "learning" the rules of music, it has a bad habit: it has a photographic memory.

  • The Analogy: Imagine you teach a parrot to talk by reading it a million books. Eventually, the parrot doesn't just learn how to speak; it starts reciting specific sentences from those books word-for-word.
  • The Legal Risk: If you ask the AI to "write a song in the style of Artist X," it might accidentally spit out a song that is 90% identical to a real song by Artist X.
  • Why this matters: The authors explain that this "memorization" is a legal problem on its own. Even if the AI was allowed to eat the data to learn, if it vomits (regurgitates) a copyrighted song back out, that is copyright infringement.

Real-world example: The paper mentions a lawsuit where a music company sued an AI because the AI was singing lyrics that were almost identical to real songs. The court in Germany recently said: "Yes, the AI memorized the lyrics, and that is illegal."

3. The "Recipe Book" Dilemma

So, what do we do? The authors suggest a new way forward called Documentation.

  • The Current Situation: AI companies often say, "We scraped the internet." It's like a chef saying, "I used ingredients from everywhere, but I don't know which farm the tomatoes came from."
  • The Proposed Solution: The authors want AI developers to keep a detailed recipe card for every model they build.
    • Basic Level: "We used these websites."
    • Intermediate Level: "We used these specific songs (with ID numbers)."
    • Advanced Level: "We used fingerprinting technology to match every audio clip to its original owner."

Why this helps:

  1. For Artists: They can look at the recipe card and say, "Hey, you used my song! I didn't agree to that. Take it out!"
  2. For AI Companies: If they have a clean recipe card, they can prove they tried to be fair.
  3. For Researchers: It stops the "black box" mystery. We can finally see exactly what the AI learned.

4. The "Music Research" Superpower

The paper is written for the MIR (Music Information Retrieval) community. These are the scientists who build tools to identify songs, find similar melodies, and organize music libraries.

The Call to Action:
The authors say, "You guys are the experts at finding needles in haystacks!"

  • Instead of just letting AI companies scrape music blindly, the music research community should build tools to fingerprint the data.
  • They should create a system where every dataset used to train an AI is tagged with the owner's name and permission status.
  • They should refuse to work with companies that don't follow these rules, just like a chef who refuses to use ingredients from an unlicensed supplier.

Summary: The Three Takeaways

  1. The "Fair Use" Defense is Shaky: Just because an AI is "learning" doesn't mean it's automatically legal to use everyone's copyrighted work without permission.
  2. Memorization is Dangerous: Even if the training is legal, if the AI accidentally spits out a copyrighted song, the company is in trouble.
  3. Transparency is the Key: We need to stop guessing what data AI uses. We need a "nutrition label" for AI models that lists exactly what they ate, so artists can protect their work and researchers can build better, safer systems.

The Bottom Line:
Generative AI is a powerful tool, but it can't just eat the world's creativity for free. To keep the future of music and art healthy, we need to treat the original creators with respect, know exactly what data is being used, and build systems that don't accidentally steal the show.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →