Position: No Retroactive Cure for Infringement during Training
This paper argues that post-hoc mitigation strategies like machine unlearning cannot retroactively cure liability for copyright infringement in generative AI because legal responsibility hinges on the unauthorized acquisition and training process rather than the final outputs, necessitating a shift toward verifiable ex-ante compliance.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Core Message: You Can't "Un-Bake" a Cake
Imagine a baker who steals flour, sugar, and eggs from a neighbor to bake a giant, delicious cake. The baker then sells the cake. Later, the neighbor sues. The baker says, "Don't worry! I've installed a special filter on the cake so you can't taste the stolen ingredients, and I've even promised to throw away the recipe book."
This paper argues that this defense doesn't work.
The legal problem wasn't the cake (the AI's output); the problem was the theft (the training data). Once the baker stole the ingredients and mixed them into the batter, the crime was already committed. You can't "un-steal" the ingredients just by changing how the cake tastes later.
The Three Main Arguments
1. The "Finished Crime" Rule (The Completed Act)
The Analogy: Think of copyright infringement like a bank robbery.
If a robber breaks into a bank and steals $1 million, the crime is complete the moment the money leaves the vault. It doesn't matter if the robber later feels guilty, returns the money, or paints the bank to look different. The robbery happened.
The Paper's Point:
AI developers often argue that if they can "unlearn" the stolen data or filter the AI's answers, they are safe. The authors say no.
- The Moment of Theft: When an AI downloads and copies a book or image to train on it, that is the "robbery."
- No Retroactive Cure: Deleting the file later or filtering the AI's output is like the robber returning the money after the trial has started. It might lower the fine, but it doesn't erase the fact that the theft happened. The liability is "locked in" the moment the data was copied.
2. The "Ghost in the Machine" (Model Weights)
The Analogy: Imagine you memorize a secret recipe by reading it once. Even if you burn the physical piece of paper (delete the data), the recipe still exists in your brain (the AI's weights).
If you are a chef who memorized a stolen recipe, you can't say, "I burned the paper, so I didn't steal the recipe." The knowledge is still inside you.
The Paper's Point:
Even if an AI is "unlearned" or filtered, the "ghost" of the stolen data remains in the model's mathematical structure (its weights).
- Fixed Copies: The law says a "copy" exists if it's stored in a stable place (like a hard drive) and can be read by a machine. The AI's brain is that stable place.
- Recoverable: Even if the AI refuses to say "Hello" to a specific user, a hacker could potentially use special tools to extract the stolen information from the AI's brain. Because the information is still technically "there," the infringement hasn't truly vanished.
3. The "Contract Trap" (It's Not Just About Copyright)
The Analogy: Imagine you enter a private club. The sign at the door says, "No photos allowed." You take a photo anyway. Later, you say, "But the photo isn't copyrighted, so I'm fine!"
The club owner says, "You broke the contract by entering and ignoring the sign."
The Paper's Point:
Even if the data isn't copyrighted (or if "Fair Use" might apply), developers can still get sued for breaking contracts.
- Terms of Service: Many websites say, "Do not use our content for AI training." If an AI ignores this and scrapes the data anyway, it's a breach of contract.
- Unfair Competition: If a company spends millions curating a special dataset, and an AI steals it for free to build a competing product, that's "free-riding." It's like someone stealing your blueprints to build a cheaper house next door. You can be sued for this even if the blueprints themselves aren't copyrighted.
The "Head Start" Problem
The Analogy: Imagine a race. One runner cheats by taking a shortcut through a private field. They get to the finish line 10 minutes early.
Later, the judge says, "Okay, you have to walk back to the start line and delete the map you used."
Does that fix it? No. The runner still has the memory of the shortcut and the advantage of having finished early. They still have the "head start."
The Paper's Point:
AI companies gain a massive advantage by training on stolen data. They save millions of dollars and months of time.
- Unjust Enrichment: The law says you can't keep the profits you made from cheating.
- The Remedy: Sometimes, the only fair punishment is to force the company to throw away the entire AI model (the "Head Start" is gone) and start over with clean data. Just deleting the source files isn't enough because the AI has already learned the shortcuts.
The Solution: "Ex-Ante" Compliance (Check Before You Cook)
The paper concludes that we need to stop trying to "fix" the mess after the AI is built. Instead, we need to check the ingredients before we start baking.
For Engineers (The Chefs):
- The Glass Pipeline: Instead of a black box, build AI systems where you can prove exactly where every ingredient came from. Use digital tags (like a receipt) to show you have permission for every piece of data.
- Git for Models: If you accidentally use a stolen ingredient, don't try to "unlearn" it. Instead, go back to the version of the model before you added that ingredient and start fresh from there.
For Policymakers (The Club Owners):
- Data Trusts: Create a central "super-market" where AI companies can buy licenses for data easily, rather than trying to ask permission from millions of individual artists.
- Opt-In, Not Opt-Out: Currently, artists have to say "No" to AI. The paper suggests the default should be "No" unless the artist explicitly says "Yes." This puts the burden on the AI company to ask, not on the artist to chase them down.
The Bottom Line
"True compliance is not about what a model says, but about how it came to know."
You cannot clean up a dirty foundation by painting the walls. If the AI was trained on stolen data, the model itself is "tainted." The only way to be truly safe is to ensure you have the right to use the data before you train the model.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.