How Do Data Owners Say No? A Case Study of Data Consent Mechanisms in Web-Scraped Vision-Language AI Training Datasets
This paper investigates DataComp, a massive vision-language dataset, to reveal that a significant portion of its web-scraped content violates data owners' consent through copyright notices, restrictive Terms of Service, and watermarks, thereby exposing critical gaps in current AI data collection practices and the urgent need for a unified consent framework.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the internet as a massive, global library where anyone can walk in and take a book off the shelf. In recent years, Artificial Intelligence (AI) companies have been trying to build "super-brains" by reading billions of these books (images and text) all at once. They call this process "scraping."
This paper is like a detective report investigating whether the people who wrote those books (the data owners) actually said, "You can read my book to build your robot."
Here is the story of what the researchers found, broken down into simple concepts:
1. The Big Library vs. The AI's Recipe Book
The researchers looked at a specific, huge collection of data called CommonPool. Think of this not as a box of photos, but as a giant recipe card that just lists the address (URL) of where a photo is stored on the internet, along with a short description.
- The Problem: The AI builders download this recipe card. They don't have the actual photos yet. To get the photos, they have to go back to the internet addresses listed on the card and download them themselves.
- The Analogy: Imagine a librarian hands you a list of 12 billion addresses to houses. The librarian says, "Here is a list of houses; go knock on the doors and take a picture of whatever is inside." The librarian didn't ask the homeowners if it was okay; they just gave you the list.
2. The "Do Not Disturb" Signs (Data Consent)
The researchers asked: "How do homeowners tell the AI to stop?" They looked for three types of signs:
- The Copyright Notice (The "©" Sign): Just like a book has a copyright page, many images have a little "©" symbol or a watermark saying "Property of Jane Doe."
- The Finding: They found that about 122 million images in this giant library have these signs. The AI builders often ignore them.
- The Watermark (The Invisible Ink): Many artists put invisible or visible watermarks on their art to prove they own it.
- The Finding: The AI's current tools are terrible at spotting these. It's like trying to find a needle in a haystack while wearing blindfolded glasses. The AI thinks it's safe to use the image, but it's actually stolen.
- The "No Trespassing" Fence (Robots.txt & Terms of Service): Every website has a digital fence. Some say, "You can look, but don't take anything." Others say, "You can take photos, but only if you are a student or a researcher, not a robot."
- The Finding: They checked the top 50 websites that provided the most images. 60% of them had fences saying, "No scraping allowed!" Yet, the AI builders went right over the fence anyway.
3. The "Blind" Delivery Truck
Here is the tricky part that makes this a legal and ethical nightmare.
The AI builders say, "We didn't steal the photos; we just downloaded a list of addresses from a public archive (CommonCrawl)."
- The Analogy: Imagine a delivery truck driver drops off a list of addresses to a construction crew. The crew goes to those houses and starts taking bricks.
- The Reality: The delivery truck driver (the dataset curator) claims, "I didn't take the bricks; I just gave you the list." But the construction crew (the AI user) is the one actually breaking into the houses.
- The Glitch: Because the list only gives the address of the photo (which might be on a generic server like Amazon's cloud), the AI builders often can't see the real house owner's "No Trespassing" sign. They are looking at the wrong door.
4. The "Fair Use" Loophole
Some AI companies argue, "We are just learning from the books; that's fair use."
- The Counter-Argument: The paper points out that even if "learning" is okay, how you get the books matters. If you steal the books from a library that says "No photocopying," your "learning" is built on a lie.
- The Court Cases: The paper mentions that courts are currently deciding if this is okay. Some judges are saying, "Maybe it's fair use to read, but maybe it's not fair use to steal the books from a library that explicitly said 'No'."
5. The Solution: A Unified "Consent Language"
Right now, the internet is a mess of different languages for saying "No."
- Some people write "No" on the image (Watermark).
- Some put a sign on the website (Terms of Service).
- Some put a sign on the gate (Robots.txt).
- Some put a sign in the metadata (EXIF data).
The AI builders are only listening to one language and ignoring the rest.
The Paper's Recommendation:
We need a universal "Do Not Enter" sign that everyone agrees on.
- For the Creators: A simple, clear way to say, "Do not use my art for AI training."
- For the AI Builders: A rule that says, "If you see this sign, you must stop. No excuses."
- The "Opt-In" Idea: Instead of assuming everyone wants their data used unless they say "No" (Opt-Out), the paper suggests we should assume people don't want their data used unless they explicitly say "Yes" (Opt-In).
The Bottom Line
The internet has become the main ingredient for building modern AI. But the researchers found that a huge chunk of those ingredients were taken without asking the chefs (the owners) for permission.
The current system is like a chef who walks into a neighborhood, grabs ingredients from people's fridges because they didn't explicitly lock their doors, and then claims, "I didn't steal; I just took what was available." This paper argues that we need to stop and ask for permission before we start cooking.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.