SAGA: Source Attribution of Generative AI Videos
The paper introduces SAGA, a novel framework that leverages a data-efficient pretrain-and-attribute strategy and Temporal Attention Signatures to achieve state-of-the-art, multi-granular source attribution of generative AI videos, identifying specific models and developers with minimal labeled data while providing interpretable forensic insights.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the internet is a giant, bustling city where everyone is creating videos. For a long time, the only question people asked was, "Is this video real, or is it a fake?" It was like a security guard at a club checking IDs: Real (let them in) or Fake (kick them out).
But now, AI has gotten so good at making videos that they look almost perfect. The problem isn't just knowing if a video is fake anymore; it's knowing who made it. Was it made by a specific AI company? A specific artist? Or a brand new, unknown AI tool?
This is where SAGA comes in. Think of SAGA as a super-detective that doesn't just catch the criminal; it identifies the specific criminal's fingerprint, their gang, their style, and even the specific tool they used to commit the crime.
Here is how SAGA works, broken down into simple concepts:
1. The Problem: Too Many "Fakes" to Count
Imagine you have a pile of 19 different types of forgeries. A simple security guard (old AI detectors) can tell you, "This is a forgery." But they can't tell you if it was forged by "Artist A" using "Pen X" or "Artist B" using "Pen Y."
As AI video generators (like Sora, Runway, Pika) multiply, we need a way to trace a video back to its exact source. This is crucial for:
- Forensics: Figuring out who spread a fake news video.
- Copyright: Knowing who owns the style or the model.
- Safety: Stopping bad actors from using specific dangerous tools.
2. The Solution: SAGA's "Two-Stage" Training
SAGA is smart about how it learns. It doesn't try to learn everything from scratch, which would take forever and require millions of labeled examples. Instead, it uses a Two-Stage Strategy:
- Stage 1: The "Real vs. Fake" Boot Camp.
First, SAGA trains on a massive amount of data just to learn the difference between a real human video and any AI video. It's like teaching a detective to spot a fake ID card in general. It learns to see the subtle "glitches" or "artifacts" that AI leaves behind. - Stage 2: The "Specific Suspect" Specialization.
Once the detective knows what a fake looks like, SAGA teaches it to distinguish which AI made it. Here is the magic trick: It only needs 0.5% of the data to learn this.- Analogy: Imagine you have a library with 100,000 books. Usually, to learn the author of every book, you'd need to read them all. SAGA is like a genius who reads just 500 pages, figures out the pattern, and can then identify the author of any book in the library, even ones it's never seen before.
3. The Secret Weapon: "Temporal Attention Signatures" (T-Sigs)
How does SAGA know the difference between two AI models that look almost identical?
AI videos have a "heartbeat." When an AI generates a video, it doesn't just make one picture; it makes a sequence of pictures that move. Sometimes, the AI gets a little clumsy with how one frame moves to the next. It might flicker, warp, or move in a weird rhythm.
SAGA invented something called T-Sigs (Temporal Attention Signatures).
- Analogy: Think of every AI generator as a musician. Even if they play the same song, they have a unique "style" or "fingerprint" in how they hit the drums or slide the guitar.
- SAGA listens to the "rhythm" of the video frames. It creates a visual map (a signature) of how the AI moves from one frame to the next.
- The Cool Part: Even if SAGA has never seen a specific AI before, it can look at the video's "rhythm," see a pattern it doesn't recognize, and say, "This isn't the AI I know, but it has a unique rhythm. It's a new suspect."
4. The Five Levels of Detective Work
SAGA is a multi-level detective. It can answer questions at different levels of detail:
- Level 1 (BIN-L): Is this real or fake? (Yes/No)
- Level 2 (TASK-L): Did it turn text into video, or an image into video?
- Level 3 (SD-L): Which version of the "engine" was used? (e.g., Stable Diffusion 1.5 vs. 2.1)
- Level 4 (TEAM-L): Which company or research team built it? (e.g., Google vs. OpenAI)
- Level 5 (GEN-L): Which exact model made this? (The specific fingerprint)
5. Why This Matters
Before SAGA, if a new AI tool popped up tomorrow, old detectors would be useless until they were retrained with tons of new data. SAGA is different because:
- It's Data-Efficient: It learns incredibly fast with very little data.
- It's Explainable: It doesn't just give a number; it shows you why it thinks a video is fake (by showing you the "rhythm" signature).
- It's Future-Proof: It can spot new, unseen AI tools because it understands the fundamental "glitches" of how AI thinks, not just memorizes specific examples.
In short: SAGA is the ultimate AI video detective. It doesn't just say "This is a lie." It says, "This lie was told by [Name], using [Tool], with [Specific Style], and here is the proof." This helps us keep the internet honest in an age where seeing is no longer believing.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.