LocateEdit-Bench: A Benchmark for Instruction-Based Editing Localization
This paper introduces LocateEdit-Bench, a large-scale dataset of 231K images designed to address the critical gap in AI-generated forgery localization by benchmarking methods against emerging instruction-based image editing paradigms.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a magic paintbrush that can change a photo just by listening to your voice. You could say, "Add a tiger to the sofa," and the brush would instantly paint a tiger there, making it look so real that it blends perfectly with the room. This is what modern "instruction-based" image editing does.
However, this magic creates a new problem for digital detectives. For years, experts have been good at spotting "fake" photos because the old magic paintbrushes left obvious clues—like a jagged line where a new object was pasted in. But this new, voice-controlled brush is so smooth that it leaves almost no visible seams. It's like a master forger who doesn't just paste a fake signature; they rewrite the whole document so perfectly that the paper looks untouched.
The Problem: The "Invisible" Edit
The authors of this paper realized that the tools we use to catch these fakes are like security guards trained only to look for torn paper. They are great at finding old-style fakes (where someone cut and pasted an image), but they are completely confused by these new, seamless edits. When you ask an AI to "change the car to red," the AI doesn't just paint over the car; it rebuilds the car from scratch, making it look like it was always there. The old detectors can't find the "cut" because there isn't one.
The Solution: A New Training Ground (LocateEdit-Bench)
To fix this, the researchers built a massive new training ground called LocateEdit-Bench. Think of this as a giant "whack-a-mole" game for AI detectives.
- The Game Board: They created 231,000 fake images.
- The Forgers: They used four of the smartest, most advanced AI editors currently available to create these fakes.
- The Moves: They practiced three main types of tricks:
- Adding: Putting a new object in (like a snowman in a living room).
- Swapping: Replacing one thing with another (like turning a wooden bridge into a stone castle).
- Changing: Altering an object's look (like turning a blue vase green).
- The Answer Key: For every single fake image, they also created a perfect "mask" (a digital outline) showing exactly where the AI changed the picture. This is the "answer key" that the detective AI needs to learn from.
The Experiment: Putting Detectives to the Test
Once they built this massive dataset, they took the best existing "fake-spotting" AI tools and put them through the wringer. They asked: Can these old tools find the new, invisible edits?
The results were a wake-up call:
- The Gap: The old tools struggled mightily. They were like security guards trying to find a ghost; they could see the room, but they couldn't spot the invisible intruder.
- The "Magic" Editor: They found that one specific AI editor (called BAGEL) was the hardest to catch. It was so good at blending its changes that even the smartest detectors got confused.
- The Generalization Problem: When they trained a detector on fakes made by one AI and then tested it on fakes made by a different AI, the detector's performance crashed. It's like teaching a student to solve math problems using only addition, and then giving them a test with multiplication. They didn't learn the underlying logic; they just memorized the specific patterns of the first teacher.
The Takeaway
This paper doesn't claim to have solved the problem of catching all fakes yet. Instead, it built the first giant, realistic "gym" where researchers can train their detectors to spot these new, seamless edits.
They showed us that the old way of looking for "glitches" isn't enough anymore. To catch these new fakes, we need detectives that understand the story of the image, not just the pixels. This new dataset, LocateEdit-Bench, is the first step in teaching those detectives how to spot the invisible changes before they become a problem for everyone.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.