← Latest papers
💻 computer science

PROVE: A Perceptual RemOVal cohErence Benchmark for Visual Media

This paper introduces the PROVE framework, comprising the perception-aligned RC-S and RC-T metrics and the PROVE-Bench dataset, to address the limitations of existing evaluation methods by providing a more accurate assessment of object removal coherence in visual media that better aligns with human judgment.

Original authors: Fuhao Li, Shaofeng You, Jiagao Hu, Yu Liu, Yuxuan Chen, Zepeng Wang, Fei Wang, Daiguo Zhou, Jian Luan

Published 2026-05-15
📖 5 min read🧠 Deep dive

Original authors: Fuhao Li, Shaofeng You, Jiagao Hu, Yu Liu, Yuxuan Chen, Zepeng Wang, Fei Wang, Daiguo Zhou, Jian Luan

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a photo editor hired to remove a photobomber from a group picture. You do a great job: the person is gone, and the background looks natural. But how do you know if your edit is actually good?

This is the problem the paper PROVE tackles. The authors argue that the current tools we use to grade these edits are like using a ruler to measure the taste of soup—they are looking at the wrong things.

Here is a breakdown of their solution using simple analogies.

1. The Problem: The "Copy-Paste" and "Blurry" Traps

The paper explains that existing grading tools (metrics) make two major mistakes:

  • The "Copy-Paste" Trap (Full-Reference Metrics): Imagine a teacher grading a student's essay by comparing it word-for-word to a sample answer. If the student copies the sample perfectly, they get an A, even if the essay is boring.
    • In image editing, old tools reward models that just "copy-paste" the background from the original photo rather than actually "erasing" the object and inventing a new, realistic background. They prefer safe, boring answers over creative, realistic ones.
  • The "Blurry is Clean" Trap (No-Reference Metrics): Imagine a judge who thinks a blurry photo is better than a sharp one because the blur hides mistakes.
    • Current tools often give high scores to blurry results because blur makes things look "smooth" and uniform. They fail to notice that the image is actually low-quality or that the object removal created weird artifacts.

The Video Problem: When editing a video, existing tools look at the whole screen. If the background is stable but the spot where the object was removed is flickering wildly, the tool might still give a high score because the rest of the video is fine. It's like a teacher grading a 10-page essay but only reading the first page and ignoring the rest.

2. The Solution: The "Spotlight" Approach (RC Metrics)

The authors propose a new way to grade edits called RC (Removal Coherence). Instead of looking at the whole picture or comparing it to a perfect reference, they use a "sliding spotlight" to zoom in on the specific area where the object was removed.

They have two main tools:

  • RC-S (Spatial Coherence): Think of this as a texture detective. It zooms in on the hole where the object was and asks: "Does the new background here look like the surrounding neighborhood?"
    • It uses a sliding window (like a magnifying glass) to check if the textures, lighting, and patterns match the area around the edit. If the new background looks like a different texture or is too blurry, the score drops.
  • RC-T (Temporal Coherence): This is the video stability inspector. It looks at the same spot across two consecutive frames.
    • It asks: "Does this spot jump around or flicker as the video plays?" If the background in the removed area is shaking or changing weirdly from frame to frame, this metric catches it, even if the rest of the video is smooth.

The Secret Sauce: They use a special "brain" (a feature extractor called DINOv2) that is very sensitive to tiny details and frequency changes, much like a human eye is sensitive to subtle visual glitches.

3. The New Playground: PROVE-Bench

To prove their new grading system works, they built a new testing ground called PROVE-Bench. They realized that old test sets were either:

  1. Too fake: Computer-generated videos that don't look like real life.
  2. Too hard to grade: Real videos where we don't know what the "perfect" background should look like.

Their new benchmark has two parts:

  • PROVE-M (The Controlled Lab): They filmed real scenes, removed an object, and filmed the scene again without the object. This gives them a "perfect" ground truth to compare against, but they added simulated camera shakes to make it feel like a real handheld video.
  • PROVE-H (The Stress Test): A collection of 100 difficult, real-world videos (crowds, fast motion, reflections) where they don't have a perfect reference. This tests if the models can handle chaos.

4. The Results

When they tested their new "Spotlight" grading system against the old tools:

  • Old tools often gave high scores to blurry, copy-paste results or missed flickering in videos.
  • PROVE (RC-S and RC-T) aligned much better with what actual humans thought looked good. If a human said, "That looks fake," the new metric agreed. If a human said, "That looks great," the new metric agreed.

Summary

The paper argues that to judge object removal, we need to stop looking at the whole picture or comparing it to a perfect reference. Instead, we need to zoom in on the edited area to check if it blends in with its neighbors (Spatial) and stays steady over time (Temporal). They built a new set of tools and a new testing ground to prove that this "zoom-in" approach is the only way to get a grade that matches human perception.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →