BeyondMasks: Evaluating Causal and Physical Consistency in Video Object Removal
The paper introduces BeyondMasks, a paired benchmark and the CORE evaluation protocol designed to assess video object removal by measuring causal and physical consistency—such as shadows and reflections—rather than just local masked region fidelity, revealing that current state-of-the-art methods fail to remove secondary physical effects despite high visual plausibility.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the world of digital video, there is a growing ability to edit reality with a simple command. Imagine watching a recording of a busy street and asking a computer to erase a specific car, leaving the rest of the scene perfectly intact. For years, the technology behind this has focused on a narrow task: filling in the empty space where an object used to be. It is like a painter carefully repainting a hole in a wall, trying to match the texture and color of the surrounding bricks. This approach works well when the object is just sitting there, but it often fails when the object has been interacting with its world. In the real physical world, things do not just occupy space; they cast shadows, reflect light, change the way a room is lit, and leave behind traces of their movement, like ripples in water or dust kicked up by a footstep. If you remove the object but leave these physical consequences behind, the scene looks wrong, as if a ghost is haunting the frame.
A team of researchers has realized that true video editing requires more than just filling a hole; it requires understanding the cause-and-effect relationships that bind an object to its environment. They have introduced a new way of testing video editing tools, called BeyondMasks, which challenges computers to remove not just the object, but also the invisible physical effects it created. To do this, they built a collection of 180 video clips, half created on computers and half filmed in the real world. In every clip, they have a "before" version with an object present and an "after" version where that object was never there, serving as the perfect answer key. This allows them to see exactly what the computer gets wrong. They found that even the most advanced video tools, which can make the empty space look very realistic, consistently fail to remove the secondary effects. A shadow might linger on the ground where a fox once stood, or a reflection might remain in a window after a lamp has been deleted. The researchers discovered that current technology treats the object and its physical footprint as separate problems, when in reality, they are one and the same.
To measure these failures, the team developed a new scoring system that acts like a critical eye. Instead of just counting how many pixels match the perfect answer, they use a sophisticated artificial intelligence to watch the videos and judge whether the scene looks physically consistent. This system checks two things: did the object disappear completely, and did the environment return to a natural state as if the object had never been there? When they tested nine of the best video editing models available today, the results were revealing. While some models scored very high on traditional measures of image quality, they scored poorly on this new test of physical logic. For instance, a tool might successfully erase a duck from a pond but leave behind the ripples it made, or remove a teapot but leave the shadow it cast on the table. The study shows that these tools are still struggling to understand the physics of light and matter. They can hallucinate a background that looks plausible, but they cannot reason about the causal chain of events that an object triggers.
The researchers also looked at why these mistakes happen. They found that many editing tools work by simply blocking out the area where the object is and asking the computer to guess what should be there. This approach throws away valuable clues. If a glass cup is sitting on a table, the computer can see the table through the glass. If the tool blocks out the cup entirely, it loses that view of the table and has to guess the pattern, often getting it wrong. The study suggests that future removal systems should incorporate context-aware reasoning rather than relying solely on masked inpainting to better recover missing information. The study highlights a significant gap between what looks visually smooth and what is actually physically correct. It suggests that for video editing to truly feel real, the technology must move beyond simple picture repair and learn to understand the invisible rules of the physical world, ensuring that when an object is removed, the world it lived in returns to a state of perfect consistency.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.