← Latest papers
🤖 AI

Domain-Grounded Candidate Selection for Agentic Image Editing: A Shadow Removal Case

This paper proposes a physics-grounded agentic pipeline that leverages commercial generative models for shadow removal while using domain-specific priors to constrain hallucinations and select optimal edits, achieving state-of-the-art performance on the ShadowRemovalRefine benchmark.

Original authors: Shilin Hu, Jingyi Xu, Dimitris Samaras, Hieu Le

Published 2026-08-07
📖 4 min read☕ Coffee break read

Original authors: Shilin Hu, Jingyi Xu, Dimitris Samaras, Hieu Le

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to fix a broken toy, but instead of using a screwdriver, you ask a super-smart robot artist to redraw the whole toy from scratch. This robot has seen millions of pictures on the internet, so it knows exactly what a toy should look like. It can make the colors pop and the details sharp. But here's the catch: because the robot is so good at guessing what things should look like, it sometimes gets too creative. It might decide the broken part wasn't broken at all, but actually a new, cool feature it invented. Or it might accidentally erase the toy's original shape while trying to fix it. This is the big question scientists are asking right now: Do these super-powerful "AI artists" replace the old, careful rules of physics and math that we used to fix pictures? Or do we still need those old rules to keep the AI from daydreaming?

This paper dives into that question using a specific, tricky problem: removing shadows from photos. Shadows are weird. They aren't physical objects; they are just areas where light is blocked. But to an AI that has only seen pictures, a dark shadow can look like a dark rock, a hole in the ground, or a piece of fabric. If you just ask a powerful AI to "remove the shadow," it might try to "fix" the shadow by inventing a new object or changing the texture of the wall, thinking the darkness was a mistake in the material. The authors found that while these AI artists can make photos look amazing, they often fail at this specific task because they don't understand the physics of how shadows work. They treat shadows like solid things instead of just missing light.

To solve this, the researchers built a smart "agent" system—a team of digital workers that don't just guess once, but think, check, and try again. Here is how their system works:

First, the system asks the AI artist to make a "guided probe," which is basically a first draft of the shadow-free image. Then, a strict "evaluator" (another AI) looks at that draft. If the evaluator sees that the AI accidentally turned the shadow into a fake rock or hallucinated a new object, it rejects the draft and says, "Try again, but remember: shadows are just blocked light, not objects!"

If the first draft passes the check, the system doesn't stop there. It asks the artist to make a few more versions (candidates) of the fix. Then, it filters through all these options, picking the one that removes the shadow best without changing the rest of the scene. Crucially, the researchers gave the AI a special "physics lesson" in its instructions. They told the generator and the evaluator explicitly: "Shadows are caused by light being blocked, not by the material of the object." This simple rule grounded the AI in reality.

The results were impressive. On a standard test called the ShadowRemovalRefine benchmark, their "physics-grounded" system achieved a score (called CDD) of 0.0075. This was at least 47% better than the previous best method. The paper suggests that while these commercial AI models are incredibly powerful, they don't replace the need for classic physics rules. Instead, those rules are now needed to "steer" the AI, keeping it from getting too creative and ensuring the final photo looks real, not just plausible.

The authors also found that their system wasn't perfect. Sometimes, even with the physics rules, the AI would still slightly change the colors or smooth out textures in areas that weren't supposed to be touched. This suggests that while the "agent" approach is a huge step forward, we still need more work to make these edits perfectly safe for every part of a photo. But the main takeaway is clear: we don't need to throw away the old physics textbooks; we just need to use them to teach the new AI artists how to be more careful.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →