← Latest papers
💻 computer science

PhysEditBench: A Protocol-Conditioned Benchmark for Dense Physical-Map Prediction with Image Editors

This paper introduces PhysEditBench, a novel protocol-conditioned benchmark that standardizes the evaluation of general-purpose image editors against specialized models for predicting five dense physical maps (depth, normal, albedo, roughness, and metallic) from single RGB images, revealing that while specialized models generally outperform editors, editors show promise on certain scalar metrics despite suffering from structural and lighting-related errors.

Original authors: Jiaxin Yang, Yu Hou, Muxin Liu, Weixuan Liu, Ze Yuan, Zeming Chen, Zhongrui Wang, Xiaojuan Qi

Published 2026-05-14
📖 5 min read🧠 Deep dive

Original authors: Jiaxin Yang, Yu Hou, Muxin Liu, Weixuan Liu, Ze Yuan, Zeming Chen, Zhongrui Wang, Xiaojuan Qi

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a very talented, general-purpose artist who is famous for taking a photo and painting over it to change the mood, add a hat to a person, or turn a sunny day into a rainy one. This artist is great at following instructions like "make the sky purple" or "add a cat."

Now, imagine you ask this artist a different kind of question: "Can you look at this photo of a living room and tell me exactly how far away every object is? Can you tell me which surfaces are rough like sandpaper and which are smooth like glass? Can you tell me which parts are made of metal?"

This is the core question the paper PhysEditBench asks. The authors wanted to see if these "general-purpose image editors" (the talented artists) could act like specialized scientists who map out the physical properties of a room just by looking at a single picture.

The Problem: The "Generalist" vs. The "Specialist"

Think of a Specialist as a master carpenter who has spent 20 years learning exactly how to measure wood, find the grain, and calculate angles. They have a specific tool for every job.
Think of a Generalist (the image editor) as a versatile handyman who can fix a leak, paint a wall, or build a shelf, but they haven't spent years studying the physics of light and material.

The paper asks: If you give the handyman a photo and a simple instruction, can they figure out the physics of the room as well as the master carpenter?

The Solution: A Strict "Test Protocol"

The authors realized that you can't just ask the handyman, "Show me the depth map," because the handyman might not know what a "depth map" is. They might just draw a picture that looks deep but isn't accurate.

So, the authors created PhysEditBench, which is like a standardized testing center.

  • The Rules: They gave every model (both the handyman and the carpenter) the exact same photo and the exact same set of instructions.
  • The Targets: They tested five specific "physical maps":
    1. Depth: How far away things are.
    2. Normal: Which way a surface is facing (like the direction a wall is leaning).
    3. Albedo: The true color of an object, ignoring shadows and lighting.
    4. Roughness: How smooth or bumpy a surface is.
    5. Metallic: Where the metal is.

The Results: Who Won?

The paper ran the tests and found some interesting things:

  1. The Specialists Still Rule the Classroom: For figuring out Depth, Surface Direction (Normal), and True Color (Albedo), the specialized "carpenters" (models trained specifically for these tasks) were much better. They were like students who had studied the textbook; they got the right answers. The general "handymen" (image editors) often produced results that looked cool but were physically wrong.

  2. The Handymen Are Getting Smarter: The general image editors are getting better. For Roughness and Metallic (the harder, more abstract concepts), the best image editors actually did as well as, or sometimes even better than, the specialists on simple number-crunching tests. They could guess "this looks shiny" or "this looks rough" quite well.

  3. The "Look" vs. The "Truth": The paper found that image editors are great at making things look like a physical map (the visual style is right), but they often fail to get the physics right. It's like an artist painting a perfect-looking 3D cube on a flat piece of paper; it looks 3D, but if you tried to measure it, the math wouldn't add up.

The Catch: Lighting and Confusion

The paper also tested these models in "stress tests"—like trying to read a map in the dark or in blinding sunlight.

  • The specialists stayed strong even in bad lighting.
  • The general image editors got confused. They struggled to tell the difference between a shadow and a dark object, or a reflection and a shiny metal surface.

The Big Takeaway

The paper concludes that while general-purpose image editors are becoming very impressive and can produce outputs that look like scientific maps, they are not yet ready to replace the specialized tools for accurate physical measurement.

Think of it this way: The image editor is a fantastic improvisational actor who can pretend to be a scientist and give a convincing performance. But if you need the actual, precise data to build a real bridge or a robot, you still need the specialized engineer who has done the actual calculations.

The paper doesn't claim these tools will soon be used to build self-driving cars or diagnose medical conditions. It simply says: "Here is a test to see how good these new AI artists are at understanding the physical world, and here is exactly where they succeed and where they still need to learn."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →