MatReplace: A Reference-Free, Conditioning-Aligned Benchmark for Material Replacement in Interior Scenes
This paper introduces MatReplace, a reference-free benchmark designed to evaluate material replacement in interior scenes across four verifiable dimensions and three conditioning tracks, revealing that while top closed-source models excel at rendering named materials, grounding materials from visual references remains a significant challenge.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine walking into a furniture store and asking to see a sofa upholstered in a specific navy blue velvet, or a kitchen island topped with a particular kind of green marble. In the real world, this requires sourcing materials, shipping them, and physically rebuilding the object. In the digital world, designers have long hoped to simply type a request or show a picture of a material and have a computer instantly repaint a room with that new texture, while keeping the shape of the furniture, the lighting, and the rest of the room exactly the same. This is the promise of generative image editing: tools that can alter a photograph as easily as a painter changes a brushstroke. However, while these tools have become common, no one had a reliable way to test if they were actually doing the job correctly. The challenge is that a perfect edit is not just about making the new material look right; it is about ensuring the old parts of the image do not change, the shadows fall naturally on the new surface, and the object does not warp or melt into something else. Without a clear standard, it is impossible to know if a computer is truly understanding the request or just guessing.
A team of researchers from the National University of Singapore, Nanyang Technological University, and VinUniversity has created a new way to measure this ability, called MatReplace. Instead of asking a computer to compare its new image to a single "correct" answer stored in a database, they designed a test that checks four specific things about every edit. First, does the painted area actually look like the requested material? Second, did the rest of the room stay exactly the same? Third, did the shape of the object remain intact, or did the computer accidentally reshape the furniture? And fourth, does the new surface look like it belongs in the room, with the correct lighting and shadows? By checking these four points without needing a pre-existing perfect photo to compare against, the researchers built a benchmark that can fairly judge different computer programs, even if those programs were built by different companies.
The researchers tested this system on over 1,400 different interior design tasks, ranging from changing a headboard to navy suede to swapping a floor for herringbone oak. They organized the tests into three different ways of giving instructions to the computer. In the first setup, the computer only received a text description. In the second, it received the text plus a digital outline showing exactly which part of the image to change. In the third, it received the outline plus a reference photograph of the material to be used. The results revealed a clear split between the most advanced, closed-source systems and the open-source ones. The strongest commercial tools, which are not available to the public, performed exceptionally well when given just a text description. They could change the material perfectly while keeping the rest of the scene untouched, proving that for the best systems, simply naming the material is enough.
However, the study found that adding more information did not always help. When the researchers gave the computer a digital outline of the area to change, it only improved the results for systems that were already struggling to keep the rest of the room stable. For the strongest systems, the outline made no difference, and for some weaker systems, it actually made the results worse. The most surprising finding came from the third setup, where the computer was shown a picture of the material instead of a text description. In this case, every single system performed worse. Instead of using the picture to understand the texture, many of the computers simply copied the picture onto the screen without adjusting the lighting or shape, or they failed to apply the material at all. One system even painted the reference picture itself instead of the room. This suggests that while computers have become very good at understanding words like "navy suede," they still struggle to understand what a picture of a material means when it needs to be applied to a new object.
The researchers also compared their new scoring method against older ways of judging these edits, which often relied on comparing the result to a stored reference image or using text-matching tools. Their new method matched the judgments of human experts much more closely than the old methods did. Human raters agreed with the new system's ranking of the tools, confirming that the four-part check was a reliable way to measure quality. The study concludes that the problem of changing a material based on a text description is largely solved for the most advanced tools, but the problem of using a picture to guide that change remains unsolved. Until computers can better translate a visual reference into a realistic, three-dimensional change, designers will likely get better results by naming the material in words rather than showing a picture. This work provides a clear, fair way to track progress in the field, ensuring that future tools are judged on whether they truly preserve the scene and lighting, rather than just on how closely they mimic a specific style.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.