SR-Prominence: A Crowdsourced Protocol and Dataset Suite for Perceptually-Weighted Super-Resolution Artifact Evaluation
This paper introduces SR-Prominence, a crowdsourced dataset and protocol suite that evaluates super-resolution artifacts based on their perceptual impact (prominence) rather than binary presence, revealing that many existing artifacts are imperceptible to most viewers and demonstrating that classical full-reference metrics often outperform specialized detectors in predicting these perceptual effects.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a photo editor trying to fix a blurry, low-quality picture. You use a fancy AI tool to make it "super-resolution" (sharper and bigger). Sometimes, the AI does a great job. Other times, it gets creative and invents things that aren't there—like turning a smooth wall into a weird, wavy pattern or making a face look like plastic. These mistakes are called artifacts.
For a long time, scientists trying to fix these AI tools have had a blind spot: they treated all mistakes as if they were equally bad. If an AI messed up a tiny patch of grass or a huge, obvious face, the old systems gave them the same "bad score."
This paper, SR-Prominence, argues that this is like grading a student's test by counting every wrong answer, even if one was a tiny typo and the other was a completely wrong essay. The authors say we need to care about how annoying the mistake actually looks to a human eye.
Here is the breakdown of their work using simple analogies:
1. The Core Idea: "Prominence" vs. "Presence"
The authors introduce a new concept called Artifact Prominence.
- Old Way (Presence): "Is there a mistake?" (Yes/No).
- New Way (Prominence): "How many people would notice and be bothered by this mistake?"
The Analogy: Imagine you are painting a room.
- Low Prominence: You accidentally paint a tiny, slightly weird dot on a patch of grass in the background. Most people walking by won't even see it. It's a mistake, but it doesn't ruin the room.
- High Prominence: You accidentally paint a weird, wavy pattern on the front door or a person's face. Everyone walking by stops and says, "Hey, that looks wrong!"
The paper argues that we should focus on fixing the "front door" mistakes, not the "grass dot" mistakes.
2. The Experiment: The "Crowd Test"
To measure this "annoyance level," the researchers didn't just use computers. They used humans.
- They took thousands of images and the "mistake maps" (masks) created by AI tools.
- They showed these images to 30 different people on a crowdsourcing website.
- They highlighted the mistake area and asked: "Does this look weird or distorted to you?"
- The Score: If 90% of people said "Yes, that's weird," the artifact has 90% Prominence. If only 10% said "Yes," it has 10% Prominence.
The Big Surprise: They tested an existing dataset of "mistakes" (called DeSRA) and found that nearly half (48.2%) of the mistakes the experts had marked were actually invisible to the average person. The old "Yes/No" lists were full of false alarms.
3. The Toolkit: SR-Prominence
The authors built a massive new library called SR-Prominence. Think of it as a giant "Hall of Shame" and "Hall of Fame" for AI photo fixers.
- It contains 3,935 specific examples of mistakes.
- Each example has a score telling you exactly how noticeable it is to humans.
- It covers different types of photos: nature, city buildings, and faces.
4. What They Discovered (The "Audit")
They used this new library to test the current tools used to detect AI mistakes and the AI tools themselves.
- The "No-Reference" Tools Failed: Many current tools try to judge quality without seeing the original perfect photo (like a teacher grading a test without the answer key). The authors found these tools are often confused. They flag the "grass dots" as bad and miss the "front door" disasters.
- The "Old School" Tools Won: Surprisingly, some older, simpler math tools (like SSIM and DISTS) that compare the new photo to the original perfect photo actually did a better job at spotting the noticeable mistakes. They act like a strict editor who knows exactly what the original text was supposed to look like.
- AI Training Doesn't Always Help: Some AI models are specifically trained to avoid mistakes. The authors found that even these "careful" models still make the most annoying, high-prominence mistakes. Being "trained to be safe" doesn't guarantee you won't make the worst errors.
5. The Solution: A New Way to Score
The paper provides a new rulebook for the future.
- Don't just count mistakes: Measure how noticeable they are.
- New Scoring System: They created a way for researchers to test new AI tools against their human-voted library without needing to hire 30 people for every single test. They use a "pseudo-teacher" (a lightweight AI) to simulate the original photo so computers can do the math quickly.
Summary
In short, this paper says: "Stop counting every tiny speck of dust as a disaster. Start measuring how much the dust actually bothers the people looking at the picture."
They built a new database where every mistake is rated by how much it annoys a crowd of humans, proving that the old way of judging AI photo tools was missing the most important part: human perception.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.