Visual Distortion Detection in UGC Images Using Large Multimodal Models
The paper proposes VIGIL, a novel framework leveraging Large Multimodal Models with multi-level feature detectors and a rigorously curated 140K-sample dataset to address the limitations of existing text-driven approaches and bridge the synthetic-to-authentic generalization gap in visual distortion detection.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are looking at a photo on your phone. Maybe it's a selfie, a picture of your dog, or a sunset you snapped while hiking. Sometimes, the whole picture looks great. Other times, the sky is perfect, but your friend's face is blurry, or there's a weird smudge of color on the grass. For a long time, computers that analyze photos were like a strict teacher who only gave a single grade for the whole test. They could tell you if a picture was "good" or "bad" overall, but they couldn't point to the specific spot where the mistake happened. This is a big problem because in the real world, we often want to fix just the blurry part, not the whole photo.
The field trying to solve this is called Image Quality Assessment. Think of it as teaching computers to be art critics. Recently, scientists have started using "Large Multimodal Models" (LMMs). You can think of these as super-smart robots that can see pictures and read text at the same time. They are like a genius art student who has read every book on photography. However, just because a robot is smart doesn't mean it's good at finding tiny, specific errors. Most of these robots were trained by being shown pictures with text descriptions like "the sky is blurry," but they struggle to actually draw a box around the blurry sky. They also have a hard time when moving from "fake" practice pictures (where scientists artificially add blur) to real, messy photos taken by people in the wild.
This paper introduces a new robot named VIGIL (Visual Intelligence for Generalized Image Localization). The researchers wanted to build a system that doesn't just say "this photo is bad," but actually points a finger and says, "Hey, the blur is right here, and it's a compression error." They found that the old way of teaching these robots—just having them write text descriptions—wasn't precise enough. So, they tried a different approach: they taught the robot to act like a detective looking for clues, rather than a writer describing a scene.
The team started by gathering a massive library of over 1 million high-quality images from the internet. They carefully filtered these to make sure they were clear and bright, then they played a game of "spot the difference" by artificially adding 8 different types of distortions (like blur, noise, or weird colors) to random parts of the images. They created a training set called VIGIL-140K, which contains over 140,000 distorted images with more than 200,000 labeled "bad spots."
Instead of asking the robot to write sentences, the researchers turned the robot's brain into a team of detectives. They used different layers of the robot's internal thinking process as separate "detectors." Imagine a classroom where instead of one student raising their hand to answer, five different students look at the same picture at the same time, each using a slightly different pair of glasses. One might be good at seeing fuzzy edges, while another is great at spotting color shifts. They all shout out their guesses simultaneously. The robot then combines all these guesses to make a final, very accurate decision.
The researchers also solved a tricky problem called the "Synthetic-to-Authentic" gap. This is like practicing archery with paper targets that have perfect, round bullseyes, but then trying to hit a wobbly, irregular target in a real forest. The paper targets (synthetic data) look too neat, so the robot gets confused when it sees real, messy distortions. To fix this, the team taught the robot to pay attention even when it thinks a spot is "clean." If the robot is almost sure a spot is clean but has a tiny doubt, they let that doubt count as a possible clue. This helps the robot stay alert when the real world gets messy.
When they tested VIGIL, it was a huge success. On practice tests with the fake distortions, it found the errors much better than other top models. But the real magic happened when they tested it on real-world photos it had never seen before. While other models got confused and missed the errors, VIGIL kept its cool, accurately pointing out the blurry or noisy spots in real user photos. The paper shows that by using a team of internal detectors and keeping an open mind about "clean" areas, we can teach computers to be much better at spotting exactly where a photo has gone wrong, paving the way for smarter tools that can fix our pictures automatically.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.