PreResQ-R1: Response-Preference Disentangled Ranking-and-Scoring Reinforcement Optimization for Robust Visual Quality Assessment
PreResQ-R1 is a novel reinforcement learning framework that disentangles response and preference objectives to unify absolute score regression and relative ranking, achieving state-of-the-art performance in both image and video quality assessment with strong generalization and human-aligned reasoning capabilities.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
When we look at a photograph or a video, our brains instantly judge its quality. We notice if the colors are dull, if the edges are blurry, or if the image looks grainy. This instinctive judgment is what researchers call visual quality assessment. For decades, computers have struggled to replicate this human ability. Early attempts relied on rigid mathematical formulas that measured pixel errors, but these often failed to match how people actually perceive beauty and clarity. More recently, powerful artificial intelligence models known as large multimodal models have been trained to look at images and describe them. However, when these models are asked to assign a specific quality score, they often become inconsistent. If you ask the same model to rate the same picture twice, it might give two very different numbers, or its reasoning might contradict its final score. This instability makes it difficult to trust these systems for real-world applications, from improving smartphone cameras to ensuring the quality of streaming videos.
A team of researchers at Shanghai Jiao Tong University has developed a new approach to fix this problem, creating a system they call PreResQ-R1. Instead of forcing the artificial intelligence to simply guess a number or blindly follow a ranking, they taught the model to separate two distinct tasks: stabilizing its own reasoning and aligning its preferences with human judgment. The researchers found that the instability in previous systems came from the model's tendency to generate wildly different explanations for the same image. To solve this, they designed a training process that rewards the model for being consistent in its internal logic while also ensuring its final scores match what a human would likely say. The system breaks down the assessment of an image into five specific parts: how vivid the colors are, how much noise or grain is visible, how sharp the details are, how clear the main subject is, and how clear the background is. By evaluating these five aspects individually before combining them, the model learns to build a more reliable and detailed picture of quality.
The researchers tested this new method on a wide variety of images and videos, including those with natural distortions like blur or low light, as well as those created by other artificial intelligence tools. They found that their system significantly outperformed existing methods. On tests involving still images, the new model improved accuracy by 5.60 percent compared to the best previous approaches. For video quality, where the challenge is even greater because the system must judge movement and time, it improved accuracy by 2.53 percent. Crucially, the model did not just produce better numbers; it produced better explanations. When shown a blurry photo of a landscape, the model correctly identified that the leaves were out of focus and the background lacked detail, rather than just giving a vague comment about the image being "bad." This ability to provide a structured, evidence-based reason for its score makes the system far more trustworthy.
The success of this work relies on a two-stage training strategy that mimics a learning process. First, the model is encouraged to explore many different ways of looking at an image, generating a wide range of possible scores and reasons. During this phase, the system penalizes answers that are too scattered or inconsistent, gently guiding the model toward a stable way of thinking. Once the model has found a reliable internal rhythm, the second stage focuses on fine-tuning its preferences. Here, the system compares its judgments against human ratings, learning to adjust its scores so they align with human perception. This separation of "thinking clearly" from "scoring correctly" allows the model to handle difficult cases where the image quality is subjective or the distortions are complex. The researchers also extended this method to video by analyzing both the flow of movement over time and the details within individual frames, without needing to reconstruct every single motion in the video.
The implications of this research extend beyond just getting a better score. By creating a system that can explain its reasoning in a way that matches human perception, the technology opens the door for more robust applications in digital media. It suggests that the key to making artificial intelligence understand visual quality is not just to feed it more data, but to teach it to be consistent in its own reasoning before asking it to make a final judgment. The study demonstrates that when a model is trained to stabilize its internal logic and then align that logic with human preference, it can achieve a level of reliability that was previously out of reach. This approach offers a promising path forward for developing tools that can automatically enhance images, curate video content, and ensure that what we see on our screens truly reflects the quality we expect.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.