← Latest papers
💬 NLP

QQJ: Quantifying Qualitative Judgment for Scalable and Human-Aligned Evaluation of Generative AI

This paper introduces QQJ, a scalable and human-aligned evaluation framework that bridges the gap between automated assessment and human judgment by anchoring large language model evaluators in expert-designed, multi-dimensional rubrics, thereby achieving superior consistency, interpretability, and diagnostic capability for generative AI systems compared to traditional metrics and unconstrained LLM evaluators.

Original authors: Marjan Veysi, Pirooz Shamsinejadbabaki, Mohammad Zare, Mohammad Sabouri

Published 2026-05-19
📖 4 min read☕ Coffee break read

Original authors: Marjan Veysi, Pirooz Shamsinejadbabaki, Mohammad Zare, Mohammad Sabouri

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are running a massive art gallery filled with paintings and stories created by robots. The robots are getting better every day, but you have a big problem: How do you decide which pieces are actually good?

This paper introduces a new system called QQJ (Quantifying Qualitative Judgment) to solve this problem. Here is how it works, explained through simple analogies:

The Problem: The "Surface-Level" vs. The "Real Deal"

Currently, there are two main ways people try to judge these robot creations, and both have flaws:

  1. The "Math Check" (Traditional Metrics): Imagine a robot judge that only counts how many words or colors match between the robot's art and a perfect example. It's like grading a student's essay only by counting how many times they used the word "the." It's fast and easy, but it doesn't care if the story makes sense or if the painting is actually beautiful. It misses the feeling of quality.
  2. The "Human Crowd" (Human Evaluation): Imagine hiring thousands of art critics to look at every single piece. They are great at spotting deep quality, but it's incredibly expensive, slow, and they might disagree with each other. You can't do this for millions of robot creations.
  3. The "Smart Robot Judge" (LLM Evaluators): Recently, people started using super-smart AI (Large Language Models) to act as judges. They are faster than humans, but they often get confused, biased, or inconsistent. They might give a high score to a pretty picture that is actually nonsense because they aren't following a strict rulebook.

The Solution: The "Master Chef's Recipe Book" (QQJ)

The authors created QQJ to get the best of both worlds: the speed of a robot and the wisdom of a human expert. They do this by separating what we are looking for from how we check it.

Here is the step-by-step process, using a cooking analogy:

Step 1: The Expert Recipe (Rubric Construction)
Instead of letting the robot judge guess what "good" looks like, human experts write a strict recipe book (called a rubric).

  • Analogy: Think of a Michelin-star chef writing a checklist for a food critic. The checklist doesn't just say "tasty." It breaks it down: "Is the salt level right? Is the meat cooked to the exact temperature? Is the plating neat?"
  • In QQJ, humans define these specific dimensions (like "Is the fact true?" or "Did it follow the instructions?") before any judging happens.

Step 2: The Training Camp (Calibration)
The robot judge (the AI) is given a small set of examples that the human experts have already graded using that recipe book.

  • Analogy: The robot judge goes to a training camp where the Master Chef shows it: "Here is a dish that got a 5/5 because it followed the recipe perfectly. Here is a dish that got a 2/5 because it burned the garlic. Now, you try to grade these using the same logic."
  • This teaches the robot to think like the expert, not just guess.

Step 3: The Mass Production Line (Scalable Evaluation)
Once the robot judge has learned the recipe, it can grade thousands of robot creations instantly.

  • Analogy: Now, the robot judge can taste-test 10,000 dishes in an hour. Because it is following the Master Chef's specific checklist, it doesn't get tired, it doesn't get biased, and it gives a detailed report (e.g., "The flavor was good, but the texture was wrong") rather than just a single vague score.

Why is this better?

The paper tested QQJ against the old methods on both text and image generation. Here is what they found:

  • It thinks like a human: QQJ agreed with human experts much more often than the "Math Check" or the untrained "Smart Robot Judge."
  • It's consistent: If you ask the QQJ robot to judge the same picture twice, it gives the same score. The other robot judges often change their minds.
  • It catches the sneaky mistakes: QQJ is really good at spotting "hallucinations" (when the robot makes up facts) or "intent mismatches" (when the robot ignores what you asked it to do). The old math-based methods often missed these completely because the robot's answer looked similar to a real answer, even if it was wrong.

The Bottom Line

QQJ is like giving a robot a strict, human-written rulebook and a short training session. This allows us to evaluate millions of AI creations quickly and cheaply, while still keeping the deep, human understanding of what actually makes something "good." It turns the messy, subjective art of judging quality into a clear, repeatable science.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →