← Latest papers
🤖 AI

APEX: Assumption-free Projection-based Embedding eXamination Metric for Image Quality Assessment

APEX is a novel, assumption-free image quality assessment framework that leverages the Sliced Wasserstein Distance with open-vocabulary foundation models (CLIP and DINOv2) to overcome the limitations of traditional feature-distribution metrics, offering superior robustness and stability across diverse datasets.

Original authors: Caterina Gallegati, Monica Bianchini, Franco Scarselli, Vittorio Murino, Barbara Toniella Corradini

Published 2026-05-11
📖 4 min read☕ Coffee break read

Original authors: Caterina Gallegati, Monica Bianchini, Franco Scarselli, Vittorio Murino, Barbara Toniella Corradini

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a judge in a cooking competition. The chefs (generative AI models) are creating incredible new dishes (images) that look almost indistinguishable from real food. Your job is to taste them and decide: "Is this dish good? Is it better than the last one?"

For a long time, the judges used a very old, rigid checklist to grade these dishes. They would look at the ingredients (pixels) and check if they matched a specific recipe from a 1990s cookbook (Inception-v3 trained on ImageNet).

The Problem:
The paper argues that this old checklist has two big flaws:

  1. The "Closed Menu" Problem: The old checklist only knows about ingredients found in its specific 1990s cookbook. If a chef makes a futuristic dish with ingredients the old book never saw, the checklist gets confused or gives a bad score, even if the dish is delicious.
  2. The "Rigid Math" Problem: The old checklist assumes all dishes follow a perfect, predictable bell-curve shape. But real cooking (and real AI images) is messy and complex. Forcing a messy reality into a rigid shape leads to wrong scores.

Recently, some judges tried using a newer, fancier cookbook (Foundation Models like CLIP and DINOv2) to understand modern ingredients. But they kept using the same rigid math to grade them. It's like having a modern menu but still using a 1990s ruler that bends the wrong way.

The Solution: APEX

The authors introduce APEX (Assumption-free Projection-based Embedding eXamination). Think of APEX as a new, super-smart judging system with two main superpowers:

1. The "Shadow Puppet" Trick (Projection-based)
Instead of trying to measure the whole 3D shape of a dish at once (which is hard and requires rigid assumptions), APEX uses a clever trick. Imagine shining a light on a complex sculpture to cast its shadow on a wall.

  • APEX shines light from thousands of different angles (projections).
  • It measures the shadow (the 1D slice) each time.
  • It averages all those shadows to get a true sense of the shape.
  • Why this matters: This method doesn't assume the sculpture is a perfect sphere or cube. It just measures the shape as it actually is, no matter how weird or complex. It's "assumption-free."

2. The "Universal Translator" (Foundation Models)
Instead of using the old 1990s cookbook, APEX uses two modern, super-intelligent translators:

  • CLIP: A translator that understands how images relate to language (great for "Does this look like a cat?").
  • DINOv2: A translator that understands textures, shapes, and fine details (great for "Does this fur look realistic?").
    APEX uses these translators to understand the "flavor" of the image, then uses the "Shadow Puppet" trick to compare the generated image against real ones.

What the Paper Found (The Taste Test)

The authors put APEX to the test against the old judges (FID, KID) and the newer-but-still-rigid judges (CMMD). They tested them on:

  • Natural photos (like a park).
  • Faces (like a celebrity).
  • Medical scans (like an X-ray).
  • Satellite photos (like a view of a city from space).

The Results:

  • Stability: The old judges were like a shaky hand; if you showed them the same picture twice with a tiny change, they might give wildly different scores. APEX was steady and consistent, like a rock.
  • Fairness Across Domains: The old judges were great at grading faces but terrible at grading X-rays or satellite photos. APEX was equally good at judging all of them. It didn't get confused by the change in subject matter.
  • Sensitivity: When the AI chefs slowly improved their dishes (adding more detail step-by-step), APEX was the only one that noticed the tiny, subtle improvements at the very end. The old judges often got stuck or even thought the dish got worse when it actually got better.
  • Human Agreement: When real humans rated the images, APEX's scores matched human opinions almost perfectly, often better than the established "gold standard" (FID).

The Bottom Line

The paper claims that APEX is a better way to grade AI-generated images because it stops forcing images into old, rigid boxes. Instead, it uses modern, flexible tools to look at images from every possible angle, ensuring that the score reflects what the image actually looks like, whether it's a face, a tumor, or a satellite view. It's a more honest, stable, and human-aligned judge for the future of AI art.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →