← Latest papers
🤖 AI

Relational Visual Similarity

This paper addresses the critical gap in existing visual similarity metrics by introducing a new framework and an anonymized 114k image-caption dataset to train a Vision-Language model capable of measuring relational similarity—matching images based on their internal structural logic rather than just their surface-level visual attributes.

Original authors: Thao Nguyen, Sicheng Mo, Krishna Kumar Singh, Yilin Wang, Jing Shi, Nicholas Kolkin, Eli Shechtman, Yong Jae Lee, Yuheng Li

Published 2026-04-13
📖 4 min read☕ Coffee break read

Original authors: Thao Nguyen, Sicheng Mo, Krishna Kumar Singh, Yilin Wang, Jing Shi, Nicholas Kolkin, Eli Shechtman, Yong Jae Lee, Yuheng Li

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are looking at a photo of a banana that is slowly turning brown and rotting. Now, imagine a photo of a matchstick that is slowly burning down to ash.

If you ask a standard computer program (like the ones used by Google Images or social media algorithms) if these two pictures are similar, it will say "No."

Why? Because the computer is looking at the surface details:

  • One is yellow and curved; the other is thin and wooden.
  • One is food; the other is a tool.
  • They don't share the same color, shape, or "category."

But if you ask a human, we might say, "Yes, they are very similar!"
Why? Because we see a hidden story connecting them. Both are undergoing a transformation over time. Both are going from "whole" to "gone." We aren't just comparing their looks; we are comparing their logic.

This paper introduces a new way for computers to see the world, not just by what things look like, but by how they work.

The Core Idea: "Relational Visual Similarity"

The authors call this "Relational Visual Similarity."

Think of it like this:

  • Attribute Similarity (The Old Way): "These two things are alike because they are both red and round." (Like comparing an apple to a peach).
  • Relational Similarity (The New Way): "These two things are alike because they both tell the same story." (Like comparing the Earth's layers to a peach's layers: skin/crust, flesh/mantle, pit/core).

The paper argues that current AI is great at the first type but terrible at the second. It misses the "aha!" moments of human creativity and reasoning.

How Did They Teach the Computer?

To teach the computer this new way of thinking, the researchers had to build a special "school" for it. Here is how they did it, step-by-step:

1. The "Interesting" Filter
They started with a massive library of millions of random internet photos (LAION-2B). Most of these photos are boring (e.g., just a picture of a chair). The researchers trained a smart AI to act as a librarian, scanning the library and picking out only the "interesting" photos—those that show patterns, transformations, or clever connections (like the banana and the match).

2. The "Anonymous" Translator
This is the cleverest part. Usually, when we describe a photo, we say, "A banana turning brown."
But the researchers wanted the AI to ignore the banana and focus on the story. So, they trained an AI to write "Anonymous Captions."

Instead of saying "A banana turning brown," the AI learned to say:

"A {Subject} undergoing a {Transformation} over time."

By replacing specific objects with placeholders (like {Subject}), the AI learned to strip away the "costume" and look at the "plot" of the image.

3. The Matchmaker
Once they had thousands of photos paired with these "plot summaries," they taught a new model (called RelSim) to match images based on their plots, not their costumes.

  • If Image A has the plot "Transformation over time," and Image B has the plot "Transformation over time," the model says, "These are a match!" even if one is a banana and the other is a burning match.

Why Does This Matter?

The paper shows that this new model can do things old models can't:

  • Better Search: Imagine you are an artist looking for inspiration. You want to find an image that feels like "a bird flying into a cage." Old search engines would only show you pictures of birds and cages. The new model could find a picture of a "fish swimming into a net" or a "person walking into a trap," because they share the same relational logic.
  • Creative Generation: If you ask an AI to "make a new image with the same idea as this one," the old AI might just change the colors. The new AI understands the concept. If the input is a visual pun (like the word "Ice-cream" made of melting ice), the new AI can generate a different visual pun (like the word "Fire" made of burning wood) because it understands the underlying joke, not just the pixels.

The Bottom Line

Current AI is like a child who only recognizes things by their color and shape.
This new research teaches AI to be a philosopher. It learns to look past the surface and understand the invisible connections, the stories, and the patterns that make two completely different things feel like they belong together.

It's a giant leap from asking "What is this?" to asking "What is happening here?"

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →