← Latest papers
💬 NLP

Would you still call this Dax? Novel Visual References in VLMs and Humans

This paper introduces the Novel Visual References Dataset (NVRD) to evaluate how vision-language models and humans map novel, pre-training-contradictory visual concepts to language, revealing that models struggle with in-context learning of such concepts and significantly overgeneralize compared to human judgments.

Original authors: Ada Defne Tür, Gaurav Kamath, Joyce Chai, Siva Reddy, Benno Krojer

Published 2026-06-05✓ Author reviewed
📖 5 min read🧠 Deep dive

Original authors: Ada Defne Tür, Gaurav Kamath, Joyce Chai, Siva Reddy, Benno Krojer

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a child (or a robot) a new word for a toy. You show them a red, blocky robot and say, "This is a Dax."

Now, imagine you show them a slightly different toy: maybe it’s blue instead of red, or it has an extra antenna, or it’s made of wood instead of plastic. The big question is: Do they still call it a "Dax"?

This paper investigates exactly that question, but instead of children, it tests advanced AI vision models (like the ones inside chatbots that can see images). The researchers created a special dataset called NVRD (Novel Visual References Dataset) to see how well AI learns new visual concepts compared to humans.

Here is the breakdown of what they found, explained simply:

1. The Setup: The "Dax" Test

The researchers didn’t use normal objects like chairs or dogs, because the AI already knows what those are. Instead, they created 90 completely made-up objects (some were hybrids, like a "boar-toaster," and some were totally alien). They gave each object a nonsense name (like "glemture" or "floundge").

Then, they took those original images and created 20 slightly different versions of each one. These changes ranged from tiny tweaks (like adding a little digital noise) to major changes (like removing an arm or changing the shape entirely).

They asked both humans and AI models to look at the original "Dax" and the changed version, and decide: "Is this changed thing still a Dax?"

2. Finding #1: AI is Stubborn About Old Knowledge

The AI struggled most when the new object looked like something it already knew.

  • The Analogy: Imagine showing the AI a picture of a standard chair and telling it, "This is a Blomwich." Then you show it a slightly different chair. The AI will likely ignore your new word and just call it a "chair."
  • Why? The AI has strong "prior knowledge." It’s like a student who already knows the answer to a math problem; it’s hard to convince them to use a different, made-up method. The AI prefers its old labels over the new ones you’re trying to teach it.

3. Finding #2: Shape Matters More Than Color

Both humans and AI care about the shape of the object more than its color or texture.

  • The Analogy: If you paint a red car blue, you still call it a car. But if you remove the wheels and turn it into a boat, you probably won’t call it a car anymore.
  • The Result: The AI and humans agreed on this. If you changed the object’s shape (like removing a part or deforming it), both groups were much quicker to say, "No, that’s not a Dax anymore." If you just changed the color or added some background noise, both groups were more likely to say, "Yes, it’s still a Dax."

4. Finding #3: AI is Too Generous (The "Over-Generalization" Problem)

This is the biggest difference between humans and AI. While they agreed on what matters (shape vs. color), they disagreed on how strict to be.

  • The Analogy: Imagine you teach a dog to sit. If the dog sits on a cushion, it’s a "sit." If the dog sits on a rock, it’s a "sit." But if the dog stands on its hind legs and waves, a human would say, "That’s not sitting!" An AI, however, might shrug and say, "Well, it’s close enough to sitting, so I’ll call it a sit."
  • The Result: The AI over-generalized. It kept calling heavily distorted objects by the new name long after humans had rejected it. Humans were picky; if the object changed too much, they said, "That’s not it." The AI was loose; it kept saying, "Yeah, that’s still the same thing," even when the object was barely recognizable.

5. The "Essentialist" Gap

The paper suggests that humans have an "essentialist" view. We think objects have a hidden "core identity." If you change the surface (color/texture), the core is still there. But if you change the structure (shape/parts), the core is gone.

  • The AI’s View: The AI seems to rely more on visual similarity. It thinks, "It looks 60% like the original, so it’s the same." It doesn’t have that deep sense of "core identity" that humans do.

In Summary

  • Can AI learn new words? Yes, but it’s harder if the object looks like something it already knows.
  • Does AI understand shape vs. color? Yes, it prioritizes shape, just like humans do.
  • Is AI as picky as humans? No. AI is too forgiving. It will call a heavily mutated object by the same name long after a human would have said, "No, that’s something else entirely."

The researchers released their dataset (NVRD) so other scientists can keep studying this gap between how humans and machines understand the world.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →