← Latest papers
💬 NLP

Vision-Language Models Align with Human Neural Representations in Concept Processing

This study demonstrates that vision-language models generally align better with human neural representations of concepts than language-only models, though the extent of this alignment depends on the specific architecture and whether it stems from genuine concept learning or sensitivity to inference-time context.

Original authors: Anna Bavaresco, Marianne de Heer Kloots, Sandro Pezzelle, Raquel Fernández

Published 2026-01-23
📖 5 min read🧠 Deep dive

Original authors: Anna Bavaresco, Marianne de Heer Kloots, Sandro Pezzelle, Raquel Fernández

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine your brain as a massive, bustling library where every concept you know—like "bird," "dance," or "crazy"—has its own unique shelf. When you see a picture of a bird or read a sentence about one, your brain lights up in a specific pattern, like a unique constellation of stars forming in the dark.

For a long time, scientists have been trying to build computer models that can "think" like humans. Recently, we've built Vision-Language Models (VLMs). Think of these as super-smart students who have studied both books (text) and picture books (images) simultaneously, whereas older models only read books.

This paper asks a simple question: Do these new, multi-sensory students understand concepts the same way our brains do?

Here is the breakdown of their findings, using some everyday analogies:

1. The Experiment: The "Guess the Pattern" Game

The researchers didn't just ask the models to describe a picture. Instead, they played a matching game.

  • The Human Side: They looked at brain scans (fMRI) of people who were shown words like "bird" either inside a sentence ("The bird flew around the cage") or next to a picture of a bird. They mapped out the "constellation" of brain activity for each word.
  • The Model Side: They fed the exact same words and pictures into various AI models and looked at the model's internal "thought patterns" (mathematical representations).
  • The Match: They compared the human brain constellations with the AI's constellations. The closer the match, the more "human-like" the AI's understanding is.

2. The Main Finding: Two Heads (and Eyes) Are Better Than One

The study found that the Vision-Language Models generally matched human brain patterns better than the text-only models.

  • The Analogy: Imagine trying to guess what a "dog" is.
    • The Text-Only Model is like someone who has only read a dictionary definition of a dog. They know the words, but they've never seen a dog.
    • The Vision-Language Model is like someone who has read the definition and seen a thousand photos of dogs.
    • When the researchers asked, "Which one thinks about a dog more like a human brain does?" the answer was usually the one who had seen the photos. The models that could "see" and "read" together created concept maps that looked more like our own neural maps.

3. The Twist: Not All "Smart" Models Are Created Equal

This is where it gets interesting. The researchers tested different types of these multi-sensory models, and they didn't all perform the same.

  • The "Encoders" (The Careful Students): Models like LXMERT and VisualBERT were the best at matching human brains. These models are designed to carefully analyze the relationship between a word and an image, layer by layer. They seem to have learned to build concepts in a way that feels very human.
  • The "Generators" (The Creative Writers): Newer, more powerful models like LLaVA and IDEFICS2 are famous for being able to write stories or generate images. You might think they would be the smartest, but they actually matched human brain patterns less well than the "Encoders."
    • The Analogy: Think of the "Encoders" as a museum curator who deeply understands the history and meaning of every artifact. Think of the "Generators" as a talented improv comedian who can make up a story about the artifact on the spot. The comedian is impressive and useful, but their internal "understanding" of the object doesn't quite match the curator's (or the human brain's) deep, structural knowledge.

4. The Context Trap: Are They Learning or Just Copying?

The researchers wanted to know: Did these models actually learn human-like concepts during their training, or are they just good at using the picture they are shown right now?

They ran a "surgery" on the models (called an ablation study):

  • The Test: They took the models that usually need a picture and forced them to work without one, giving them only the word.
  • The Result:
    • Some models (like CLIP and VisualBERT) crashed. Their "human-like" understanding disappeared when the picture was gone. This suggests they were relying heavily on the immediate visual input, almost like a parrot mimicking a sound only when it sees the speaker.
    • Other models (like LXMERT and IDEFICS2) stayed strong. Even without the picture, their internal concept maps still looked very human. This suggests they actually learned a deep, human-like understanding of concepts during their training, not just a trick for using pictures.

5. The Bottom Line

The paper concludes that multimodality (combining sight and text) helps AI models align with human brains, but it's not a magic bullet.

  • Simply adding a picture to a text model doesn't automatically make it "human."
  • The architecture (the internal design) matters most. Some designs (the "Encoders") naturally build human-like concept maps.
  • Being good at generating text or images (the "Generators") doesn't necessarily mean the model understands concepts the way a human brain does.

In short: To build an AI that thinks like a human, you don't just need to give it eyes; you need to build its brain in a specific way that allows it to truly integrate what it sees with what it reads.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →