← Latest papers
🧬 biology

Revealing the core dimensions underlying representations in brains, behavior and AI

This paper introduces Similarity-Based Representation Factorization (SRF), a general computational method that extracts low-dimensional, interpretable, and non-negative embeddings from similarity matrices across neuroscience, psychology, and AI datasets, thereby overcoming limitations in current approaches to uncover the core dimensions underlying representations.

Original authors: Florian P. Mahner, Ka Chun Lam, Francisco Pereira, Martin N. Hebart

Published 2026-05-27
📖 5 min read🧠 Deep dive

Original authors: Florian P. Mahner, Ka Chun Lam, Francisco Pereira, Martin N. Hebart

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Imagine you are trying to understand how a library organizes its books. You don't have the catalog; you only have a list of which books people say are "similar" to each other. Maybe people say "a cat is like a dog" and "a car is like a truck," but they don't say "a cat is like a car."

For a long time, scientists studying the brain, human behavior, and artificial intelligence have been stuck with just that list of similarities. They could see that things were grouped together, but they couldn't easily see why. They knew a cat and a dog were neighbors in the library, but they didn't know if they were neighbors because they are both animals, because they have fur, or because they are small.

This paper introduces a new tool called SRF (Similarity-Based Representation Factorization). Think of SRF as a magical "similarity decoder" that takes that messy list of "what looks like what" and breaks it down into the specific, understandable reasons why they look alike.

Here is how it works, using some everyday analogies:

1. The Problem: The "Black Box" of Similarity

Imagine you have a giant, tangled ball of yarn. You know that certain strands are knotted together, but you can't see the individual threads.

  • Old Methods: Previous tools could untangle the ball a little bit to show you the general shape (like a big knot of "animals" and a knot of "vehicles"), but they couldn't tell you which specific thread was "fur" and which was "wheels." Or, if they did find the threads, they were often so abstract (like "Thread #4") that no one knew what they meant.
  • The Missing Pieces: Often, the list of similarities is incomplete. Maybe you only asked people about 10% of the possible book pairs. Old tools would either get confused by the missing data or try to guess (impute) the missing links, which often led to wrong conclusions.

2. The Solution: SRF as a "Community Detective"

The authors created SRF to solve this. Imagine SRF is a detective who looks at the ball of yarn and says, "I don't need to see every single knot to figure out the pattern."

  • Finding the Communities: SRF looks at the web of similarities and finds "communities." If a lion loads heavily onto a "furry" dimension and a "wild" dimension, SRF sees that. If a ball loads onto "round" and "playful," it sees that too.
  • Soft Membership: Unlike a strict rule where a book must be only in the "Fiction" section, SRF allows for "soft membership." A book can be 80% "Mystery" and 20% "Historical." This matches how our brains actually work; things aren't always black and white.
  • Handling Missing Data: This is SRF's superpower. Even if you only have 1% of the similarity data (like a very sparse map), SRF can still figure out the underlying structure without guessing the missing pieces. It learns the pattern from what is there and fills in the gaps logically, rather than just making up numbers.

3. Why It's Better Than the Old Way (RSA)

Scientists often used a method called RSA (Representational Similarity Analysis) to test ideas.

  • The RSA Analogy: Imagine you want to prove that "Animacy" (living vs. non-living) is important. With RSA, you compare your whole tangled ball of yarn against a hypothesis. But the ball is also tangled with "Size," "Color," and "Shape." The signal for "Animacy" gets drowned out by all the other noise. It's like trying to hear a whisper in a rock concert.
  • The SRF Advantage: SRF first untangles the yarn into separate threads (dimensions). Now, you can pick up just the "Animacy" thread and test it. Because you isolated it from the noise, you can hear the whisper clearly. The paper shows that SRF is much better at finding these specific, hidden patterns, especially when the data is noisy or incomplete.

4. What They Found (The Results)

The researchers tested SRF on many different types of data:

  • Human Brains: Looking at how brain cells react to pictures.
  • Human Behavior: Asking people what words or objects feel similar.
  • AI Models: Looking at how computer programs (like the ones that generate images) "see" the world.

The Results:

  • It Works Everywhere: SRF found clear, understandable dimensions in all these areas. For example, in word associations, it found dimensions like "food-related," "animal-related," and even abstract ones like "physical strength" or "affection."
  • It Predicts Behavior: When they took the dimensions SRF found from word associations, they could accurately predict how humans would rate words on things like "how happy a word sounds" or "how big an object is." This proves the dimensions aren't just math tricks; they capture real meaning.
  • It's Stable: Even if you run the math 30 times with different starting points, SRF finds the same core dimensions every time.

The Bottom Line

This paper presents a new way to look at how brains, people, and machines organize information. Instead of just saying "these things are similar," SRF tells you the specific reasons they are similar. It turns a blurry, tangled picture of similarity into a clear, organized map of the dimensions that matter—like animacy, size, color, or emotional tone—even when you only have a tiny, incomplete amount of data to start with.

It's like going from having a blurry photo of a crowd to having a high-definition list of exactly who is in the crowd and what they are wearing.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →