← Latest papers
💬 NLP

From Associations to Activations: Comparing Behavioral and Hidden-State Semantic Geometry in LLMs

This paper demonstrates that behavioral similarity data, particularly from forced-choice tasks, can effectively recover and predict the hidden-state semantic geometry of large language models, suggesting that external behavioral measurements retain significant information about internal cognitive representations.

Original authors: Louis Schiekiera, Max Zimmer, Christophe Roux, Sebastian Pokutta, Fritz Günther

Published 2026-02-17
📖 5 min read🧠 Deep dive

Original authors: Louis Schiekiera, Max Zimmer, Christophe Roux, Sebastian Pokutta, Fritz Günther

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Question: Can We Read a Robot's Mind by Watching What It Says?

Imagine you have a very smart robot (a Large Language Model, or LLM) that you can't open up. You can't see its brain, its circuits, or its internal "thoughts." You can only talk to it and see what it says back.

The researchers asked: If we watch how this robot answers questions, can we figure out how its internal "brain" is organized?

In the world of psychology, scientists have long used word games to guess how human brains are wired. If you ask a person, "What comes to mind when you hear 'dog'?", they might say "cat," "leash," or "bark." By collecting thousands of these answers, scientists can draw a map of how the human mind connects ideas.

This paper asks: Does the same trick work for AI? Can we map the AI's hidden internal "geometry" (how it actually stores concepts) just by looking at its behavior?


The Experiment: Two Ways to Play the Game

The researchers tested this on 8 different AI models using a shared list of 5,000 words. They used two different "games" to see which one gave a better peek inside the AI's brain.

Game 1: The "Multiple Choice" Quiz (Forced Choice)

  • The Setup: The AI is given a word (e.g., "Dog") and a list of 16 other words. It must pick the two that fit best.
  • The Analogy: Imagine you are at a party, and someone points to a picture of a dog and asks, "Which two of these people are most like a dog?" You have to choose from a specific lineup.
  • The Result: This game worked amazingly well. The way the AI chose its answers matched its internal brain structure almost perfectly. It was like the AI was taking a test that perfectly revealed its internal logic.

Game 2: The "Free-Flow" Chat (Free Association)

  • The Setup: The AI is given a word (e.g., "Dog") and asked to just spit out five words that come to mind. No list, no rules.
  • The Analogy: Imagine the same party, but now someone just yells, "Tell me anything about dogs!" The AI might say "bark," "fur," "pizza," "moon," "blue." It's chaotic and open-ended.
  • The Result: This game was noisy and confusing. The answers were all over the place. While it showed some connection to the AI's brain, the signal was weak and hard to read. It was like trying to hear a specific conversation in a crowded, shouting room.

The "Brain Scan" Comparison

To prove their point, the researchers didn't just guess. They actually looked inside the AI's "brain" (its hidden layers of code) while it was processing these words.

  1. The Internal Map: They took a snapshot of the AI's internal math (its "hidden states") to see how it actually organized the word "Dog" internally.
  2. The Behavioral Map: They built a map based only on the answers the AI gave in the games above.
  3. The Match: They compared the two maps.

The Discovery:

  • When the AI played the Multiple Choice game, the "Behavioral Map" looked almost identical to the "Internal Map." The external behavior was a perfect reflection of the internal reality.
  • When the AI played the Free Chat game, the maps were fuzzy and didn't match up well.

Why Does This Matter?

Think of the AI's internal brain as a library.

  • Free Association is like asking a librarian to "shout out some books about dogs." They might shout "Puppy," "Fire," "Pizza," and "Blue." It's hard to tell how the library is actually organized based on that chaos.
  • Forced Choice is like asking the librarian, "Here are 16 books. Which two are most similar to 'Dog'?" Because the librarian has to compare specific options, their choices reveal the true, organized structure of the library shelves.

The Takeaway

  1. Behavior is a Window: You can learn a lot about an AI's internal "thoughts" just by watching how it behaves, without needing to hack its code.
  2. Structure Matters: The way you ask a question matters more than you think. Constrained questions (like multiple choice) are much better at revealing the truth than open-ended questions (like free chat).
  3. The "Universal" Brain: The researchers also found that different AI models seem to share a similar "mental map." If you ask one AI a multiple-choice question, you can often guess how a different AI would answer, because they all seem to organize their knowledge in a similar way.

In a Nutshell

If you want to understand how an AI thinks, don't just ask it to "chat." Give it a structured test with limited options. That's the key to unlocking the secret map of its mind.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →