← Latest papers
💬 NLP

Beyond Decodability: Reconstructing Language Model Representations with an Encoding Probe

This paper introduces an Encoding Probe that reconstructs language model representations from interpretable features, offering a complementary perspective to traditional decoding probes by enabling direct feature comparison and revealing how different linguistic and acoustic attributes contribute to internal model states.

Original authors: Gaofei Shen, Martijn Bentum, Tom Lentz, Afra Alishahi, Grzegorz Chrupała

Published 2026-05-04
📖 5 min read🧠 Deep dive

Original authors: Gaofei Shen, Martijn Bentum, Tom Lentz, Afra Alishahi, Grzegorz Chrupała

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a super-smart robot that reads books and listens to stories. You want to know what's happening inside its brain. For a long time, scientists used a method called a "Decoding Probe."

Think of the Decoding Probe like a detective interrogating a suspect. The detective looks at the robot's brain activity and asks, "Can you tell me who the speaker is?" or "Can you tell me what part of speech this word is?" If the detective gets the right answer, they say, "Aha! The robot knows this!"

But the authors of this paper say this detective method has two big flaws:

  1. The "Score Comparison" Problem: Imagine the detective asks two questions.

    • Question A: "Who is speaking?" (The robot gets this right 95% of the time).
    • Question B: "What is the sound of this letter?" (The robot gets this right 58% of the time).
    • The old method says, "The robot knows the speaker better!" But wait! Maybe Question A was just an easier test to begin with. Maybe there are only 10 speakers but 100 different sounds. You can't compare the scores directly because the tests weren't fair. It's like comparing a runner's time in a 100-meter sprint to a marathon; the numbers don't tell you who is "more important" to the runner's overall fitness.
  2. The "Copycat" Problem: Imagine the detective asks, "What is the grammar of this sentence?" The robot answers correctly 85% of the time. But what if the robot isn't actually thinking about grammar? What if it's just recognizing the words? Since certain words almost always appear with certain grammar (like "the" is usually a noun), the robot might be cheating by just looking at the word. The detective thinks the robot understands grammar, but it's just a lucky guess based on the vocabulary.

The New Solution: The "Encoding Probe"

To fix this, the authors invented a new tool called the Encoding Probe. Instead of a detective interrogating the robot, imagine you are a chef trying to recreate a secret soup.

  • The Old Way (Decoding): You taste the soup and guess, "Is there salt in here?" "Is there pepper?"
  • The New Way (Encoding): You take a list of ingredients (Speaker, Sound, Grammar, Words) and try to rebuild the soup from scratch using a recipe. You mix the ingredients together and see how close your new soup tastes to the original robot's "soup" (its brain activity).

Here is how the Encoding Probe solves the problems:

  1. Fair Comparisons: Now, instead of guessing scores, you measure how much of the soup's flavor is missing if you leave out an ingredient.

    • If you leave out "Speaker" and the soup tastes almost the same, the Speaker ingredient wasn't very important.
    • If you leave out "Sound" and the soup tastes terrible, the Sound ingredient was crucial.
    • This gives you a direct, fair comparison of how much each ingredient contributes to the final dish.
  2. Stopping the Copycat: Because you are rebuilding the soup with all the ingredients at once, you can see what happens if you remove just one.

    • If you remove "Grammar" but the soup still tastes perfect because "Words" are still there, you know the robot was just relying on the words, not the grammar.
    • If you remove "Grammar" and the soup gets worse even though the words are still there, you know the robot actually does use grammar independently.

What They Found

The authors tested this new "soup recipe" method on robots that read text and robots that listen to speech.

  • The Speaker Test: They looked at whether the robots cared about who was speaking or what they were saying.

    • The Result: In a robot trained to identify speakers, the "Speaker" ingredient became the main flavor of the soup. But in a robot trained to recognize speech sounds, the "Speaker" ingredient was barely noticeable, even though the robot could still guess the speaker if asked (the old detective method). The new method showed that for the speech robot, the sound of the words mattered way more than who said them.
  • The Grammar Test: They looked at whether the robots cared about sentence structure (grammar) or just vocabulary (words).

    • The Result: Both text and speech robots used words and grammar to build their understanding. However, the "Words" ingredient was the biggest flavor contributor. But crucially, even after accounting for the words, the "Grammar" ingredient still added a unique flavor. This proved the robots do understand grammar, not just the words.

The Big Takeaway

The paper concludes that the old "Detective" method (Decoding) is good at telling you if a robot has a piece of information. But the new "Chef" method (Encoding) is much better at telling you how important that information is compared to everything else, and whether the robot is truly using it or just copying it from something else. It gives a clearer, fairer picture of what's really cooking inside the robot's brain.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →