Hyperdimensional Probe: Decoding LLM Representations via Vector Symbolic Architectures
This paper introduces the Hyperdimensional Probe, a hybrid method leveraging Vector Symbolic Architectures to unify input-focused feature extraction and output-oriented analysis, thereby overcoming the limitations of existing interpretability techniques to provide deeper semantic insights into Large Language Model representations.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you're trying to figure out how a super-smart robot thinks. You've probably heard of Large Language Models (LLMs), the AI brains behind chatbots that can write stories, solve math problems, and chat like humans. But here's the catch: even though these robots are amazing, they are also "black boxes." We can see what goes in (a question) and what comes out (an answer), but the messy, magical thinking happening inside? That's a mystery. Scientists have been trying to peek inside using two main tools. One tool, called a "probe," acts like a translator, trying to guess what the robot is thinking based on the words it sees. The other tool, a "Sparse Autoencoder" (SAE), tries to break the robot's thoughts down into a long list of simple, separate ingredients, like a recipe. But both tools have a blind spot. The translator focuses on input-oriented feature extraction but doesn't fully integrate the bigger picture of how the robot actually answers, and the recipe-maker often gets confused by noise or misses the connection to the final output. We care about this because if we don't understand how these AI brains work, we can't fully trust them or fix them when they get things wrong. We need a way to see the whole picture, not just the start or the finish.
Enter the "Hyperdimensional Probe," a new, hybrid detective tool introduced in this paper. Think of it as a super-powered translator that combines the best parts of the old tools while fixing their flaws. The researchers built this tool using something called "Vector Symbolic Architectures" (VSAs). If you imagine the robot's brain as a giant library of ideas, previous tools were either trying to read the book titles (input) or count the pages (output), but they couldn't see how the story connected. The Hyperdimensional Probe, however, uses a special kind of "hypervector algebra"—a fancy way of saying it uses high-dimensional math to mix and match ideas like building blocks. It unifies the "top-down" view of traditional probes with the "sparsity" (or neat, organized list-making) of SAEs.
In their experiments, the authors found that this new approach consistently pulls out meaningful semantic information—basically, it actually understands what the robot is thinking about—across different types of LLMs, different sizes, and various settings. They tested it in two specific scenarios: when the robot was finishing a sentence (input-completion) and when it was answering questions (QA-focused text generation). The results suggest that this method is better at finding clear concepts than the old ways. For instance, it overcomes the limits of looking only at the robot's final word choices (logits), which are stuck with the robot's limited vocabulary. At the same time, it avoids the "noisier" results that often happen when using SAEs in situations where the concepts are supposed to be clear and bounded. By letting researchers look at both the input and the output features at the same time, this work suggests a path toward a deeper, more unified understanding of how neural representations actually work, bridging the gap between how the robot sees the world and how it speaks about it.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.