Retrieval is Enough: Training-Free Interpretability with a Tool-Using Agent
The paper introduces HARP, a training-free, tool-using agent that leverages retrieval from activation databases to generate and validate hypotheses, demonstrating that it can outperform expensive training-based interpretability methods in concept discovery, detection, steering, and secret elicitation.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to understand how a giant, super-smart robot brain works. This brain, called a neural network, doesn't think in words like we do; it thinks in invisible clouds of numbers called "activations." For a long time, scientists have tried to peek inside these clouds to see what the robot is actually thinking about. Some researchers tried to build a massive, expensive library of these number-clouds, training a new, specialized robot to memorize them and explain them. Others tried simpler tricks, like using a magnifying glass to find patterns without any training at all. The big question everyone is asking is: Do we really need to build those expensive, trained libraries to understand the brain? Or is it possible that the "smart" trained robots are just really good at looking up answers in a book they already memorized, rather than actually discovering anything new?
This paper, titled "Retrieval is Enough," proposes a clever new way to peek inside the robot's brain without building any expensive libraries. The authors created a tool called HARP (Hypothesis-driven Agentic Retrieval and Probing). Think of HARP not as a student who has to study for years to learn the material, but as a super-smart detective with a magic phone book. This phone book contains millions of examples of the robot's thoughts paired with the sentences that caused them. When HARP wants to know what a specific thought means, it doesn't guess; it looks up similar thoughts in the phone book, reads the surrounding sentences, and figures out the pattern. It then uses simple math tools to "subtract" that pattern from the thought to see what's left, repeating the process until it has peeled back every layer of meaning.
The paper finds that this "detective with a phone book" is surprisingly powerful. In fact, HARP often does a better job than the expensive, trained robots at discovering what concepts are hidden inside the neural network's activations. Whether the task is figuring out the main topics in a paragraph, spotting a specific hidden idea, or even tricking a model into revealing a secret word it was trained to hide, HARP performs just as well, and sometimes better, than the heavy-duty trained methods. The authors suggest that many of the "insights" we thought we were getting from expensive training might just be the result of the system retrieving and recombining patterns it saw during its training, rather than unlocking deep, new secrets. By showing that a simple, training-free retrieval system can do the job, the paper argues that we might not need to spend so much time and money training complex interpreters if we can just look up the answers in a well-organized database of examples.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.