← Latest papers
💬 NLP

Circuit Fingerprints: How Answer Tokens Encode Their Geometrical Path

The paper proposes the "Circuit Fingerprint" hypothesis, which posits that answer tokens geometrically encode the directions of the circuits that produce them, allowing for circuit discovery and activation steering through geometric alignment rather than gradient-based or causal intervention methods.

Original authors: Andres Saurez, Neha Sengar, Dongsoo Har

Published 2026-02-11
📖 3 min read☕ Coffee break read

Original authors: Andres Saurez, Neha Sengar, Dongsoo Har

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to understand how a massive, complex city works. There are two main ways you could study it:

  1. The Detective Approach (Circuit Discovery): You follow people around, watch where they walk, and see which intersections they use to get from point A to point B. You are trying to map out the "roads" (the circuits) that people use to complete specific tasks.
  2. The Architect Approach (Activation Steering): You don't care about the history; you just want to change the city. You want to know, "If I want everyone to feel happy, which streets should I decorate with flowers?" You are trying to "steer" the city's vibe.

For a long time, AI researchers thought these were two totally different jobs. But this paper, "Circuit Fingerprints," argues that they are actually two sides of the same coin.

The Big Idea: The "Breadcrumb" Theory

The researchers discovered something fascinating: The destination itself contains a map of the journey.

Think of it like this: If you find a pile of sand in the middle of a desert, you can look at the shape and texture of that sand to figure out exactly which wind currents and dunes it traveled through to get there. The sand carries a "fingerprint" of its path.

In an AI model, the "answer" (like the word "Paris") isn't just a random word that pops out. The way the model's internal "brain" represents the word "Paris" actually contains a geometric signature of all the mathematical "roads" the model traveled to reach that conclusion.

The "Read-Write" Duality

The paper proposes that because the answer holds the map, we can do two things with it:

  • READING (The Detective): We can look at the "fingerprint" of an answer to work backward and find the exact "roads" (the neurons and connections) the model used to get there. This allows us to find the "circuit" without having to do the exhausting work of manually testing every single connection one by one.
  • WRITING (The Architect): Because we now know exactly which "roads" lead to a certain feeling or fact, we can "write" on those roads. If we want the AI to be more joyful, we don't just ask it nicely (prompting); we actually go into its internal "map" and nudge the activations along the "joyful" roads.

Why does this matter? (The Results)

The researchers tested this on several AI models and found two big wins:

  1. It’s Efficient: They could find the "roads" the AI uses to solve puzzles (like identifying names or agreeing with verbs) almost as well as the expensive, slow, traditional methods, but much more simply.
  2. It’s More Powerful: When they tried to make the AI act "emotional," their "steering" method worked much better than just telling the AI, "Please act happy." Their method got the emotion right about 70% of the time, while just asking the AI via text only worked about 53% of the time. Most importantly, the AI stayed smart and factual while it was being emotional.

The Takeaway

This paper suggests that an AI isn't just a black box of random numbers. Instead, it is a geometric landscape. Every thought the AI has leaves a footprint, and if we learn to read those footprints, we don't just understand how the AI thinks—we gain the steering wheel to guide it.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →