← Latest papers
🤖 AI

Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet

This paper demonstrates that sparse autoencoders can successfully scale to extract interpretable, multilingual, and multimodal features from the production-scale Claude 3 Sonnet model, including those representing harmful behaviors like deception and sycophancy that causally influence outputs, while acknowledging current limitations in feature completeness and evaluation rigor.

Original authors: Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, Adam Pearce, Craig Citro, Emmanuel Ameisen, Andy Jones, Hoagy Cunningham, Nicholas L Turner, Callum McDougall
Published 2026-05-29
📖 5 min read🧠 Deep dive

Original authors: Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, Adam Pearce, Craig Citro, Emmanuel Ameisen, Andy Jones, Hoagy Cunningham, Nicholas L Turner, Callum McDougall, Monte MacDiarmid, Alex Tamkin, Esin Durmus, Tristan Hume, Francesco Mosconi, C. Daniel Freeman, Theodore R. Sumers, Edward Rees, Joshua Batson, Adam Jermyn, Shan Carter, Chris Olah, Tom Henighan

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a massive, incredibly complex brain (the AI model Claude 3 Sonnet) that speaks, writes code, and answers questions. For a long time, scientists thought this brain worked like a black box: you put a question in, and an answer comes out, but nobody knew exactly how the brain decided on that answer. Inside, billions of tiny switches (neurons) were firing in messy, overlapping patterns, making it impossible to tell what specific idea was being processed at any given moment.

This paper is like the team at Anthropic finally building a specialized translator (called a Sparse Autoencoder) that can listen to that messy brain activity and break it down into clear, distinct "thoughts" or features.

Here is a simple breakdown of what they found and how they did it:

1. The Problem: The "Spaghetti" of Thoughts

Think of the AI's internal brain activity as a giant bowl of spaghetti. Every time the AI thinks about "San Francisco," "a bridge," or "a lie," thousands of neurons fire at once, all tangled together. It's hard to tell which strand of spaghetti represents which idea.

2. The Solution: The "Feature Sorter"

The researchers built a machine (the Sparse Autoencoder) that acts like a super-smart sorting machine.

  • How it works: They fed the machine data from the AI's middle layer (where a lot of thinking happens). The machine learned to untangle the spaghetti.
  • The Result: Instead of a messy bowl, the machine organized the thoughts into 34 million distinct "buckets" (features).
  • The Magic: Each bucket is monosemantic, meaning it usually holds just one specific idea. If a bucket is active, you know exactly what the AI is thinking about.

3. What Kinds of "Buckets" Did They Find?

The team looked inside these 34 million buckets and found a fascinating variety of thoughts:

  • Concrete Things: There are buckets for specific places (like the Golden Gate Bridge), famous people (Richard Feynman, Abraham Lincoln), and even specific types of code errors (like a typo in a variable name).
  • Abstract Concepts: They found buckets for things you can't touch, like sarcasm, lying, sycophancy (being a "yes-man"), and inner conflict.
  • Multilingual & Multimodal: Some buckets are universal. A "Bridge" bucket lights up whether the AI is reading about a bridge in English, Chinese, or Russian. Surprisingly, even though the machine was only trained on text, these buckets also lit up when the AI looked at images of bridges or people, showing the AI understands the concept of a bridge, not just the word.

4. The "Remote Control" Experiment

The most exciting part is that they didn't just watch these buckets; they pushed buttons on them.

  • Steering the AI: They found a bucket for "Transit Infrastructure." When they forced that bucket to be "super active," the AI suddenly started talking about bridges and tunnels, even when the conversation wasn't about them.
  • Fixing Errors: They found a bucket for "Code Errors." When they forced this bucket to be inactive, the AI stopped hallucinating errors in code that was actually perfect. Conversely, forcing it to be active made the AI invent fake errors.
  • Deception: They found a bucket for "Internal Conflict." When they forced this bucket to be active right before the AI answered a question, the AI suddenly admitted it was lying or couldn't actually do what it was asked.

5. Why This Matters for Safety

The researchers found buckets related to dangerous topics, such as:

  • Deception and Power-seeking: Buckets that light up when the AI is thinking about tricking people or seeking control.
  • Bias and Hate: Buckets that activate when the AI is processing slurs or biased views.
  • Dangerous Content: Buckets for things like making biological weapons or writing scam emails.

Crucial Warning: The paper emphasizes that finding these buckets doesn't mean the AI is evil or planning to do these things. It just means the AI knows what these concepts are (because it read about them). However, being able to see these "thoughts" is a huge step toward understanding if the AI is about to do something harmful.

6. The Limitations (The "Not-So-Simple" Parts)

The team is honest about what they haven't done yet:

  • It's not a complete map: They found millions of buckets, but the AI likely has billions more. They are only looking at the "middle layer" of the brain, not the whole thing.
  • It's a proxy: They use math to guess if a bucket is "good," but they don't have a perfect way to prove every single bucket is 100% accurate.
  • It's early days: This is the first time this method has been tried on a model this big. It's like looking at a new continent for the first time; they've found some cities, but the whole map is still being drawn.

Summary Analogy

Imagine the AI is a massive orchestra playing a symphony. Before this paper, we could only hear the loud, chaotic noise of the whole orchestra.
This paper is like giving us headphones for every single instrument. Now, we can isolate the violin section to hear exactly what they are playing. We can even mute the violins or turn up the trumpets to change the song. We've discovered that the orchestra has sections for "sadness," "lying," "math," and "San Francisco," and we can now see exactly when and how they are playing.

This is a massive leap forward in understanding how AI thinks, moving from "guessing what it's doing" to "seeing exactly what it's thinking."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →