← Latest papers
💬 NLP

Persistent Sparse Autoencoders: Learning Feature Timescales in Language Models

This paper introduces Persistent Sparse Autoencoders, an extension of standard SAEs that learns feature-specific persistence coefficients to capture both fast, local detectors and slow, topic-level semantic information, thereby enabling more effective interpretation and monitoring of language models over long contexts.

Original authors: Haoyan Luo, Mateo Espinosa Zarlenga, Mateja Jamnik

Published 2026-07-21
📖 3 min read☕ Coffee break read

Original authors: Haoyan Luo, Mateo Espinosa Zarlenga, Mateja Jamnik

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to understand a giant, super-smart robot that writes stories, answers questions, and solves problems. This robot, called a Large Language Model (LLM), doesn't just think one word at a time; it holds a massive amount of information in its "brain" as it reads a sentence. Scientists have been trying to peek inside this brain to see what the robot is thinking about. They use a tool called a "Sparse Autoencoder" (SAE). Think of an SAE like a high-tech translator that breaks the robot's complex, messy thoughts into a list of simple, distinct "switches." When the robot thinks about "math," a specific math-switch flips on. When it thinks about "cats," a cat-switch flips.

For a long time, these translators had a weird blind spot. They treated every single word the robot read as if it were brand new, completely ignoring the words that came before it. It was like watching a movie but only looking at one frame at a time, forgetting the plot you just saw. This made it hard to understand things that last a while, like a story's theme, a character's mood, or a specific topic that the robot is discussing for several paragraphs. The big question was: Can we teach these translators to remember the "vibe" of the conversation, not just the individual words?

This is where a new team of researchers steps in with a clever upgrade they call "Persistent Sparse Autoencoders" (Persistent SAEs). They realized that while the standard translators were ignoring the past, the robot's brain actually was remembering it. The new tool doesn't just look at the current word; it learns to keep a "memory state" for each switch. Some switches are designed to be "fast" and forgetful, only caring about the word right in front of them. Others are "slow" and persistent, acting like a sticky note that stays on the robot's desk for a long time, tracking the overall topic.

The researchers found that this new method works beautifully. It keeps the ability to explain individual words just as well as the old tools, but it also creates a special group of "slow switches" that hold onto the big picture. In their tests, these slow switches were so good at tracking topics that they could tell the difference between a history lesson and a math problem just by looking at the robot's internal state, even without being told what the topics were.

Perhaps most excitingly, they tested this on a safety scenario called "prompt injection." Imagine someone tries to trick the robot into doing something bad by hiding a secret instruction inside a long email. The robot might read the email, ignore the secret for a while, and then act on it hundreds of words later. The old tools would forget the secret instruction as soon as the robot moved past the email. But the new "slow" switches kept the warning signal alive, acting like a persistent alarm that stayed loud even after the danger had passed. This suggests that by giving these translators the ability to learn how long to remember things, we can finally start to understand and monitor the long-term thoughts of AI, keeping them safe and predictable even in very long conversations.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →