← Latest papers
💬 NLP

Patches of Nonlinearity: Instruction Vectors in Large Language Models

This paper investigates how large language models process instructions by identifying "Instruction Vectors" (IVs) that act as circuit selectors, demonstrating that while these representations are localized and linearly separable, they interact with the model's internal pathways through complex non-linear causal mechanisms.

Original authors: Irina Bigoulaeva, Jonas Rohweder, Subhabrata Dutta, Iryna Gurevych

Published 2026-02-10
📖 4 min read☕ Coffee break read

Original authors: Irina Bigoulaeva, Jonas Rohweder, Subhabrata Dutta, Iryna Gurevych

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are watching a world-class chef prepare a complex dish. You see them chop onions, sear meat, and simmer a sauce. You know they are making "Beef Bourguignon," but you don't quite know how their brain is organizing the steps. Are they following a mental recipe card, or are they just reacting to the smell of the garlic as it hits the pan?

This paper, "Patches of Nonlinearity," is essentially a "brain scan" for Large Language Models (like ChatGPT). The researchers wanted to know: When you give an AI an instruction (like "Translate this to French"), how does it actually "hold" that instruction in its mind while it works?

Here is the breakdown of their discovery using three simple analogies.


1. The "Mental Sticky Note" (Instruction Vectors)

The researchers found that as soon as an AI reads an instruction, it creates a tiny, concentrated "summary" of that task.

The Analogy: Imagine you walk into a library. Before you even look for a book, you scribble "Find a mystery novel" on a sticky note and slap it on your forehead. You don't need to keep re-reading the sign at the entrance; the "essence" of your mission is right there, attached to you, ready to guide every step you take.

In AI terms, they call these Instruction Vectors (IVs). They are localized "digests" of the task that the model creates early on, which then guide the rest of the process.

2. The "Synergistic Jazz Band" (Superadditivity)

This is the most surprising part of the paper. Usually, scientists assume that if Part A of an AI does 5 units of work and Part B does 5 units of work, together they do 10 units. This is called "linear" thinking.

But the researchers found that instruction-following is non-linear (superadditive).

The Analogy: Think of a Jazz Band. If a drummer plays alone, it’s just a beat. If a pianist plays alone, it’s just a melody. But when they play together, they don't just produce "Beat + Melody." They create a "Groove"—a third, magical thing that is much greater than the sum of its parts.

The AI's internal layers work like that jazz band. One layer might hold a piece of the instruction, and another might hold a piece of the logic, but they only truly "click" and solve the task when they interact in a complex, non-linear way. You can't understand the "groove" by looking at the drummer and pianist in isolation.

3. The "Circuit Selector" (The Mechanism)

Finally, the researchers wanted to know what those "sticky notes" actually do. Do they just sit there, or do they change how the AI thinks?

The Analogy: Imagine a massive Electrical Switchboard in a city. There are millions of wires (paths) that can carry electricity. Without an instruction, the power is just flowing randomly. But the Instruction Vector acts like a Master Switch.

Once the "sticky note" is created, it tells the model: "Hey, we are doing a math task! Flip switches 5, 12, and 89, and ignore the rest!" The instruction vector acts as a Circuit Selector, picking out the specific "electrical pathways" needed to solve that specific problem while turning off the pathways used for poetry or coding.


Why does this matter?

If we want to make AI safer and more reliable, we can't just treat it like a "black box" that spits out answers. We need to know how it "thinks."

By proving that instructions are stored as these "superadditive" sticky notes that select specific circuits, the researchers have given us a new map. It tells us that if we want to fix an AI's mistakes, we can't just tweak one single "neuron"—we have to understand how the whole "jazz band" is playing together.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →