Shared Lexical Task Representations Explain Behavioral Variability In LLMs
The paper demonstrates that LLM prompt sensitivity can be explained by the activation levels of "lexical task heads"—shared internal mechanisms that represent the task regardless of whether the prompt uses instructions or examples—and that performance variability stems from how strongly these heads are triggered or diluted by competing representations.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a very talented, but slightly moody, assistant how to perform different tasks.
Sometimes, if you give them a clear instruction (like a recipe), they do a perfect job. Other times, if you just show them a few examples (like showing them pictures of finished dishes), they might get confused or perform inconsistently. You might think the assistant is just being unpredictable, but this paper reveals that there is actually a very logical "internal wiring" causing this behavior.
Here is the breakdown of the paper using a "Kitchen & Chef" analogy.
1. The Discovery: The "Recipe Labels" (Lexical Task Heads)
The researchers discovered that inside the AI's "brain," there are specific parts called Lexical Task Heads.
Think of these as sticky notes that the chef (the AI) slaps onto their tools. If the task is to translate words, the chef doesn't just start translating; first, a specific part of their brain grabs a sticky note that literally says "Translation Mode" or "Language Switch."
The researchers found that these "sticky notes" are written in actual words that humans can read. When they looked at the AI's internal signals, they saw the AI's brain literally shouting words like "Opposite," "Plural," or "Capital City" before it even gave the final answer.
2. The Shared Secret: One Brain, Many Styles
The big mystery was: Does the AI use different brain parts for a written instruction versus a visual example?
It turns out, no. Whether you give the AI a written command ("Translate this...") or a list of examples ("Apple Manzana"), the AI eventually reaches for the exact same sticky notes.
It’s like a chef: whether you give them a written recipe or just show them a photo of a cake, they eventually reach for the same "Baking Mode" mental switch. This means the AI has a "core understanding" of a task that exists independently of how you ask it.
3. Why the AI Fails: The "Weak Signal" and "Distractions"
If the AI has these shared "sticky notes," why does it still mess up? The paper offers two main reasons:
- The Faint Sticky Note (Insufficient Activation): Sometimes, the prompt you use is "weak." It doesn't trigger the "Baking Mode" sticky note strongly enough. The note is there, but it's a tiny, blurry scribble. Because the signal is faint, the chef doesn't quite "get into the zone," and they make mistakes. The researchers proved this by "turning up the volume" on these sticky notes, which actually fixed the AI's mistakes!
- The Wrong Recipe (Competing Tasks): Sometimes, a prompt is ambiguous. Imagine you tell a chef, "Make something with flour." They might make bread, or they might make a cake. If the prompt is slightly confusing, the AI might accidentally grab the "Cake Mode" sticky note when you actually wanted "Bread Mode." The "wrong" signal drowns out the "right" one.
4. The "Many-Shot" Magic: Why more examples help
You’ve probably noticed that giving an AI 20 examples works better than giving it just one. The researchers found the reason: Accumulation.
Every time you provide an example, it’s like hitting the "Task Mode" button harder and harder. Each example adds more "ink" to that sticky note, making the signal louder and clearer until the chef is 100% certain which mode they are in.
Summary in a Nutshell
The paper tells us that AI "prompt sensitivity" isn't just random magic. It’s a mechanical process where different ways of asking a question trigger internal task-labels with different levels of strength.
The takeaway for humans: If your AI is acting up, it’s likely because your prompt isn't "loud" enough to trigger the right internal sticky note, or it's accidentally triggering a different one!
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.