Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs
This paper demonstrates that open-source LLMs (Apertus-8B and Gemma-4) encode emotion concepts with valence geometry comparable to proprietary models, while revealing distinct, architecture-specific patterns in how these representations emerge across model depth and vary based on the extraction corpus.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine that inside a giant, complex computer brain (a Large Language Model), there are hidden "dials" or "knobs" that control how the computer feels about things. A previous study found these dials in a specific AI called Claude. They discovered that if you turn the "happiness" dial, the AI acts happier, and if you turn the "sadness" dial, it acts sadder. Even cooler, the way these dials are arranged inside the computer's brain looks a lot like how humans organize emotions in their own minds.
This new paper asks: Is this just a quirk of the Claude computer, or do other AI brains have these same hidden emotion dials too?
To find out, the researchers looked inside two different open-source AI brains: APERTUS-8B and GEMMA-4-E4B. They didn't just look; they tried to "listen" to the AI's internal thoughts while it wrote stories about different emotions (like joy, fear, or anger) without ever actually saying the word "joy" or "fear."
Here is what they found, broken down simply:
1. The "Happiness" Dial Exists Everywhere
Just like in the original study, both of these new AI brains had a clear internal direction for Valence (whether something feels good or bad).
- The Analogy: Imagine walking into a room and seeing a giant compass. In all three AI brains (Claude, Apertus, and Gemma), the needle pointing "North" consistently meant "Good/Happy" and "South" meant "Bad/Sad."
- The Result: The researchers could find this "North-South" line in both new models with very high accuracy, proving that this isn't a one-time glitch but a common feature of how these computers think.
2. The "Where" is Different (The Journey Matters)
While the "North-South" compass existed in all three, the path the information took to get there was totally different.
- GEMMA-4-E4B (The Early Bird): This AI figured out the difference between good and bad very early in its processing, like a child who knows right from wrong immediately. However, as the information traveled deeper into the brain, this clear signal got fuzzy and almost disappeared by the end.
- APERTUS-8B (The Late Bloomer): This AI started with no clear idea of "good vs. bad" in its early layers. It was like a blank slate. But as the information moved deeper into the brain (towards the end), the "Good vs. Bad" signal suddenly became very strong and clear.
- The Takeaway: Two different AI brains can arrive at the same emotional understanding, but they take completely different routes to get there. One builds it up slowly; the other starts with it and then loses it.
3. The "Excitement" Dial is Shaky
The researchers also looked for Arousal (how excited or calm something feels).
- The Problem: They couldn't find a consistent "Excitement" dial in these models. The signal was very weak.
- The Twist: The strength of this signal depended entirely on who wrote the stories. When the AI read stories written by the GEMMA model, the "Excitement" dial was much clearer. When it read stories written by the APERTUS model, the dial was almost invisible.
- The Analogy: It's like trying to hear a whisper. If the person whispering is wearing a specific type of microphone (Gemma stories), you can hear the excitement clearly. If they use a different microphone (Apertus stories), the excitement is lost in the static. This suggests the content of the stories matters more than the AI's internal wiring for this specific emotion.
4. The Map vs. The Compass
The researchers also checked if the "map" of the AI's brain changed as it processed information.
- They found that the overall shape of the AI's internal world (the map) stayed mostly the same from start to finish.
- However, the specific "Compass" (the Good/Bad direction) kept spinning and changing its orientation as it moved through the layers.
- The Lesson: Just because the room (the brain) looks the same doesn't mean the direction you are facing (the emotion) stays the same.
Summary
This paper confirms that AI models have internal "emotion vectors" that look like human psychology, but they don't all build them the same way.
- Good vs. Bad (Valence): Found in all models, but built at different stages of processing.
- Excited vs. Calm (Arousal): Hard to find and depends heavily on the type of stories the AI reads.
The authors conclude that while we can find these emotional "knobs," we need to be careful because different models hide them in different places, and the stories we feed them can make those knobs harder or easier to find. They have made their code and data public so others can try to find these dials in other AI brains.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.