Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study
This study analyzes how large language models represent self-harm content across four architectures and two datasets, revealing that such information crystallizes in the final 3–7% of network layers and that the most accurate detection probes do not necessarily rely on the most linearly separable directions, with Gemma-3-4B exhibiting a distinct representational pattern.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a giant, magical library where every book ever written is stored, but the books are written in a secret code that only the library's robots can read. This library is a "Large Language Model" (AI), and its job is to understand human language, from writing poems to answering questions. But here's the tricky part: these robots don't just "know" things like we do; they have a hidden internal map made of thousands of layers of mathematical connections. Scientists are trying to figure out how these robots "think" about dangerous topics, like self-harm. Think of the robot's brain like a multi-story skyscraper. When a message enters the building, it travels up through the floors. At the bottom, the message is just raw noise. As it climbs, the robot starts to make sense of it. The big question researchers are asking is: At exactly which floor does the robot finally realize, "Oh, this is a cry for help"? And once it realizes that, does it store that thought in a simple, straight line, or is the idea hidden in a complex, twisted maze? Understanding this is crucial because if we want to build AI that can safely spot people in crisis and help them, we need to know exactly where and how that "danger signal" lights up inside the machine.
In this study, the researchers acted like detectives peering into the inner workings of four different AI skyscrapers (specifically models named Qwen3-0.6B, Llama 3.2-1B, Llama 3.2-3B, and Gemma-3-4B). They wanted to see how these models represent posts about self-harm. They ran two main experiments. First, they built a simple "detector" (called a linear probe) at every single floor of the building to see if the AI could tell the difference between a post about self-harm and a normal post. They found that the AI doesn't really "get it" until the very top of the building. Specifically, the information about self-harm "crystallizes"—meaning it becomes clear and solid—in the final 3% to 7% of the layers. For a building with 34 floors, that means the AI only fully understands the danger in the top two or three floors. Before that, the signal is too fuzzy to catch reliably.
The second experiment was even more interesting. The researchers tried to find a single "direction" in the AI's math-space that points directly at self-harm, kind of like finding a compass needle that always points to "danger." They expected that the AI models that were best at spotting self-harm would also have the clearest, straightest compass needle. But they found something surprising: that wasn't true. The model called Gemma-3-4B was actually the best at spotting self-harm (getting the highest accuracy scores), yet its "danger compass" was the least straight and simple. It turns out that Gemma doesn't store the idea of self-harm in one single, easy-to-find line. Instead, it spreads the signal out across many different directions in its brain, making it harder to find with a simple compass but actually better at recognizing the real thing.
The team also discovered that the way the AI represents this danger changes as the message moves up the building. The direction of the "danger signal" at the bottom of the building is almost completely different from the direction at the top. It's as if the message starts as a red arrow at the bottom, but by the time it reaches the top, it has rotated and turned into a blue spiral. This means you can't just look at the early layers of an AI to understand its safety features; you have to look at the very end.
Finally, the researchers compared two different sets of data they used to train their detectors. One set was a bit messy, with posts that used metaphors or talked about self-harm in news stories, while the other set was very clear and direct. They found that the AI performed much better and more consistently on the clear, direct data. This suggests that the quality of the examples we give the AI matters a huge amount. If the training data is confusing or full of tricky metaphors, the AI's internal map becomes fuzzier.
In short, the paper suggests that for these AI models, the concept of self-harm is a late-stage realization that lives in the top few layers of the network. It also shows that the most accurate AI isn't necessarily the one with the simplest "danger signal," but rather the one that encodes the danger in a more complex, distributed way. This gives us a new way to think about how to build safer AI: instead of just hoping the robot knows better, we might be able to engineer specific parts of its brain to handle these critical moments, provided we understand exactly where and how those thoughts form.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.