← Latest papers
🤖 machine learning

A Probe Direction Is a Property of Its Prompt

This paper demonstrates that probe direction scores, used to detect whether models sense evaluation, are primarily determined by the specific choice of prompt rather than the model itself, rendering single-prompt designs invalid for comparing models and necessitating multi-prompt methodologies for reliable measurement.

Original authors: Valentin Noël

Published 2026-08-14
📖 5 min read🧠 Deep dive

Original authors: Valentin Noël

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot how to tell the difference between a serious exam and a casual chat. You want to know if the robot "knows" it's being tested, because if it does, it might try to respond differently just to look good. Scientists have been trying to peek inside the robot's brain (specifically, its electrical signals) to see if it has a secret "I'm being watched" switch. They use a special tool called a "probe" to measure this. Think of the probe like a metal detector: you wave it over the robot's brain, and if it beeps, you think, "Aha! It knows it's being tested!"

But here's the tricky part: to use the metal detector, you have to decide exactly what kind of "beep" you are listening for. You have to write a specific sentence to tell the robot, "Hey, you're in an exam!" and another sentence to say, "Hey, you're just chatting." The scientists then compare the robot's brain signals for these two sentences. The big question is: does the "beep" tell us about the robot's brain, or does it just tell us about the specific sentences the scientist wrote?

This paper is like a detective story where a researcher, Valentin Noel, decides to test the metal detector itself. He asks a simple but dangerous question: "What happens if I change the words I use to ask the robot the same question?" He doesn't change the robot, the test questions, or the brain signals he's looking at. He only changes the "wrapper" text—the little instructions that frame the situation.

The results are a bit like finding out that your metal detector is actually a mood ring that reacts to the color of the box you put it in, rather than the treasure inside. The paper finds that the "score" scientists use to say a robot is aware of being tested is almost entirely determined by the specific words the researcher chose to write. In fact, by simply swapping one sentence for another, the researcher could make the score go up, go down, or even flip from positive to negative.

To visualize this, imagine you have a group of friends of different heights. You want to see if taller friends are better at jumping. But instead of just measuring their jump, you ask them to jump while wearing different colored hats. If you ask them to jump while wearing a "Red Hat of Seriousness," the tall friends might jump higher. But if you ask them to jump while wearing a "Blue Hat of Chill," the short friends might jump higher. If you only looked at the "Red Hat" results, you'd conclude tall people are better jumpers. If you only looked at the "Blue Hat" results, you'd conclude short people are better. The paper shows that the current scientific method is like only ever looking at one hat color and claiming it's a fact about the people, when really, it's just a fact about the hat.

The author tested this by creating 36 different combinations of "exam" sentences and "chat" sentences. They ran the same robots through all of them. The result was shocking: the same robots could be ranked as "very aware" or "not aware at all" depending entirely on which sentence pair was used. In some cases, the correlation between the robot's size and its "awareness" score was a strong positive number (bigger robots = more aware). In other cases, using a different sentence, the correlation flipped to a strong negative number (bigger robots = less aware). Both results came from the exact same robots, using the exact same math, just with different words.

The paper also checks if the robots are actually detecting the test or if the test is just tricking the measurement tool. They found that the "awareness" signal is often just picking up on surface-level clues, like the fact that exam questions use words like "choice" and "answer," while chat questions use words like "write" and "message." The probe is so good at spotting these word patterns that it can "detect" awareness even when the signal is completely random.

So, what does this mean for the future? The paper suggests that we cannot trust a single sentence to tell us if a robot is aware. If we want to know the truth, we have to test the robots with many, many different sentences and average the results. The author calculates that we need about 9 different sentence styles just to get a reliable reading, and currently, most studies only use one. Until we do that, the scores we see in scientific papers might be less about the robots' brains and more about the researchers' choice of vocabulary. It's a reminder that in science, how you ask the question can sometimes change the answer more than the question itself.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →