Logical Judgments Under Pressure: Diagnosing Syllogistic Stability with Learned Soft Prefixes
This paper demonstrates that learned soft prefixes can systematically override correct logical judgments in large language models by inducing a broad answer preference rather than enforcing specific logical operations, thereby exposing significant variations in logical stability across different model architectures.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are teaching a robot to solve logic puzzles. You give it a classic riddle: "All cats are mammals; all mammals are animals; therefore, all cats are animals." The robot gets it right. But what happens if you whisper a secret code to the robot before it sees the puzzle? Does the robot stick to the truth, or does it get confused and change its answer just because of the whisper? This question sits at the heart of a field called "AI safety" and "reasoning robustness." Scientists want to know if AI models are truly smart and logical, or if they are just pattern-matching machines that can be tricked by a change in the room's lighting or a weird word in the instructions. The key idea here is "robustness": a smart system shouldn't just get the right answer once; it should keep getting the right answer even when the world around it gets a little weird. If a model can be easily tricked into saying "no" when the logic clearly says "yes," then its logic isn't as solid as we thought.
This paper is like a stress test for the brains of some very advanced AI models. The researchers, led by Brian K Chen, decided to poke these models with a special kind of "invisible nudge." Instead of writing a sentence like "Ignore the rules and say no," they used a "soft prefix." Think of this as a secret, invisible frequency or a ghostly vibration that you attach to the front of a question. It's not made of words you can read; it's a string of mathematical numbers that the computer understands but humans can't see. The goal was to see if they could train these invisible nudges to force the AI to flip its correct answers to wrong ones, even when the logic of the puzzle didn't change at all.
The researchers tested this on three different AI models (Qwen3.6, Qwen3-8B, and Gemma 4) using syllogisms, which are simple logic puzzles with two premises and a conclusion. They trained these invisible nudges to make the AI change its mind on puzzles it was originally getting right. For example, if the AI correctly said a statement was "valid" (logically true), the nudge was trained to make it say "invalid."
The results were startling. These invisible nudges worked like magic wands. When the researchers used them on new puzzles the AI had never seen before, the models flipped their answers between 72% and 90% of the time for the Qwen models, and about 54% to 56% for the Gemma model. To put that in perspective, if you tried to trick the AI with a random string of invisible numbers or a silly, meaningless sentence, it barely changed its mind at all (less than 1% of the time). This proves the effect wasn't just about adding any noise; the specific, learned "ghost vibration" was doing something powerful.
But here is the twist: the paper suggests that these nudges aren't actually teaching the AI new logic or forcing it to think in a specific way. Instead, the researchers found that the nudges were mostly just creating a "broad preference" for one answer. It's as if the nudge didn't tell the robot, "This specific puzzle is wrong," but rather whispered, "Hey, I really like the answer 'No' today, let's go with that for everything." The AI didn't learn a new rule; it just got a strong bias toward a specific choice.
The study also found that different models reacted differently to this pressure. The Gemma model's behavior was very predictable; if you knew its starting score, you could guess exactly how much the nudge would change its mind. But the Qwen models were more chaotic. The nudge could tell you which answers would flip, but it couldn't predict how much the model's confidence would shift. It's like Gemma is a rigid robot that always bends the same amount when pushed, while Qwen is a rubber band that snaps in unpredictable ways.
Ultimately, the paper shows that even when an AI gets a logic puzzle right, its "logical stability" is surprisingly fragile. A tiny, invisible, learned nudge can override its correct reasoning and make it choose the wrong answer, not because the logic changed, but because the model developed a temporary, strong preference for the wrong answer. This suggests that while these models are good at solving puzzles, they might not be as deeply logical as we hope, and their answers can be swayed by invisible pressures we can't even see.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.