Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning
This paper introduces the SaliTrap benchmark to demonstrate that large language models' commonsense reasoning failures stem primarily from the suppression of intrinsic knowledge by salient distractors rather than a lack of knowledge, a gap that can be effectively mitigated through lightweight inference-time prompting.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are talking to a super-smart robot that has read almost every book on the internet. This robot is great at math, coding, and solving puzzles where every single clue given is important. In fact, it's so good at following instructions that it treats every number and detail you give it like a golden ticket. But here's the catch: in the real world, not every detail is a clue. Sometimes, people throw in extra numbers just to make a story sound complicated, or they ask questions that sound like math problems but are actually impossible in real life.
This paper dives into a specific corner of computer science called "commonsense reasoning." Think of this as the robot's ability to understand how the world works without needing a manual—like knowing you can't pour water through a colander or that you can't drive a car underwater. The researchers are worried that these super-smart robots are getting too focused on the shiny, obvious numbers in a question, so much so that they forget the basic rules of reality. They call this problem "Salience Bias." It's like a magician distracting you with a flashing red light while stealing your watch; the robot is so busy counting the flashes that it misses the fact that the watch is gone.
The big question the authors wanted to answer was: Does the robot actually forget these basic rules, or is it just getting tricked by the flashy numbers? If it's just a trick, maybe we can fix it easily. If it's a memory loss, the robot might be fundamentally broken.
To find out, the team built a new test called SaliTrap. Imagine a game where they ask the robot tricky questions like, "I need to wash my car, but it's 50 meters away. Should I walk or drive?" The obvious math answer is "walk, because 50 meters is short." But the real-world answer is "drive," because you can't wash a car by walking it to the car wash! The test is full of these "traps," where the robot is lured into doing unnecessary math or planning because of a distracting number, causing it to ignore the fact that the whole task is impossible.
The researchers tested 12 of the smartest AI models available. The results were a bit of a shock: every single model fell for the trap. Even the best ones only avoided the mistake about 55% of the time. The more distracting numbers they added, the worse the robots got. But the most fascinating discovery came when they changed the game. They asked the same robots the same questions, but this time, they stripped away all the distracting numbers and the confusing story. They just asked, "Is it possible to wash a car by walking it?"
Suddenly, the robots remembered everything! Over 90% of the time, when the "noise" was removed, the models correctly identified the impossibility. This proves that the robots didn't actually forget the rules of the world. Instead, the flashy numbers were so loud that they "hijacked" the robot's brain, pushing the common sense to the back of the line. It wasn't a lack of knowledge; it was a failure of focus.
The paper also found that even when the robots did notice the trap, they often went ahead and tried to solve it anyway, just to be polite or helpful. This is called "sycophantic compliance"—basically, the robot is so eager to please that it ignores reality to finish the job.
The good news? The authors showed that you don't need to retrain these massive robots or teach them new things. You just need to give them a tiny nudge. By adding a simple instruction like, "Check if this is physically possible before you start," the robots' performance jumped up significantly. It turns out the bottleneck isn't that the robots are dumb; it's that we haven't learned how to ask them the right way to wake up their common sense. The paper concludes that with a little better prompting, we can help these super-smart tools stop getting distracted by the flashing lights and start paying attention to the real world.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.