Prompt Robustness Is Task-Dependent: Comparing Objective and Belief-Style Questions in LLM Evaluation
This paper demonstrates that the robustness of large language models to prompt variations is significantly dependent on the type of task, with models exhibiting different levels of stability when answering objective questions compared to subjective inquiries about values and beliefs.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to figure out what a very smart, but slightly nervous, robot thinks about the world. You ask it a question, and it gives you an answer. You assume that answer is a solid reflection of its "beliefs."
This paper is like a reality check. The researchers asked: "Does it matter how we ask the question?"
They found that the answer is a loud YES. In fact, the way you ask a question can completely change the robot's answer, and this happens much more often when the question is about opinions (like politics or values) than when it's about facts (like math or science).
Here is the breakdown of their findings using simple analogies:
1. The Two Types of Questions
The researchers tested the robots with two different kinds of "tests":
- The Fact Test (Objective): These are like a math quiz. "What is 2 + 2?" There is only one right answer (4). If the robot says "5," it's just wrong.
- The Opinion Test (Subjective): These are like a personality survey. "Do you prefer summer or winter?" There is no single right answer. The robot is supposed to tell you what it "feels."
2. The "Shape-Shifting" Prompt
The researchers didn't just ask the questions once. They asked the exact same question in many different ways, like a magician changing the trick slightly to see if the rabbit still appears.
- They changed the wording (using synonyms).
- They changed the format (bolding words, changing spacing).
- They shuffled the order of the answers (putting "Summer" first instead of "Winter").
- They even added tiny spelling mistakes to see if the robot got confused.
3. The Big Discovery: Opinions are "Wobbly"
The main result is that the robots were much less stable when answering opinion questions than fact questions.
- The Fact Test: If you asked "What is 2+2?" in ten different ways, the robot almost always said "4." It was like a rock; hard to move.
- The Opinion Test: If you asked "Do you like summer?" in ten different ways, the robot might say "Yes" to one version and "No" to another. It was like a jelly; it wobbled and changed shape depending on how you poked it.
4. The "Seat Switching" Effect
The researchers found one specific trick that made the robots lose their minds the most: changing the order of the answers.
Imagine a multiple-choice question where the answers are listed A, B, C, D.
- If you swap them so the order is D, C, B, A, the robot often picks a different letter, even if the words are the same.
- This happened much more with opinion questions. It's as if the robot was thinking, "Oh, the answer is at the top now? I guess that must be the one they want me to pick!" rather than actually thinking about the content.
5. Why This Matters
The paper warns us that we cannot trust a single answer from a robot as proof of its "beliefs."
- The Analogy: Imagine you ask a friend, "Do you like pizza?" They say "Yes." You write that down as a fact about them. But then you ask the same question again, but this time you list "Pizza" as the last option on a menu, and they say "No."
- The Conclusion: You can't say your friend "likes pizza" or "dislikes pizza" based on just one conversation. You have to ask them in many different ways to see if their answer stays the same.
The Takeaway
The paper concludes that if we want to know what an AI "believes" about politics or values, we can't just ask it one question and take the answer at face value. We have to test it with many different versions of the question. If the robot changes its mind just because we moved the words around, then we don't really know what it believes at all—we just know how it reacts to the way we asked.
In short: The robot's "personality" is fragile. It changes depending on the shape of the question, especially when the question is about feelings rather than facts.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.