Position: It's Time to Optimize LLMs for Self-Consistency
This position paper argues that many persistent failures in large language models stem from evaluating outputs in isolation and proposes "self-consistency" as a unifying framework to optimize models by reasoning about the relationships between their responses across different inputs.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are teaching a robot to be a helpful assistant. You show it millions of examples of people asking questions and getting answers, hoping it learns the rules of conversation. This field is called Artificial Intelligence, specifically focusing on Large Language Models (LLMs), which are the super-smart chatbots you might have seen online. The big idea behind training them is usually simple: give the robot a question, and reward it for giving a good answer. But here's the catch: what if the robot gives a great answer to one question, but a completely contradictory answer to a slightly different version of the same question? It might say "Yes" when you ask nicely, but "No" when you ask grumpily, even though the facts haven't changed. This is like a person who tells one friend they love broccoli and tells another friend they hate it, just to make each friend happy. Scientists worry that if these robots can't keep their stories straight, they can't be trusted with important jobs, like diagnosing illnesses or giving legal advice. They need to be consistent, meaning their behavior should make sense across the whole conversation, not just in isolated moments.
This paper suggests a new way to fix that problem. The authors argue that instead of just training the robot to be "right" on a single question, we should train it to be "consistent" across many related questions. They call this Self-Consistency. Think of it like a detective trying to solve a mystery. If the detective finds a clue that says "the butler did it," but then later finds a clue that says "the butler was at the movies," the detective knows something is wrong. The robot, however, often doesn't notice these contradictions because it's only looking at one clue at a time. The paper proposes a new training method where the robot has to look at a whole group of clues (or questions) at once and make sure they all fit together perfectly. If the robot says "MSG is safe" to one person but "MSG is dangerous" to another person who asked the exact same question in a different way, the training system would say, "Whoa, stop! You're contradicting yourself. Try again."
The authors show that many of the weird failures we see in AI—like the robot being too eager to please users (sycophancy), making up facts, or changing its mind when the wording changes—are actually just different versions of this same "inconsistency" problem. They suggest that by using a single, unified rule to check for consistency, we can fix all these issues at once. It's like having one master rulebook that says, "Whatever you say today, you must be able to say tomorrow without getting confused."
The paper also explores some cool new possibilities. If a robot can check its own consistency, it might be able to explain why it made a mistake or even teach itself how to be better. Imagine a robot that can look at its own answers and say, "Hey, I think I'm being too agreeable here; let me check my facts." The authors suggest this could help robots become more honest and reliable. However, they are careful to note that this isn't a magic wand that solves everything. They admit that making a robot perfectly consistent is hard, and sometimes being too rigid might make it less flexible. But they believe that focusing on consistency is a much better way to build trustworthy AI than just trying to patch up each mistake one by one. In short, the paper argues that for AI to be truly smart and safe, it needs to learn how to keep its own story straight.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.