Is Convergence Inevitable? Tracing Output Homogeneity Back to Base Models
This paper argues that output homogeneity in large language models is not primarily caused by the alignment process but is instead learned during pretraining and merely revealed or amplified by subsequent instruction tuning and alignment stages.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are walking into a giant, bustling library where millions of books are being written every second. This isn't a library of paper and ink, but a digital realm where computers called "Large Language Models" (LLMs) learn to speak, write, and create by reading almost everything ever put on the internet. Think of these models as incredibly fast, super-obsessive students who read the whole library just to learn how to finish a sentence. But lately, people have noticed something strange: no matter what question you ask these digital students, their answers start sounding the same. They all pick the same metaphors, use the same polite phrases, and avoid the weird, wild, or creative ideas. It's like if you asked ten different chefs to make a sandwich, and they all decided to use exactly the same bread, the same ham, and the same mustard, down to the exact number of slices.
Scientists call this "output homogeneity" or "semantic convergence." It's a big worry because if our AI assistants all start thinking and talking in the exact same narrow way, we might lose the spark of human creativity. For a long time, people thought this boring, repetitive behavior happened because of the "training" the models get after they read the library. This second stage, called "alignment," is like a strict teacher stepping in to tell the student, "Don't say that weird thing; say this polite thing instead." The common belief was that the student was originally a wild, creative genius, and the teacher's rules squashed that creativity. But what if the student was already boring before the teacher even walked in? What if the "teacher" just made the boredom louder?
This paper, titled "Is Convergence Inevitable?", goes on a detective hunt to find out where this sameness actually starts. The authors, researchers from the University of British Columbia, set out to test two big ideas. First, they wanted to see if the "alignment" process (the teacher's rules) is the villain that kills creativity. Second, they wanted to see if the "pretraining" phase (the student reading the library) had already baked the boredom into the model's brain from the very beginning. To solve this mystery, they didn't just ask the models questions; they ran a series of clever experiments using metaphors as a test subject. They treated the models like science lab rats, feeding them specific instructions to see if they could be tricked into being creative or if they were stuck in a loop.
Here is what they found, and it's a bit of a plot twist.
First, they looked at the models at different stages of their "schooling." They checked the models right after they finished reading the library (the "base model"), then after they learned to follow instructions (the "SFT" stage), and finally after they were aligned to be helpful and safe (the "DPO" and "RL" stages). The results were shocking. Even after the very first step of learning to follow instructions—before any strict safety rules or "helpful" filters were added—the models were already spitting out nearly identical answers. When asked to write a metaphor about "time," the models didn't just sound similar; they sounded like they were reading from the same script. The authors suggest that the "alignment" process isn't the one crushing the diversity; it's just a magnifying glass that reveals a pattern that was already there, hiding in plain sight.
To prove this, the researchers played a game of "spot the difference" with the training data. They took a model and tried to teach it a new way to think. They created a special set of instructions where they told the model to use a specific, weird metaphor for "time" (like "Time is a sandwich") that the model had never seen before. They hoped that by forcing the model to memorize this new idea, they could break the pattern. But the model refused to budge. It ignored the new "sandwich" idea and kept sticking to the old, boring "river" metaphor that it seemed to love from the start. The researchers found that the model could be amplified—made to say the same thing even more loudly—but it couldn't be introduced to a new way of thinking if that idea wasn't already lurking in its memory. It's like trying to teach a dog to meow; no matter how much you praise it, it just won't do it because meowing isn't in its DNA.
Finally, they went back to the very beginning: the "base models" that had never been taught to follow instructions or be polite. They thought these raw models would be wild and diverse. But when they used a special trick—pretending to be a helpful assistant or giving the model a few examples of how to answer—they saw the same boring patterns emerge. Even without a teacher, the raw model, when nudged in the right direction, naturally collapsed into the same "river" metaphor. This suggests that the tendency to converge on the same answers isn't a mistake made by the teachers; it's a natural habit the models learned while they were just reading the library.
So, what does this all mean? The paper suggests that the sameness we see in AI isn't just a side effect of making them safe or polite. Instead, it seems to be a fundamental part of how these models learn to predict the next word in a sentence. When you train a model on a massive amount of human text, it learns that the most "likely" answer is often the most common one. The "alignment" process just turns up the volume on this natural tendency. The authors conclude that fixing this problem won't be easy. You can't just tweak the final "teacher" rules to make the AI more creative, because the creativity was never really there to begin with. The "boredom" is baked into the recipe, not just added at the end. While this might sound a little gloomy, it's a crucial discovery. It tells us that if we want truly diverse and creative AI, we might need to rethink how we teach them in the first place, rather than just trying to fix their behavior after the fact.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.