The Asymmetric Harms of LLM Compression
This paper reveals that standard aggregate metrics fail to capture the asymmetric harms of LLM compression, which disproportionately degrades head knowledge retention, maintains high confidence in newly lost information, and masks opposing shifts in social biases across demographic subgroups.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Large language models are the powerful computer programs behind many of today's most advanced artificial intelligence tools, capable of writing stories, answering complex questions, and solving problems. To make these massive programs useful on everyday devices like smartphones or laptops, engineers often shrink them down, a process known as compression. This is done to save space and speed up how fast the computer thinks, much like packing a heavy suitcase more tightly for a trip. For years, the standard way to check if this shrinking process worked was to look at the overall score of the program. If the compressed version still answered most questions correctly and sounded natural, it was considered safe to use. However, this approach treats the model like a single, uniform entity, assuming that if the average performance looks good, everything inside is functioning well.
A team of researchers set out to look deeper than these average scores. They wanted to know if the process of shrinking these models hides specific, uneven problems that could be dangerous or unfair. They examined three different large language models and tested them against eleven different ways of compressing them. Instead of just checking if the models got the right answer, they investigated what happened to the specific facts the models knew, how sure the models were when they got things wrong, and whether the models treated different groups of people differently. Their work reveals that compression does not simply make a model slightly worse at everything; instead, it creates a strange, uneven landscape where some knowledge is lost while the model remains dangerously confident, and where hidden biases can shift in ways that average scores completely miss.
The researchers began by asking how compression affects the model's memory of different types of facts. In the real world, some facts are common knowledge, like the capital of a country, while others are rare, like the population of a small, obscure town. The team found that when a model is compressed, it does not lose these rare facts more often than the common ones, as one might expect. Instead, the opposite happens in a subtle way. While the model still answers questions about common facts correctly more often than questions about rare facts, it actually loses a larger proportion of its ability to recall those common facts. The rare facts, surprisingly, hold on to their original accuracy better than the common ones do. This means that the standard way of measuring performance, which looks at the total number of correct answers, hides the fact that the model's grasp on well-known information is weakening more than its grasp on obscure details.
Even more concerning is what happens when the model forgets something. The researchers discovered that when a compressed model loses a piece of knowledge it once had, it often does not realize it has forgotten. Instead of becoming unsure or admitting it does not know the answer, the model frequently remains very confident in its new, incorrect response. In many cases, even after compression has caused the model to fail on a question it previously answered correctly, the model still speaks with a high degree of certainty. This overconfidence is not consistent across all types of compression or all models, but it is a frequent and troubling pattern. It suggests that a compressed model might give a wrong answer with the same authority as a correct one, making it difficult for a user to tell when the machine is mistaken.
The study also looked at how compression affects fairness and bias, specifically how the models treat different groups of people based on gender or other characteristics. The researchers found that looking at an overall bias score can be misleading. A compressed model might show almost no change in its overall bias score, leading observers to believe it is still fair. However, when they looked closer at specific groups, they found that the model's preferences were shifting dramatically in opposite directions for different people. For example, a model might become significantly more biased against one group while simultaneously becoming less biased against another, and these two changes cancel each other out in the final average. This means that a model could appear perfectly balanced on paper while actually becoming much more harmful to specific individuals or communities.
These findings challenge the idea that a compressed model is simply a smaller, faster version of the original. The research shows that the process of shrinking these models creates asymmetric changes that standard tests fail to catch. The models do not just degrade evenly; they lose their grip on common knowledge while holding onto rare facts, they become confidently wrong, and they hide deep, uneven shifts in bias behind stable average scores. The authors conclude that before these compressed models are deployed in the real world, they must be tested with much more detailed, granular checks. Relying on average performance is no longer enough, because it leaves practitioners blind to the specific, uneven harms that compression introduces.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.