← Latest papers
📊 statistics

LLM Evaluation as Tensor Completion: Low Rank Structure and Semiparametric Efficiency

This paper establishes a semiparametric inference framework for low-rank tensor completion in LLM evaluation, deriving efficiency bounds and introducing a score-whitening method to enable asymptotically normal, uncertainty-quantified estimation of model abilities from noisy, sparse pairwise comparisons.

Original authors: Jiachun Li, David Simchi-Levi, Will Wei Sun

Published 2026-09-04
📖 4 min read☕ Coffee break read

Original authors: Jiachun Li, David Simchi-Levi, Will Wei Sun

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the rapidly evolving world of artificial intelligence, a new kind of statistical challenge has emerged. As large language models become more capable, the question is no longer just about building them, but about knowing exactly how good they are. To answer this, platforms have turned to a method that mirrors how humans learn: asking people to compare two responses side by side and vote for the better one. This process generates a massive stream of data, but it is messy. Some models are compared thousands of times, while others are rarely seen. Some questions are asked constantly, while others are ignored. The result is a leaderboard that looks like a clear ranking but is actually built on a foundation of uneven, noisy, and incomplete information. Without a way to measure the uncertainty in these rankings, it is difficult to know if a model has truly improved or if the result is just a fluke of the data.

Researchers at MIT and Purdue University have tackled this problem by treating the evaluation of these AI models as a puzzle of missing pieces. They realized that a model's ability is not a single number, but a complex profile that changes depending on the task, the language, or the user. To make sense of this, they imagined the performance of every model across every possible scenario as a vast, multi-dimensional grid of hidden scores. Most of these scores are never directly observed because we cannot ask a human to compare every model against every other model in every situation. Instead, we only see the outcome of specific, random matchups. The researchers developed a new mathematical framework to fill in these missing scores, not just by guessing, but by calculating the most probable values while rigorously measuring how confident we can be in those guesses.

The core of their work addresses a specific difficulty that previous methods missed: the uneven nature of the data. In a perfect world, every comparison would be equally informative, but in reality, a matchup between two very similar models provides a lot of useful information, while a matchup between a top-tier model and a weak one tells us very little. Furthermore, the way data is collected is often biased toward popular models or trending topics. The researchers found that standard statistical tools fail in this environment because they assume the data is uniform. They discovered that the mathematical machinery used to estimate these scores becomes unstable when the information is lopsided, leading to unreliable rankings and misleading confidence intervals.

To solve this, the team introduced a technique they call "score whitening." Think of it as adjusting the volume on a radio so that every station, whether it is broadcasting clearly or with static, contributes equally to the final sound. By mathematically rescaling each comparison based on how much information it actually carries, they transformed the chaotic, uneven data into a balanced stream. This allowed them to construct a new type of estimator that can handle the real-world messiness of AI evaluation. They proved that this method allows for the creation of confidence intervals that are both accurate and efficient, meaning they are as tight as possible without sacrificing reliability.

Their findings show that by accounting for the low-rank structure of the data—meaning that model performance is driven by a few underlying factors like reasoning ability or coding skill rather than random noise—it is possible to borrow strength across different tasks and models. This borrowing of information is crucial when data is sparse. The researchers demonstrated that their approach works not only for simple questions, like "how much better is Model A than Model B?", but also for complex, non-linear questions, such as "what is the probability that Model A will win in a specific category?".

In experiments using both synthetic data and real-world data from the popular Arena platform, the new method proved its worth. When compared to existing approaches that treat each task separately or ignore the uneven sampling, the new estimator produced rankings with significantly narrower margins of error. It successfully calibrated its confidence levels, meaning that when it claimed a 95% chance of being correct, it was correct about 95% of the time. The study confirms that while the data from human preferences is inherently noisy and biased, it is possible to extract a clear, statistically sound picture of model performance. This provides a principled way for developers and users to understand not just who is winning, but how certain we can be about the victory, turning a noisy collection of votes into a reliable measure of artificial intelligence capability.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →