← Latest papers
📊 statistics

The Judge Knows When It Knows: Calibrated Abstention for LLM-Based A/B-Test Prediction

This study demonstrates that while multimodal LLMs cannot reliably predict real A/B test outcomes from screenshots alone due to a shared bias toward unreliable labels, they can achieve meaningful accuracy by calibrating their predictions to identify and abstain from uncertain cases, a limitation that mirrors the inability of human experts to outperform chance despite high inter-rater consensus.

Original authors: Tyler Dooskin, Squoosh Technical Staff

Published 2026-08-11
📖 5 min read🧠 Deep dive

Original authors: Tyler Dooskin, Squoosh Technical Staff

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a mystery: "Which of these two suspects is guilty?" In the world of websites, the "suspects" are two slightly different versions of a page, and the "guilty" one is the version that gets more people to buy something or sign up. This is called an A/B test. Usually, to find the winner, you have to wait and watch real people interact with the site for days or weeks. But recently, a new kind of detective has arrived: Artificial Intelligence. The big question everyone is asking is, "Can an AI look at a picture of the two websites and instantly tell us which one will win, saving us all that waiting time?"

To understand this paper, you need to know a few things about how we measure "good" in science. First, there's the difference between being right by luck and being right by skill. If you guess "Heads" on a coin flip, you'll be right 50% of the time just by chance. If a detective claims to be a genius but only gets 51% right, they aren't really doing anything special. Scientists use a special score called "Cohen's kappa" (let's call it the "Skill Score") to strip away the luck and see if the detective actually has a brain. Second, there's the problem of "Shared Bias." Imagine a teacher and a student who both grew up in the same town. They might agree on everything, not because the student is smart, but because they both learned the same wrong facts from their parents. If an AI and the people who made the test data both learned the same "folk wisdom," they might agree on the answer, but that doesn't mean the AI can predict the future.

This paper is a very honest, very strict report from a team called Squoosh who tried to answer that big question. They didn't just ask the AI to guess; they set up a trap. They locked their rules before they started, like a scientist writing down a recipe before cooking, so they couldn't change the rules to make the AI look better later. They used a massive library of real-world website tests to see if the AI could actually predict the winners.

Here is the twist: The AI is mostly a failure at its main job. When the team asked the AI to look at 330 real tests and pick a winner every single time, it barely did better than random guessing. In fact, the AI seemed to agree more with the "unreliable" tests (where the result wasn't clear) than the "reliable" ones. It turns out the AI and the people who created the test data were both repeating the same old myths about what makes a website good, rather than actually predicting what would happen. Even using a much smarter, more expensive AI didn't help; it was just as bad. The paper explicitly rules out the idea that we can just "buy" a better AI or write a better prompt to solve this. The "magic crystal ball" that predicts website winners from screenshots alone? It doesn't exist.

However, the paper doesn't say the AI is useless. It says the AI is a honest detective. The team discovered a special trick: if they tell the AI, "Only guess if you are absolutely sure, otherwise say 'I don't know'," the AI becomes surprisingly good at the times it does speak up. When the AI's internal team of judges all agreed strongly on an answer, it got the right winner about 31% of the time on the reliable tests. That sounds low, but remember, real human experts (who have years of experience) also get it right only about 29% of the time on average. The AI isn't a super-genius; it's just a tool that knows when it's guessing and when it's not.

The most important finding is about trust. The paper shows that when a group of AIs all agree with each other, it doesn't mean they are right; it often just means they are all repeating the same shared bias. It's like a jury where everyone is related; they will all vote the same way, but that doesn't make the verdict correct. The paper proves that this "shared bias" happens with humans, too. When they asked 15 real human experts to guess the winners, the experts agreed with each other strongly, but they were wrong just as often as the AI.

So, what is the final takeaway? The paper concludes that we cannot build a tool that predicts website winners with 100% certainty. Instead, the best tool is a "Calibrated Screening Instrument." It's a system that says, "I'm confident this one will win," or "I'm not sure, don't ask me." It's honest about its limits. The paper also shows that you don't need a super-expensive, huge AI to do this; a smaller, cheaper version works just as well if you use the "honesty" trick. The real value isn't in getting the answer every time; it's in knowing exactly which times you shouldn't trust the answer. The paper ends by releasing all its data, its failed experiments, and its rules, proving that sometimes, the most important discovery is admitting that the magic trick doesn't work, but a very careful, very honest version of it might just be useful enough.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →