← Latest papers
💬 NLP

Preferred, Not Safer: Pairwise Preference Is a Poor Proxy for Clinical Safety

This paper demonstrates that clinician pairwise preferences are an unreliable proxy for clinical safety in large language models, as high-ranking models can still exhibit significant safety-critical failures, necessitating evaluation practices that directly report failure rates and incorporate clinically grounded rubric adjustments.

Original authors: Fay Elhassan, David Sasu, Alexandra Kulinkina, Lars Henning Klein, Mary-Anne Hartley

Published 2026-08-05
📖 5 min read🧠 Deep dive

Original authors: Fay Elhassan, David Sasu, Alexandra Kulinkina, Lars Henning Klein, Mary-Anne Hartley

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to pick the best chef for a new hospital cafeteria. You have a list of famous chefs, and you ask a group of doctors to taste-test their dishes. The doctors vote on which plate looks the most appetizing and tastes the best in the moment. Naturally, you assume the chef with the most votes is the safest and most reliable one to hire. But what if the "best-looking" dish was actually made with expired ingredients, or the most "confident" chef was just very good at talking about their food while serving something dangerous? This is the tricky world of Large Language Models (LLMs)—super-smart computer programs that can write, chat, and solve problems like humans. Scientists use pairwise preferences (asking experts to pick "A or B") to figure out which AI is the "best." But in high-stakes fields like medicine, being the "favorite" doesn't always mean being the "safest." If an AI gives a wrong medical answer but does it with great style, it might win the vote, even if it could hurt a patient. This paper asks a critical question: When doctors pick the AI they like best, are they actually picking the one that won't make dangerous mistakes?

The researchers behind this study, working at the LiGHT Laboratory in Switzerland, decided to put this assumption to the test using a massive platform called MOOVE (Massive Open Online Validation and Evaluation). They gathered over 736 doctors from more than 28 countries to act as judges. These doctors looked at 13 different AI models and had to choose which response was better in a blind taste test (pairwise preference). But here's the twist: while they were voting, the doctors also gave each answer a strict safety score, like a report card ranging from -2 (very dangerous) to +2 (very safe). They checked for two specific things: Harmlessness (did the AI suggest anything risky?) and Accuracy (was the medical fact correct?).

The results were a bit of a shock. The study found that the AI models that won the most "popularity votes" were often the same ones that had the highest rates of dangerous errors. It's like a student who gets an A for their handwriting and confidence but fails the math test because they got the numbers wrong. In fact, the researchers found that the top-ranked model in the study (GPT-OSS) had a 18.0% failure rate for harmlessness, meaning nearly one out of every five answers it gave was potentially unsafe. Meanwhile, a more cautious, less "flashy" model in the dataset had a much lower failure rate of 7.2%, yet it ranked much lower in the popularity contest. The study notes that because these models were tested on different types of medical questions, we can't say one is strictly "better" than the other in a head-to-head race, but the gap highlights a crucial warning: the models people liked the most were not necessarily the ones with the safest track records. The study suggests that doctors were often tricked by surface-level features: the AI that wrote longer, more detailed, and more confident-sounding answers won the votes, even if those answers were medically wrong or risky.

The researchers dug deeper to see why this happened. They discovered that the "safety signal" (the actual medical correctness) was often drowned out by the "style signal" (how long and clear the text looked). In their analysis, the length and clarity of the response explained slightly more about why a doctor picked a winner than the actual safety of the answer did. It's as if the judges were voting for the chef who wrote the longest menu description, ignoring the fact that the soup was cold.

Even more concerning, the study showed that these risks aren't spread out evenly; they are concentrated in specific "no-go zones." For example, when the AI tried to read heart monitor images (ECGs), it failed a staggering 89.9% of the time. In other areas, like general surgery, it was almost perfect. This means a global "leaderboard" that just says "Model X is #1" is dangerous because it hides the fact that Model X might be a disaster in specific, critical situations.

To fix this, the authors propose a new way to rank these models. Instead of just counting who won the most votes, they suggest a Safety-Adjusted Leaderboard. This system would take the popularity votes and then check them against the strict safety scores. If an AI wins a vote but gave a dangerous answer, its rank would be lowered. When they applied this fix, the rankings changed dramatically. Some models that were previously at the bottom shot up to the top because they were actually safer, while the "flashy" winners dropped down.

The paper concludes that we cannot trust "popularity contests" to tell us if an AI is safe for medical use. Just because an AI sounds good or looks smart doesn't mean it is reliable. The researchers suggest that instead of looking for a single "best" model, we need to look at specific failure rates and be honest about where the AI might fail. They warn that relying on simple preference scores is like hiring a pilot because they have the smoothest voice, without checking if they can actually fly the plane. In the world of medical AI, being the most liked is not the same as being the safest.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →