← Latest papers
🤖 machine learning

CANDOR: Chance-Calibrated Discordance in Frozen Foundation Encoders

The paper introduces CANDOR, a chance-calibrated discordance metric that reveals frozen foundation encoders are not inherently blind to medical findings but suffer from weak geometric separation and poor head selection, as evidenced by their high nearest-neighbor error rates even when lightweight heads achieve strong performance.

Original authors: Soroosh Tayebi Arasteh, Sven Nebelung, Daniel Truhn

Published 2026-07-22
📖 7 min read🧠 Deep dive

Original authors: Soroosh Tayebi Arasteh, Sven Nebelung, Daniel Truhn

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot to recognize different types of birds. You show it millions of photos, and it learns to spot a "robin" or a "sparrow" by looking at the shapes and colors in the pictures. This is how modern "foundation models" work: they are giant, pre-trained brains that have seen almost everything, from cats to cars to clouds. Now, imagine you want to use this same robot to help doctors read X-rays. You don't want to retrain the whole robot; you just want to attach a tiny, simple "head" to it that says "yes, there is a broken bone" or "no, everything looks fine."

The big question is: Does the robot actually see the broken bone in its internal memory, or is it just guessing based on other clues? Usually, we check if the robot is good by seeing how often the tiny head gets the answer right. But what if the robot's internal map is actually messy? What if, deep inside its brain, a picture of a broken bone looks more like a picture of a healthy bone than it does like another picture of a broken bone? If that's true, the robot isn't "blind"—it's just confused. It's like a librarian who has organized books by color instead of by story; they can find a red book, but they can't tell you if it's a mystery or a romance just by looking at the shelf. This paper asks a crucial question: When these powerful robots are frozen in place, are they truly blind to medical details, or are they just weak at separating them?

The Problem with the Old Way

For a long time, scientists measured these frozen robots by how well a simple "head" could read their features. If the head got a high score, everyone assumed the robot was smart. But the authors of this paper, Soroosh Tayebi Arasteh and his team, realized this method was like judging a chef by how well they can serve a meal without checking if the food is actually cooked.

They found a hidden trap in how we measure "confusion." Imagine you are trying to find a friend in a crowded room. If your friend is wearing a red hat, and there are 100 people in red hats but only 1 person in a blue hat, and you ask the robot to find the person in the blue hat, the robot might just guess "blue" because it's rare. The old measurement tools were getting fooled by how common or rare a disease was. If a disease was rare, the tools would say the robot was "collapsed" (completely useless) just because the math was unbalanced, not because the robot was actually bad. It was like blaming a detective for failing to find a needle in a haystack, when the real problem was that the haystack was made of 99% needles and 1% hay.

Enter CANDOR: The Fair Judge

To fix this, the team invented a new tool called CANDOR (Chance-Calibrated Discordance). Think of CANDOR as a referee that makes sure the game is fair before the whistle blows.

Here is how it works: Instead of letting the robot look at a messy crowd, CANDOR sets up a perfectly balanced game. It takes a picture of a patient with a specific problem (like a collapsed lung, called a pneumothorax) and asks the robot to find its five closest "neighbors" in its memory. But here's the trick: CANDOR forces the robot to choose between two groups of neighbors that are exactly the same size. One group has patients with the problem, and the other group has patients without it.

If the robot is truly good, the patients with the problem should be closer to the "problem" group. If the robot is confused, the patients with the problem will be closer to the "no problem" group. Because the groups are the same size, the robot has a 50/50 chance of guessing right just by luck. This gives the scientists a perfect "chance level" to compare against.

The Shocking Discovery: Not Blind, Just Weak

When the team ran CANDOR on 22 different powerful robots across 20 different datasets (including chest X-rays, eye scans, and even bird photos), they found something surprising.

First, they proved that the robots are not blind. The old tools had made it look like the robots were completely useless for rare diseases, but CANDOR showed that the robots actually do have the information. They aren't staring at a blank wall; they are just standing in a foggy room.

However, the robots are very weak. Even the best robot in the study, called RAD-DINO, which is trained specifically on chest X-rays, still got confused about 18.4% of the time on cases of pneumothorax. This means that for nearly one out of every five patients with a collapsed lung, the robot's internal map placed that patient closer to a healthy person than to another person with a collapsed lung.

The numbers were even starker for other conditions. For glaucoma (an eye disease), the best robot performed at the level of pure chance (50%), meaning it was no better than flipping a coin. For bird species, the same robot was a genius, getting confused only 4.5% of the time. But for chest X-rays, it was confused 42.8% of the time. It's as if the robot is a master chef who can perfectly identify a rare spice but gets confused when asked to tell the difference between salt and sugar.

Why Does This Happen?

The team also asked: "What makes a robot get confused?" They tested many ideas. Is it because the robot is too small? No. Is it because it was trained on old data? No. Is it because the disease is rare? No.

They found one thing that strongly predicts confusion: Erasure Retention. Imagine you take a photo of a broken bone and then use a magic eraser to wipe out the broken part, leaving just the healthy bone. If you show this "erased" photo to the robot, does its internal map change?

  • If the robot is smart, its map should shift dramatically because the most important part is gone.
  • If the robot is weak, its map barely moves. It's like a person who is looking at a picture of a dog but is actually focusing on the background grass; if you erase the dog, they don't notice.

The team found that the robots that barely moved when the evidence was erased were the same ones that got confused the most. This suggests the robots aren't looking at the disease itself; they are looking at something else in the image, like the shape of the rib cage or the lighting, and missing the actual problem.

The Good News and the Bad News

The paper ends with a hopeful but realistic twist. Even though a single robot might get confused, the team found that if you have a panel of 11 different robots, there is almost always one of them that gets it right.

Think of it like a group of detectives. If one detective misses a clue, another might catch it. The team showed that while a single robot might miss 35.9% of the cases, a smart system that picks the best robot for each specific image could miss only 2.8%. The information is there; it's just scattered across different robots. The problem isn't that the data is missing; it's that we haven't built a system good enough to pick the right robot for the job.

The Takeaway

This paper doesn't say we should throw away our AI doctors. Instead, it gives us a new, fairer way to check if they are working. It tells us that these frozen robots are not blind, but they are fragile. They can be tricked easily, and they often miss the fine details that matter most. By using CANDOR, doctors and scientists can now spot exactly which diseases a robot is weak at before they even try to use it. It's a reminder that just because a robot is big and powerful doesn't mean it sees everything clearly; sometimes, it just needs a better map.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →