← Latest papers
🔬 pathology

Understanding Human AI Discrepancy in Breast Cancer TIL Assessment: A Multi-Rater and Perceptual Bias Study

This study reveals that while inter-pathologist agreement in breast cancer TIL assessment is high, there is substantial discrepancy between pathologists and AI models, suggesting that AI integration requires further refinement to match human reliability, particularly when using multi-category classifications rather than dichotomized thresholds.

Original authors: Capar, A., Aloglu, I., Aker, F., Ertano, M., Mese, Y. E., Ungor, A., Yildiz, B. E.

Published 2026-06-04
📖 5 min read🧠 Deep dive

Original authors: Capar, A., Aloglu, I., Aker, F., Ertano, M., Mese, Y. E., Ungor, A., Yildiz, B. E.

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

The Big Picture: Counting Tiny Invaders

Imagine breast cancer cells as a fortress. Inside the walls of this fortress, there are tiny "invaders" called lymphocytes (a type of immune cell). Doctors call these Tumor-Infiltrating Lymphocytes (TILs).

Think of TILs as the "good guys" fighting the cancer. The more of them you see, the better the patient's chances of surviving and responding to treatment. To figure this out, pathologists (specialist doctors who look at microscope slides) have to estimate what percentage of the space around the cancer cells is filled with these tiny invaders.

The Problem: Humans Are Bad at Estimating Crowds

The study started with a known problem: Humans are terrible at guessing numbers when things are scattered.

Imagine you are looking at a jar filled with jellybeans. If they are all clumped in one corner, you can guess the number easily. But if the jellybeans are spread out evenly across the whole jar, your brain struggles. You might guess 20% when it's actually 40%, or vice versa. This is exactly what happens when doctors look at cancer slides. Even with strict rules, two different doctors looking at the same slide often give different answers because their brains process "scattered dots" differently.

The Experiment: Humans vs. Robots

The researchers wanted to see if Artificial Intelligence (AI) could do a better job than humans. They set up a "taste test" with three human pathologists and two different AI models (let's call them Robot A and Robot B).

They gave all of them the exact same microscope slides of breast cancer and asked them to count the percentage of "invaders" (TILs).

Here is what they found:

1. The Humans Agreed with Each Other

When the three human doctors looked at the slides, they mostly agreed with each other. Their scores were very similar.

  • Analogy: It's like three experienced chefs tasting a soup. They might not use the exact same spoon, but they all agree on whether it's "a little salty" or "very salty."

2. The Robots Disagreed with the Humans

When the researchers compared the doctors to the robots, the agreement dropped significantly. The robots and the humans were often giving very different numbers.

  • The Twist: The robots weren't just "randomly" wrong. They had a specific habit: They consistently underestimated the number of invaders.
  • Analogy: Imagine the doctors say, "This jar is 40% full of jellybeans." The robots say, "No, it's only 25% full." The robots were looking at the same jar but seeing fewer beans than the humans did.

3. The Robots Were Good at "Ranking," Bad at "Counting"

Even though the robots gave the wrong numbers, they were actually pretty good at understanding the order.

  • Analogy: If you have three jars of jellybeans (Small, Medium, Large), the robots could correctly tell you which one was the smallest and which was the biggest. However, when asked to say exactly how many beans were in each, they were off by a lot.
  • The Data: The study showed that while the robots and humans didn't agree on the exact percentage, they did agree on which cases had more lymphocytes and which had fewer.

4. Why the Difference? (The "Crowded Room" Effect)

The researchers wanted to know why the robots and humans disagreed. Was the robot broken? Or was the human brain just flawed?

To find out, they did a special side experiment. They showed the doctors simple black-and-white pictures with white dots scattered on a black background (mimicking the lymphocytes). The doctors had to guess what percentage of the picture was white.

  • The Result: Even with these simple dots, the doctors were bad at guessing the exact percentage. Their guesses varied wildly from the true number.
  • The Conclusion: The disagreement isn't just because the robots are bad. It's because human brains have a hard time visually estimating scattered dots. The robots are actually doing a very precise "pixel count," while humans are using their "visual guesswork," which is prone to error.

The Takeaway

The study concludes that:

  1. Humans are consistent with each other: Doctors generally agree on the "vibe" of the slide.
  2. Robots are precise but "blind" to human perception: The robots count every single dot perfectly, but they end up with a different total than the humans because humans perceive scattered dots differently.
  3. The robots are currently underestimating: The AI models tend to say there are fewer immune cells than the doctors think there are.

The Bottom Line:
Right now, AI is like a super-accurate calculator that counts every grain of sand on a beach, while the human doctor is a person looking at the beach and guessing the volume. They are both looking at the same beach, but they are coming up with different numbers. Before we can trust the AI to replace the doctor, the AI needs to be "calibrated" to understand how human eyes see these scattered cells, or the doctors need to learn to trust the robot's precise counting over their own visual guesses.

Note: The paper emphasizes that this study was a test of how well the AI matches human doctors, not a test of whether the AI is ready to be used in hospitals yet. More work is needed to fix the "underestimation" issue before it can be used for real patient care.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →