← Latest papers
📊 statistics

Agreement Between Fitzpatrick-17k Skin-Type Labels Is Metric- and Stratum-Dependent: A Chance-Corrected Reanalysis

This paper reanalyzes the Fitzpatrick-17k dataset using chance-corrected statistics to demonstrate that label agreement between annotation sources is highly dependent on the chosen metric and skin-type stratum, revealing significant discrepancies—particularly for darker skin tones—that necessitate transparent reporting of methodology and per-stratum confidence intervals in dermatology AI research.

Original authors: Niya Pennie

Published 2026-08-13
📖 4 min read☕ Coffee break read

Original authors: Niya Pennie

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you're trying to teach a robot to recognize different shades of human skin, like a digital art student learning to paint portraits. To do this, the robot needs a "teacher" to look at photos and say, "This is very pale," or "This is deep brown." In the world of dermatology and artificial intelligence, this teacher is often a human annotator who assigns a skin type based on the Fitzpatrick scale, a system that sorts skin into six categories, from Type I (very fair, always burns) to Type VI (deeply pigmented, rarely burns). But here's the tricky part: humans aren't perfect. One person might think a photo is "medium brown," while another says "dark brown." If the robot learns from these mixed-up instructions, it might get confused, especially when trying to help people with darker skin tones who are already underrepresented in medical care. The big question is: How much do these human teachers actually agree with each other, and does it matter if they disagree by just one step on the scale?

This paper is like a detective story where a researcher, Niya Pennie, goes back to a famous dataset called Fitzpatrick-17k to check the homework. This dataset was created by asking two different groups of crowd-workers (one from Scale AI and one from Centaur Labs) to label thousands of skin photos. The original creators of the dataset said, "Hey, 85% of the time, our two groups were pretty close, usually within one step of each other!" That sounded great. But Pennie decided to run a stricter, more mathematical test to see what was really going on. She didn't just ask, "Were they close?" She asked, "Did they pick the exact same label? And if they were off by one step, does that count as a win or a loss?"

Here is what the investigation found. First, the "close but not exact" story is true: if you count being off by just one skin type as "agreement," then yes, the two groups agreed 91% of the time. But if you demand they pick the exact same number, the agreement drops to less than half (47.9%). It's like two friends looking at a sunset and one saying "orange" and the other saying "red-orange." They are close, but they aren't saying the same thing.

The paper also discovered that this disagreement isn't spread out evenly. It's like a game where the rules change depending on the score. For the lightest skin types (Type I), the two groups agreed almost 88% of the time. But for the darker skin types (Types II through VI), the agreement plummeted, hovering between 34% and 55%. In other words, the human teachers were much more confused about darker skin than lighter skin. Furthermore, the two groups didn't just disagree randomly; they had a pattern. The Centaur Labs group tended to label photos as lighter than the Scale AI group did, especially for the darker skin tones.

Finally, the paper points out a serious "counting problem." For the darkest skin category (Type VI), there are very few photos of serious skin conditions (malignant images) in the dataset. Under one labeling system, there were 61 such images; under the other, only 45. Because these numbers are so small, the paper suggests that any study trying to measure how well an AI works on this specific group would be like trying to guess the average height of a room by measuring just three people. The results would be too wobbly and imprecise to trust.

So, the main takeaway isn't that the dataset is useless, but that we need to be much more careful about how we read it. The paper argues that researchers can't just say, "The AI is fair," without specifying exactly which labeling system they used, how they defined "agreement," and whether they have enough data to make a solid claim. If you change the rules of the game (the labeling source) or the size of the team (the number of images), the score changes. The paper concludes that to truly understand if AI is fair to everyone, we need to stop treating these labels as perfect facts and start reporting exactly how uncertain and variable they really are.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →