← Latest papers
💻 computer science

A Multi Center Breast FNAC Whole-Slide Cytology Dataset for AI-Assisted Patch-Wise Classification Using C1 to C5 Reporting Categories

This paper introduces a multi-center breast FNAC whole-slide cytology dataset comprising 470 images from 321 Indian patients, featuring 7,398 expert-annotated patches with C1–C5 reporting labels to support AI-assisted patch-wise classification.

Original authors: Garima Jain, Abhijeet Patil, Surabhi Jain, Sanghamitra Pati, Amit Sethi, Sandeep Mathur, Pulkit Verma, Nishi Halduniya, Jatin Kashyap, Sharat Kumar, Simmi Kharb, Sunita Singh, Sucheta Devi Khuraijam
Published 2026-06-30
📖 4 min read☕ Coffee break read

Original authors: Garima Jain, Abhijeet Patil, Surabhi Jain, Sanghamitra Pati, Amit Sethi, Sandeep Mathur, Pulkit Verma, Nishi Halduniya, Jatin Kashyap, Sharat Kumar, Simmi Kharb, Sunita Singh, Sucheta Devi Khuraijam, Sushma Khuraijam, Ratan Konjengbam, Arvind Kumar, Deepali Tirkey, Saurav Banerjee, Shivani Kalhan, Rakesh Kumar Gupta, Ranjana Solanki, Deepika Hemranjani, Shashank Nath Singh, Uma Handa, Manveen Kaur, B. G. Malathi, Yogender P., Niraj Kumari, Shruti Gupta, Indu R. Nair, Vidya C., Basumitra Das, Sunil Kumar Komanapalli, Ravindra Karle, Tanaya Kulkarni, Vandana Raphael, Biswajit Dey, Vaishali Gaikwad, Nilam Adhav

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a computer how to be a master detective. To do this, you need to show it millions of clues. In the world of medicine, one of the most important clues for finding breast cancer early is a tiny sample of cells taken with a needle, known as a Fine Needle Aspiration Cytology (FNAC).

This paper is essentially the "training manual" and the "giant box of clues" that researchers have built to teach Artificial Intelligence (AI) how to read these cell samples.

Here is the story of how they did it, broken down into simple parts:

1. The Big Collection (The "Giant Library")

The researchers didn't just look at one hospital; they gathered samples from 13 different medical centers all across India. Think of this as collecting puzzle pieces from 13 different boxes to make one massive, complete picture.

  • The Players: They collected data from 321 patients.
  • The Images: They turned the physical glass slides containing the cells into 470 giant digital photos (called Whole Slide Images or WSIs).
  • The Variety: Just like photos can be taken with different filters, these slides were stained with two different colors (Papanicolaou and May–Grünwald–Giemsa) to make the cells pop out. They used the same high-tech camera for all of them to ensure the pictures were consistent.

2. The "Five-Star" Rating System

When a human doctor looks at these cells, they don't just say "Cancer" or "No Cancer." They use a specific 5-step scale (C1 to C5) to describe what they see:

  • C1: The sample is empty or broken (like trying to read a book with missing pages).
  • C2: Everything looks normal and healthy (Benign).
  • C3: Something looks a little weird or unusual (Atypical).
  • C4: It looks suspicious, like it might be cancer.
  • C5: It is definitely cancer (Malignant).

The goal of this dataset is to teach the AI to recognize these five specific categories, not just a simple "yes/no" answer.

3. The "Zoom-In" Process (Patch-Wise Classification)

This is the most clever part. A single digital slide is huge—like a high-resolution photo of a whole city. If you feed the whole city to the AI at once, it gets confused.

  • The Solution: The researchers acted like editors cutting out specific "clues" from the big photo. They manually found the interesting parts of the slides and cut them out into smaller squares called patches.
  • The Result: From the 470 big photos, they created 7,398 smaller "clue" images.
  • The Labeling: Expert doctors (pathologists) looked at each of these 7,398 tiny squares and gave them a label (C1 through C5). Then, a senior doctor double-checked their work to make sure the labels were perfect.

4. What's in the Box?

The researchers are sharing this entire collection for free. If you download it, you get:

  • The Giant Photos: The 470 original high-res slides.
  • The Clues: The 7,398 cut-out patches, ready for the AI to study.
  • The Map: A digital map (GeoJSON) that tells you exactly where each clue came from on the big photo.
  • The Instructions: A dictionary explaining what every piece of data means and a list of who contributed what.

5. Important Rules for Using the Clues

The paper gives a few very important warnings about how to use this data:

  • It's a Training Tool, Not a Final Verdict: The labels on these tiny squares are "expert-verified," but they are just for training the AI. In the real world, a doctor can't diagnose a patient based on just one tiny square; they need to look at the whole picture, the patient's history, and other tests.
  • The "Missing" Clues: The dataset is "enriched," meaning the doctors only cut out the parts that were interesting. They didn't cut out every single square of the slide. So, the AI learns from the "best" parts, not the whole messy background.
  • Imbalance: There are way more "Normal" (C2) and "Cancer" (C5) clues than "Weird" (C3) or "Broken" (C1) clues. The AI needs to be trained carefully so it doesn't get confused by this imbalance.

Summary

In short, this paper says: "We have built a massive, high-quality library of breast cell images from across India. We have cut them into small, labeled pieces, and we have given the map to the world so that anyone can build an AI to help doctors spot breast cancer earlier and more accurately."

The data is available online (on a site called Zenodo) for anyone to download and use for research.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →