A systematic review of quality control efforts for histopathology annotation
This systematic review of 16 eligible studies found that while 62.5% of research on histopathology image annotation incorporates quality control efforts, the field remains fragmented with implementations varying between quality control measures, features, or a combination of both.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
In the world of medical diagnosis, a pathologist is a detective who examines tiny slices of tissue under a microscope to find signs of disease. To teach computers to do this work, scientists must first show them thousands of these tissue images, carefully drawing lines around the important parts to teach the machine what to look for. This process of drawing the lines is called annotation. However, just as a human can make a mistake when labeling a map, a computer or a human annotator can misidentify a cell or miss a tumor. If the computer learns from these mistakes, its future diagnoses will be flawed. To prevent this, researchers use specific checks and tools to ensure the labels are correct. These checks, known as quality control measures, involve humans reviewing the work, while the tools, called quality control features, are the buttons and menus that allow a reviewer to accept a label, reject it, or ask for it to be fixed. Without these safeguards, the digital pathologists of the future might be trained on faulty information.
A recent systematic review by Olivier Rukundo, a researcher at the University of Limerick and the Medical University of Vienna, takes a close look at how well these safety nets are currently being used in the specific field of histopathology, which is the study of tissue samples. The study did not create new software or run new experiments; instead, it acted as a comprehensive audit of existing research. The author searched five major scientific databases, casting a wide net for any study that discussed how to check the quality of tissue image labels. After finding one hundred and twenty-one potential records, the researcher carefully filtered them down, removing duplicates, conference papers that lacked detail, and studies that did not actually describe a method for checking quality. This rigorous process left a final group of sixteen studies that met the strict criteria for inclusion.
The review sorted these sixteen studies into three distinct groups based on what they offered. Some studies focused primarily on the human side of quality control, proposing methods for experts to review and correct annotations. Others focused on the software side, introducing digital features that help users decide whether a label is good or bad. A third group combined both approaches, offering a system where human review and digital tools work together. The findings revealed a clear split in the field. Overall, 62.5% of the studies implemented quality control efforts for histopathology image annotation, whereas 37.5% did not. This suggests that while a majority of researchers are aware of the need for these safeguards, a significant portion of the work in this area still proceeds without a formal system to catch errors.
The studies that did implement these checks showed a variety of approaches. Some relied on human collaboration, where pathologists would review each other's work or where experienced senior doctors would supervise the annotations made by junior fellows. Others introduced software tools that acted as a second pair of eyes. For instance, some programs used artificial intelligence to suggest where a structure should be drawn, allowing a human to simply accept the suggestion or correct it if it was wrong. Other tools were designed to spot inconsistencies, flagging areas where different people labeled the same image differently so that a final decision could be made. In the most advanced examples, the software and the human worked in a loop: the computer made a suggestion, the human reviewed it, and the system recorded whether the human accepted or rejected the change, creating a continuous cycle of improvement.
Despite the progress shown in the majority of the reviewed studies, the review highlights that the field is not yet uniform. The fact that over a third of the included studies did not implement these quality control efforts indicates that many projects may still be building their datasets without a robust method to verify accuracy. The review does not claim that the problem is solved or that a single perfect method exists. Instead, it maps the current landscape, showing that while many researchers are actively developing ways to ensure their data is clean, there is still a gap in adoption. The work serves as a reminder that for computers to become reliable partners in medical diagnosis, the human effort to label the training data must be supported by consistent, rigorous checks, whether those checks come from a senior pathologist or a digital tool designed to spot a mistake.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.