Cross-cohort evaluation of thyroid ultrasound deep learning across different outcome definitions: a duplicate-controlled benchmark
This study demonstrates that frozen thyroid ultrasound deep learning models exhibit significantly varying performance across different cohorts and outcome definitions, highlighting the critical need for explicit reporting of label provenance and cautious interpretation of cross-cohort benchmarks.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Technical Summary: Cross-cohort Evaluation of Thyroid Ultrasound Deep Learning Across Different Outcome Definitions
Problem Statement
The performance of artificial intelligence (AI) models for thyroid ultrasound interpretation is highly dependent on cohort characteristics and the specific evaluation endpoint used. A critical challenge in the field is the interchangeability of different diagnostic endpoints: postoperative pathology (a definitive diagnostic outcome) versus radiologist-assigned risk categories like the American College of Radiology Thyroid Imaging Reporting and Data System (TI-RADS) (a management-oriented suspicion score). These endpoints represent distinct stages of the diagnostic pathway and are not interchangeable. Furthermore, deep learning models trained on one dataset often fail to generalize to independent cohorts due to shifts in population, equipment, acquisition protocols, and label definitions. This study addresses the need to benchmark thyroid ultrasound classifiers while explicitly separating evaluation against malignancy outcomes from agreement with TI-RADS-derived risk categories, all while controlling for image-level data leakage.
Methodology
The study employed a duplicate-controlled, image-level benchmarking pipeline using a fixed analysis environment (Python 3.10, PyTorch 2.5.1).
Datasets:
- Development Archive: A public Mendeley Data release containing 387 B-mode ultrasound images and 1,332 fine-needle aspiration cytology (FNA) blocks derived from 385 source cases. Labels were binary (Papillary Thyroid Carcinoma [PTC] vs. Benign) based on postoperative diagnosis. Patient-level identifiers were unavailable, preventing verified patient-level splits.
- External Cohorts: Two frozen models were evaluated without retraining or threshold adjustment on:
- TN3K: A subset of 1,228 images with dataset-provided benign/malignant labels.
- DDTI: 637 images evaluated against a binary endpoint derived from radiologist TI-RADS categories (grouping low-suspicion TI-RADS 2–3 vs. high-suspicion TI-RADS 4–5).
Data Preprocessing and Leakage Control:
- Duplicate Control: Images were grouped using MD5 hashes (exact duplicates) and perceptual hashes (near-identical images). These groups were kept within single partitions (train, validation, test) to prevent image-level data leakage.
- Preprocessing: Images were resized to 224×224 pixels. An exploratory experiment tested automatic texture-based Region of Interest (ROI) cropping to suppress peripheral interface artifacts.
Model Architecture and Training:
- Five backbone architectures were trained: ConvNeXt-Small, EfficientNet-B3, Swin-Tiny, ResNet-50, and DenseNet-121.
- Models were initialized with ImageNet weights and fine-tuned using class-weighted binary cross-entropy loss with AdamW optimization.
- An ensemble was constructed by averaging sigmoid probabilities from ConvNeXt-Small, EfficientNet-B3, Swin-Tiny, and ResNet-50.
- Calibration was assessed using temperature scaling, Brier score, and Expected Calibration Error (ECE).
Evaluation Strategy:
- Performance was measured using AUROC, AUPRC, accuracy, sensitivity, specificity, and F1 score.
- Confidence intervals were generated via non-parametric bootstrap resampling (2,000 replicates).
- The cytology and ultrasound modalities were analyzed as separate, non-paired image distributions; no within-patient comparative accuracy was estimated.
Key Contributions
- Endpoint Provenance Separation: The study explicitly distinguishes between evaluating a model's ability to predict malignancy (against pathology labels) and its agreement with radiological risk stratification (TI-RADS).
- Duplicate-Controlled Benchmarking: Implementation of exact and perceptual hashing to prevent image-level data leakage in the absence of patient-level identifiers, a common issue in public medical image archives.
- Cross-Cohort Frozen Evaluation: Evaluation of frozen models on two distinct external cohorts (TN3K and DDTI) without retraining, highlighting how performance varies based on acquisition and endpoint definition.
- Descriptive Modality Contrast: A non-paired comparison of ultrasound versus cytology performance within the development archive, clarifying that such comparisons describe image distributions rather than establishing clinical superiority.
Results
Internal Performance (Development Archive):
- Cytology: The ConvNeXt-Small model achieved an AUROC of 0.988 (95% CI 0.976–0.998), and the ensemble reached 0.986.
- Ultrasound: The EfficientNet-B3 model achieved an AUROC of 0.745 (95% CI 0.612–0.865), and the ensemble reached 0.733.
- ROI Cropping: Automatic ROI cropping improved the ultrasound ensemble's AUROC from 0.745 to 0.817.
- Calibration: Temperature scaling improved probability reliability (ECE reduced from 0.032 to 0.014) without altering rank-based metrics.
External Cohort Evaluation (Frozen Models):
- TN3K (Malignancy Labels): The ultrasound ensemble achieved an AUROC of 0.825, and ResNet-50 reached 0.888. These results indicate good discrimination against the dataset's distributed benign/malignant labels.
- DDTI (TI-RADS Endpoint): Against the TI-RADS-derived endpoint, the ensemble achieved an AUROC of 0.477 and ResNet-50 achieved 0.437. This indicates little to no agreement with the TI-RADS categorization.
- Interpretation: The authors note that the contrast between TN3K and DDTI results cannot be attributed solely to label definition, as cohort, equipment, acquisition, and disease spectrum also changed simultaneously.
Modality Comparison: A descriptive, non-paired bootstrap analysis showed a cytology-minus-ultrasound AUROC difference of 0.247 (95% CI 0.120–0.384). The authors emphasize this does not prove cytology is clinically superior, as the image sets were not patient-matched.
Significance and Claims
The paper claims that frozen thyroid ultrasound models perform differently across cohorts that differ in acquisition and outcome definition. The stark contrast between high performance on TN3K (malignancy labels) and near-chance performance on DDTI (TI-RADS labels) underscores that these endpoints are not interchangeable.
The study supports three primary implications for future research:
- Explicit Reporting: Label provenance (e.g., postoperative diagnosis vs. TI-RADS) must be explicitly reported.
- Cautious Interpretation: Cross-cohort performance contrasts should be interpreted cautiously, as they reflect a confluence of distribution shifts and endpoint definitions rather than a single causal factor.
- Patient-Level Validation: Future studies require patient-level external evaluation against a clinically appropriate reference standard to ensure transportability.
The authors modestly conclude that their findings support the need for verified patient-level splits and consistent reference standard descriptions, rather than claiming a definitive solution to the generalization problem in thyroid AI. The study does not validate malignancy diagnosis on DDTI, as the TI-RADS endpoint is not a malignancy reference standard.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.