Comparative Performance of Frontier Large Language Models for Extracting High-Risk Pathologic Features from Unstructured Gastrointestinal Oncology Reports: A Systematic Benchmarking Study with Human and Traditional NLP Baselines
This systematic benchmarking study demonstrates that frontier large multimodal models, particularly Gemini 2.5 Pro and GPT-4o, significantly outperform both time-pressured human experts and traditional NLP baselines in extracting critical perineural invasion status from unstructured gastrointestinal pathology reports, while a novel failure mode taxonomy provides actionable guidance for model selection based on specific error signatures.