ScintiGround-38K: A Grounded Visual Question Answering Benchmark for Paired Anterior–Posterior Bone Scintigraphy
The paper introduces ScintiGround-38K, a large-scale, provenance-preserving benchmark for grounded visual question answering on paired anterior-posterior bone scintigraphy images, which reveals that while models achieve high raw answer accuracy, they struggle significantly with minority-class detection and strict evidence-based localization.