VisioMath: Benchmarking Figure-based Mathematical Reasoning in LMMs
This paper introduces VisioMath, a benchmark of 1,800 K-12 mathematics problems featuring visually similar diagrams, to reveal that current Large Multimodal Models struggle with fine-grained comparative reasoning due to image-text misalignment and to demonstrate that alignment-oriented strategies can significantly improve performance.