← Latest papers
🤖 AI

SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context

This paper introduces SciFigQual-Bench, a novel full-text contextual benchmark for scientific figure quality assessment that evaluates images across five dimensions using expert-annotated data from top computer science conferences, alongside the SFQ-Agent framework which achieves state-of-the-art automated scoring performance.

Original authors: Zihan Deng, Chuanzhi Xu, Huiqi Liang, Haoyang Li, Xiaozhen Zhong, Lequan Yu

Published 2026-07-30
📖 3 min read☕ Coffee break read

Original authors: Zihan Deng, Chuanzhi Xu, Huiqi Liang, Haoyang Li, Xiaozhen Zhong, Lequan Yu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a mystery, but the clues are scattered across three different rooms. In one room, you have a photograph. In another, a handwritten note describing the photo. In the third, a diary entry where someone mentions the photo while telling a story. To solve the case, you can't just look at the photo; you have to check if the note matches the picture, and if the diary story makes sense with what's actually in the photo. This is exactly the challenge scientists face when reviewing academic papers. For a long time, computers were great at looking at photos and saying, "This is blurry," or "The colors are nice," but they were terrible at reading the story around the photo. They couldn't tell if a graph was lying about its data or if the caption was describing the wrong chart. This matters because in science, a picture isn't just art; it's proof. If the proof doesn't match the story, the whole experiment might be a bust.

Enter SciFigQual-Bench, a new tool designed to teach computers how to be better scientific detectives. The researchers behind this project realized that existing tools were like a judge who only looks at a painting without reading the title or the artist's statement. They built a massive database of over 7,600 real scientific figures from top computer science conferences, but with a twist: every single image is glued to its caption and the specific paragraphs in the paper that talk about it. They then asked human experts to grade these images on five things: how clear they are, how well they are laid out, if the caption matches the image, if the text in the paper matches the image, and if the image is trying to trick the reader.

The paper's main finding is that when you force a computer to look at the image, the caption, and the text separately before combining them, it becomes a much better judge. They created a system called SFQ-Agent that acts like a team of specialists: one looks at the pixels, one reads the text, and a third one compares the notes. When they tested this on 1,200 images, SFQ-Agent made fewer mistakes than any other method, getting the score right within one point of the human experts 93.4% of the time. The paper evaluates standard "all-in-one" models as baselines and finds that these single-pass approaches often get confused, mixing up what they see with what they read. Instead, the paper suggests that breaking the job into steps—checking the visual evidence first, then the text, then fusing them—is the secret to getting it right. The results are measured and specific: the best system had an average error of just 0.418 points on a 1-to-10 scale, proving that for scientific figures, context is king.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →