← Latest papers
🤖 AI

SciFig: Towards Automating Editable Figure Generation for Scientific Papers

The paper introduces SciFig, an end-to-end multi-agent framework that overcomes the trade-off between visual quality and editability by automatically generating high-quality, fully editable methodology figures from scientific text, alongside a new benchmark (SciFig-Bench) and evaluation protocol (SciFig-Eval) to validate its superior performance.

Original authors: Siyuan Huang, Yifan Zhou, Yutong Gao, Zi Yin, Juyang Bai, Xinxin Liu, Rama Chellappa, Chun Pong Lau, Cheng Peng, Sayan Nag, Shraman Pramanick

Published 2026-06-26
📖 4 min read☕ Coffee break read

Original authors: Siyuan Huang, Yifan Zhou, Yutong Gao, Zi Yin, Juyang Bai, Xinxin Liu, Rama Chellappa, Chun Pong Lau, Cheng Peng, Sayan Nag, Shraman Pramanick

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are writing a recipe for a complex new dish. You've written down all the steps in words, but you know that to really help a chef understand how the ingredients flow from the bowl to the stove, you need a picture.

In the world of scientific research, these "pictures" are called methodology figures. They are the diagrams that show how a new computer program or scientific method actually works.

The Problem:
Right now, making these diagrams is like trying to draw a masterpiece with a crayon while wearing oven mitts.

  • The "Crayon" Problem: If you use standard drawing tools (like PowerPoint), you can edit the picture easily, but it often looks stiff, flat, and a bit boring—like a stick-figure sketch.
  • The "Oven Mitt" Problem: If you use fancy AI image generators to make it look beautiful and realistic, you get a high-quality picture, but it's like a photograph. If you want to change just one ingredient or move one arrow, you can't just click and drag. You have to throw the whole picture away and start over.

The Solution: SciFig
The authors of this paper built a new system called SciFig. Think of SciFig as a team of four specialized chefs working together to build a custom, editable, and beautiful diagram from your written recipe.

Here is how the team works:

  1. The Planner (The Architect): First, this agent reads your text and figures out the "blueprint." It decides, "Okay, this part goes on the left, that part goes on the right, and these two need to be connected by a big arrow." It doesn't draw anything yet; it just makes the plan.
  2. The Layout Agent (The Builder): This agent takes the blueprint and builds the frame. It creates a structured skeleton (using a digital format called XML) that holds the pieces in place. Because it's a skeleton and not a finished painting, you can still move the walls around easily.
  3. The Component Agent (The Decorator): Now, this agent fills in the details. It adds the "fancy" parts—like realistic textures, colors, and icons—making the diagram look professional and rich. But, it attaches these decorations to the skeleton, so they stay editable.
  4. The Feedback Agent (The Critic): Finally, this agent looks at the result. It can listen to a human saying, "Move that box up," or it can use its own "eyes" (an AI vision model) to say, "Hey, these two arrows are crossing awkwardly; let's fix that." It then tweaks the diagram until it's perfect.

The Result:
The final product is a diagram that looks as good as a human expert drew it, but it remains fully editable. You can click on a box, change a label, or move an arrow without ruining the whole image. The system does this in about 10 minutes.

How They Tested It (SciFig-Bench & SciFig-Eval)
To prove their system works, the team didn't just guess. They:

  • Built a Library (SciFig-Bench): They gathered 435 real, high-quality diagrams from top computer science papers. These served as the "gold standard" answers.
  • Created a Grading Rubric (SciFig-Eval): Instead of just asking an AI, "Is this pretty?" they created a strict 4-point checklist:
    1. Completeness: Did it include all the important parts?
    2. Fidelity: Does it look like the original author's drawing?
    3. Quality: Is the text clear and the layout logical?
    4. Design: Are the colors and spacing pleasing to the eye?

The Verdict:
When they compared SciFig to other methods (including single AI image generators and other multi-step systems), SciFig won every category.

  • It produced diagrams that were more accurate and complete.
  • It looked better and was easier to read.
  • Most importantly, in a blind test with 60 human experts, SciFig was chosen as the best 81% of the time, beating even the original diagrams drawn by the paper authors themselves.

In Short:
SciFig is a tool that automates the hard work of drawing scientific diagrams. It combines the editability of a digital drawing tool with the visual beauty of an AI image generator, ensuring that scientists can communicate their ideas clearly without spending hours wrestling with drawing software.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →