← Latest papers
💬 NLP

SurGE: A Benchmark and Evaluation Framework for Scientific Survey Generation

This paper introduces SurGE, a comprehensive benchmark and automated evaluation framework designed to address the lack of standardized metrics for scientific survey generation by providing a large-scale dataset and assessing model performance across comprehensiveness, citation accuracy, structural organization, and content quality.

Original authors: Weihang Su, Anzhe Xie, Qingyao Ai, Jianming Long, Xuanyi Chen, Jiaxin Mao, Ziyi Ye, Yiqun Liu

Published 2026-05-06
📖 4 min read☕ Coffee break read

Original authors: Weihang Su, Anzhe Xie, Qingyao Ai, Jianming Long, Xuanyi Chen, Jiaxin Mao, Ziyi Ye, Yiqun Liu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the world of science as a massive, ever-expanding library. Every day, thousands of new books (research papers) are added to the shelves. In the past, a human librarian would have to read these new books, figure out how they fit together, and write a "guidebook" (a survey paper) to help others understand the whole topic. But with the library growing so fast, no single human can keep up.

Enter SurGE. Think of SurGE not as a robot librarian, but as a giant, super-strict referee and a massive practice arena for new AI librarians.

Here is what the paper does, broken down into simple concepts:

1. The Problem: The "Wild West" of AI Librarians

Recently, smart AI assistants (called "agents") have started trying to write these guidebooks automatically. They are getting better at sounding smooth and organized. However, there was a major problem: How do we know if they are actually doing a good job?

  • The Old Way: Researchers would ask humans to read the AI's work and give it a grade. This is slow, expensive, and hard to repeat.
  • The Other Way: Some researchers built their own custom tests just for their specific AI, which is like a coach only testing their own players with rules they invented. This isn't fair for comparing different AIs.

2. The Solution: SurGE (The Arena and the Rulebook)

The authors built SurGE, which is two things in one:

  • The Arena (The Dataset): They created a massive "practice field" containing over 1 million scientific papers. They also gathered 205 "Gold Standard" guidebooks written by real human experts. These are the perfect examples that the AI should try to match.
  • The Rulebook (The Evaluation Framework): They invented a new way to grade the AI that doesn't rely on humans reading every single word. Instead, they use a mix of:
    • Objective Checks: Did the AI find the right books? (Like checking a shopping list).
    • AI Judges: They use other smart AIs to grade the writing style, logic, and structure, but they first proved these AI judges agree with human experts.

3. How They Grade the AI (The Four Pillars)

When an AI tries to write a survey, SurGE checks it on four specific things, using a creative analogy:

  • Comprehensiveness (The "Did you miss anything?" check): Did the AI find the most important books in the library? If the human expert cited 100 key papers, did the AI find them?
  • Citation Accuracy (The "Truthfulness" check): This is the big one. If the AI says, "Book X says this," does Book X actually say that? The system checks if the AI is making things up (hallucinating) or if the book actually supports the claim.
  • Structure (The "Blueprint" check): Does the guidebook have a logical flow? Is it organized with clear chapters and sub-chapters, or is it a messy pile of paragraphs?
  • Content Quality (The "Readability" check): Is the writing smooth, clear, and free of errors?

4. What They Found (The Scoreboard)

The authors tested several different AI "librarians" using SurGE. Here is what the scoreboard showed:

  • The "Smooth Talker" Trap: Some AIs sounded very fluent and well-written (like a smooth-talking salesperson), but they were terrible at finding the right facts. They got high scores on "Content Quality" but failed miserably on "Citation Accuracy." They were lying smoothly.
  • The Power of Planning: The AIs that did the best were the ones that didn't just write from start to finish. Instead, they planned first. They made an outline, broke the topic into small pieces, and then wrote. This is like an architect drawing a blueprint before building a house. These "Agent" systems built better structures than the basic ones.
  • The Bottleneck: Even the best AIs struggled with finding the right papers. The system showed that the AI's ability to write is actually better than its ability to search the library. The main problem isn't that the AI can't write; it's that it can't find the right evidence to back up its writing.

Summary

SurGE is a new tool that lets researchers fairly compare different AI systems that try to write scientific summaries. It proved that while AI is getting good at sounding smart and organizing ideas, it still struggles to be a reliable researcher that finds the right facts and cites them correctly. The paper suggests that future AI needs to get better at "searching" and "fact-checking" before it can truly replace human experts in writing these guidebooks.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →