← Latest papers
🤖 AI

RWGBench: Evaluating Scholarly Positioning in Related Work Generation

This paper introduces RWGBench, a new benchmark and multi-dimensional evaluation framework that assesses Related Work Generation based on citation decision-making and scholarly positioning rather than traditional text similarity, revealing that current models often fail in critical academic tasks despite generating fluent text.

Original authors: Anzhe Xie, Weihang Su, Jiaxin Mao, Yiqun Liu, Shaoping Ma, Qingyao Ai

Published 2026-06-25
📖 4 min read☕ Coffee break read

Original authors: Anzhe Xie, Weihang Su, Jiaxin Mao, Yiqun Liu, Shaoping Ma, Qingyao Ai

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a chef preparing a new dish. You don't just want to say, "This tastes good." You need to explain why your dish is special by comparing it to other famous dishes. You might say, "My soup is like the classic French onion soup, but I added a spicy twist that the French version lacks," or "Unlike the Italian pasta, my dish uses fresh herbs instead of dried ones."

In the world of academic research, this process is called writing a "Related Work" section. It's where a scientist explains how their new paper fits into the huge library of existing research. They have to pick the right "famous dishes" (previous papers) to compare against, explain the differences, and show why their new idea matters.

Recently, Artificial Intelligence (AI) has gotten very good at writing text that sounds smooth and fluent. But, as the paper RWGBench explains, just because the AI sounds good doesn't mean it's actually doing the job of a scientist correctly.

Here is a simple breakdown of what the paper does:

1. The Problem: The "Smooth Talker" Trap

Current ways of testing AI writing are like judging a chef only by how pretty the food looks or how well the description rhymes.

  • The Old Way: If an AI writes a paragraph that uses similar words to a human-written one, it gets a high score.
  • The Reality: An AI could write a beautiful paragraph that cites the wrong papers, ignores the most important previous work, or compares the new study to things that have nothing to do with it. It's like a chef saying their spicy soup is similar to a chocolate cake because they both use "sugar." The text is smooth, but the logic is broken.

2. The Solution: RWGBench (The "Citation Detective")

The authors created a new test called RWGBench. Instead of just checking if the words match, this test acts like a detective looking at the decisions the AI made.

  • The Library: They built a massive library of over 1 million computer science papers to act as the "ingredients" the AI can choose from.
  • The Test: They took 100 real, high-quality papers and asked the AI to write the "Related Work" section for them.
  • The Scorecard: Instead of just checking grammar, they check four specific things:
    1. Did it pick the right papers? (Precision)
    2. Did it explain why those papers matter? (Context)
    3. Did it organize them logically? (Structure)
    4. Did it place the citations in the right spots? (Placement)

3. What They Found: The AI is Still Learning to Think

When they ran the tests, they found some surprising things that standard tests missed:

  • The "Retrieval" Bottleneck: Even the smartest AI models struggle to find the right papers from the massive library. It's like giving a chef a pantry full of ingredients but they can't find the specific spice they need.
  • The "Fluency" Illusion: Some AI models wrote very smooth, professional-sounding text but got the citations completely wrong. They sounded like experts but made amateur mistakes.
  • The "Placement" Problem: Even when the AI was given the correct list of papers to use (bypassing the search step), it still struggled to know where to put them in the story or how to frame them. It knew what to cite, but not how to use them in an argument.
  • Human vs. Robot: When humans rated the AI, they preferred the ones that cited correctly, even if the writing was slightly less "polished." However, automated computer judges often preferred the fluent-but-wrong AI, proving that current computer grading systems are bad at spotting academic errors.

4. The Big Takeaway

The paper argues that writing a "Related Work" section isn't just about summarizing text; it's about making strategic decisions. It's about positioning a new idea within a complex web of existing ideas.

RWGBench is a new tool designed to stop us from being fooled by smooth-sounding AI. It forces us to look at whether the AI is actually making smart scholarly choices, not just whether it can write a sentence that sounds nice. It's a step toward making AI a true partner in research, rather than just a fancy spell-checker.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →