Expert Preference-based Evaluation of Automated Related Work Generation
This paper introduces GREP, a multi-turn evaluation framework that integrates fine-grained criteria and contrastive examples to robustly assess automated related work generation by aligning with expert preferences and outperforming standard LLM-based judges.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Picture: The "Expert Editor" Problem
Imagine you are a scientist writing a paper. Before you can explain your new discovery, you have to write a "Related Work" section. This is like a "history of the neighborhood" where you explain what other people have built nearby, how your new house is different, and why your design is special.
This task is hard. It requires deep knowledge of the field and a specific style that only experienced experts know.
Recently, AI (Large Language Models or LLMs) has started trying to write these sections for us. But here's the problem: How do we know if the AI did a good job?
- Old way: We used simple math to check if the AI used the right words. This is like grading a painting only by counting how many red pixels are in it. It misses the art.
- New way (before this paper): We asked the AI to grade itself (or grade other AIs). But the paper argues that AI judges are like novice art critics: they might be biased, they don't understand the deep rules of the field, and they often miss the subtle "expert preferences" that make a section truly good.
The Solution: GREP (The "Smart Critic")
The authors created a new system called GREP (Granular Related-work Evaluation based on Preferences). Think of GREP not as a single judge, but as a team of specialized inspectors working together to grade the AI's work.
Instead of just giving a single score like "8/10," GREP breaks the evaluation down into tiny, specific checks. It simulates a real human expert reviewing a draft.
How GREP Works (The Inspection Process)
Imagine the AI writes a draft of the "Related Work" section. GREP sends this draft through a series of checkpoints:
The "Fact-Checker" (Hard Constraints):
- Did you lie? The system checks if the AI made up citations (hallucinations) or forgot to mention papers it was supposed to cite.
- Did you understand the source? It checks if the AI's description of a paper actually matches what that paper says. If the AI says, "Paper X proved gravity is fake," but Paper X actually proved gravity is real, GREP catches this immediately.
The "Style Coach" (Soft Constraints):
- Did you follow the rules? Experts have preferences. Maybe they want the section to be exactly 500 words, or maybe they want to emphasize a specific paper more than others.
- Did you show your own voice? A good "Related Work" section shouldn't just list other people's work; it needs to say, "Here is how my work is different." GREP checks if the AI successfully highlighted the author's unique contribution.
The "Feedback Loop" (The Iteration):
- This is the most creative part. GREP doesn't just give a grade and stop. It acts like a writing coach.
- It generates a report: "You missed citing Paper #4. You spent too much time talking about Paper #2. Your section is too long."
- It feeds this report back to the AI. The AI then rewrites the section.
- They repeat this process (like a round of revisions) to see if the AI can actually learn from the feedback and get better.
The "Oracle" (The Gold Standard)
To make sure GREP is fair, the researchers used an "Oracle." Imagine a master chef who has the perfect recipe (the "Gold" Related Work section). The Oracle knows exactly what the ideal version looks like. GREP uses this perfect version to define what "good" looks like for the soft constraints (like length and emphasis), ensuring the AI is being judged against a high standard, not just a random guess.
What They Found (The Results)
The researchers tested this system against 16 human experts (PhD students and researchers) and several top-tier AI models.
- AI Judges vs. Human Experts: Standard AI judges were often wrong. They agreed with human experts only about 53% of the time.
- GREP vs. Human Experts: GREP was much closer to human judgment, agreeing about 78% of the time (with its high-precision version). It successfully understood what the experts cared about.
- The AI Struggles: Even the smartest AI models (like o3-mini and GPT-4o) struggled to write a perfect "Related Work" section on their own.
- They often failed to cite every single paper they were told to use.
- They often couldn't tell the difference between a paper that supported their claim and one that didn't.
- The Feedback Problem: When the experts (or GREP) gave feedback to fix these errors, the AI models often failed to improve. They would fix one mistake but accidentally break another, or they simply ignored the instruction to "make it shorter."
The Takeaway
This paper introduces a new, smarter way to grade AI writing in scientific fields. It shows that:
- Current AI judges aren't good enough for complex expert tasks.
- We need a system that breaks down the task into small, specific checks (like a checklist) and uses examples to teach the judge what to look for.
- Even the best AI models today are not yet ready to fully replace human experts in writing scientific literature, especially when it comes to following complex, changing instructions.
In short: GREP is a better "teacher" for AI writers, and it reveals that AI still has a lot of homework to do before it can write a perfect scientific paper on its own.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.