← Latest papers
💬 NLP

OARelatedWork: A Large-Scale Dataset of Related Work Sections with Full-texts from Open Access Sources

This paper introduces OARelatedWork, the first large-scale dataset enabling full-text related work generation from open-access sources, which reveals that while modern LLMs struggle with synthesizing massive contexts compared to abstracts, they can outperform human baselines in evidence-grounded factuality, necessitating new statement-level evaluation frameworks to replace inadequate reference-based metrics.

Original authors: Martin Docekal, Martin Fajcik, Pavel Smrz

Published 2026-06-02
📖 4 min read☕ Coffee break read

Original authors: Martin Docekal, Martin Fajcik, Pavel Smrz

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: Writing the "Chapter of Others"

Imagine you are writing a novel. Before you tell your own story, you have to write a chapter called "Related Work." In this chapter, you have to summarize what other authors have written before you, explain how their stories are similar to yours, and show where your story is different.

For a long time, computers trying to do this were like students who were only allowed to read the blurbs on the back of books (the abstracts). They had to guess the whole story based on a tiny summary. This paper introduces a new dataset called OARelatedWork that changes the rules: now, the computer gets to read the entire library (the full text of every book) before it tries to write its chapter.

The Problem: The "Blurb" vs. The "Book"

Previously, researchers built datasets where computers only saw the short summaries of cited papers.

  • The Old Way: It's like trying to write a book review based only on the back cover. You might get the main idea, but you miss the details, the nuances, and the specific evidence.
  • The New Way (OARelatedWork): This dataset gives the computer the full text of 94,000 papers and the full text of the millions of books they reference. It forces the AI to read the whole book, not just the blurb.

What They Found: The "Smart" vs. The "Human"

The researchers tested this new dataset with various AI models and compared them to human writers. They found some surprising things:

  1. More Information Can Be a Trap: When the AI was given the full text of books instead of just the blurbs, it actually made more mistakes about where to put the citations. It was like giving a student a 1,000-page textbook and asking them to find one specific fact; they got overwhelmed and sometimes pointed to the wrong page.
  2. Humans "Make Things Up" (Abstraction): Human authors are very creative. They often read a book, close it, and then write a sentence that sounds like it came from that book, even if the exact words weren't there. They mix ideas from different parts of the book to create a smooth story.
  3. AI is the "Strict Librarian": Because the AI is trained to be very careful and only say things it can prove with a direct quote, it turned out to be more factually accurate than humans when judged strictly.
    • The Analogy: Imagine a human writer is a jazz musician who improvises a beautiful melody that feels right but might not follow the sheet music exactly. The AI is a strict conductor who checks every note against the sheet music. In this specific game of "did you follow the rules?", the AI won.

The Scoreboard: How Do We Grade?

The paper argues that the old way of grading these summaries (counting how many words overlap between the AI's work and a human's work) is broken.

  • The Old Score: Like grading a student's essay by counting how many words they used that were in the teacher's example. If the student used different words to say the same thing, they got a bad grade.
  • The New Score: The authors created a "Statement-Level Judge." Instead of looking at the whole essay, they break it down sentence by sentence and ask: "Is this fact true? Is the citation correct?"
    • They found that standard grading tools are terrible at this. You need a smart AI judge (like a senior professor) to check the facts, not just a spell-checker.

The Conclusion

This paper says: "Stop giving AI just the back-cover blurbs. Give it the whole library."

  • The Dataset: It's a massive collection of open-access papers with full texts.
  • The Lesson: When AI has to read the whole book, it struggles to synthesize the information as smoothly as humans do, but it becomes incredibly good at sticking to the facts.
  • The Future: To build better tools for researchers, we need to stop testing AI on short summaries and start testing it on the full, messy, complex reality of reading entire academic papers.

In short: This paper built a giant library for AI to practice writing literature reviews. It discovered that while AI is better at following the rules of fact-checking than humans, humans are still better at weaving those facts into a smooth, natural story.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →