A Human-in-the-Loop Corpus for LLM-Based Simplification of Scientific Summaries
This paper presents a human-in-the-loop workflow and a new corpus derived from SciSummNet that leverages GPT-4o-mini and expert feedback to create accessible scientific summaries, demonstrating that while AI-generated simplifications improve comprehensibility for non-specialists, human experts are essential for preserving critical domain terminology and claim strength.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the world of science as a massive, bustling library where every book is written in a different, highly specialized dialect. A physicist speaks in equations, a biologist in cellular codes, and a computer scientist in layers of abstract logic. While this precision is necessary for experts to build new discoveries, it creates a giant wall for anyone trying to peek over the fence. It's like trying to read a recipe written entirely in chemical formulas; you know it's about food, but you have no idea how to cook the meal. This is the problem of "interdisciplinary research"—when scientists from different fields need to talk to each other, or when the general public needs to understand a breakthrough, the language often gets in the way.
To fix this, researchers are turning to "Large Language Models" (LLMs). Think of an LLM as a super-smart, tireless translator that has read almost everything ever written. It's great at taking complex text and rewriting it in simpler words, much like a friend explaining a complicated movie plot to you over coffee. However, there's a catch: sometimes this super-translator gets too eager. It might accidentally change the meaning of a scientific fact, swap a precise technical term for a vague guess, or sound a bit too casual for a serious topic. The big question is: Can we use these AI tools to make science readable without breaking the science itself?
This paper, written by Kyuri Im and Michael Färber, tackles that exact puzzle. They didn't just ask the AI to do the work and hope for the best; instead, they built a "Human-in-the-Loop" workflow. Imagine a relay race where the AI runs the first leg, a group of curious non-experts runs the second leg to spot the stumbling blocks, and a team of expert scientists runs the final leg to fix any mistakes.
The researchers started with 1,000 summaries of computer science papers that were already written by experts but were still too dense for outsiders. First, they let an AI (specifically GPT-4o-mini) rewrite these summaries to be simpler. Then, they brought in 92 people from STEM fields (like biology or engineering) who didn't know computer science. These readers acted as the "test audience." They compared the original text with the AI's version, pointing out sentences that were still confusing and rating how easy the text was to understand, how natural it sounded, and how simple it was.
The results were a mix of good news and a reality check. The "test audience" overwhelmingly preferred the AI's versions. In 73 out of 92 cases, they said the AI made the text easier to understand, and in 70 out of 92 cases, they found it simpler. The AI was great at breaking down long, tangled sentences into bite-sized pieces. However, the human readers also noticed that the AI sometimes sounded a bit awkward or used words that felt too informal for serious science.
This is where the second part of the race began. The researchers took the feedback from the non-experts and handed it to computer science experts. These experts acted as the "editors." Their job wasn't to rewrite everything from scratch, but to carefully tweak the AI's work. They fixed the awkward phrasing and, crucially, made sure that important technical terms (like "latent variables" or "NP-hard") were kept or explained correctly, rather than being replaced with vague guesses.
The paper found that while the AI is a fantastic starting point—like a rough draft that gets the main ideas across—it isn't ready to be published on its own. The automatic computer tests showed that the AI's version looked very similar to the original text and scored high on "readability" formulas. However, when the experts stepped in to polish the text, the computer metrics actually dropped a bit because the experts changed the wording more significantly to ensure accuracy. But this drop in "similarity" scores was a good thing: it meant the experts were successfully preserving the scientific truth while making it accessible.
In the end, the authors suggest that the best way to simplify science isn't to rely solely on AI, nor to do it all by hand. Instead, it's a partnership. The AI does the heavy lifting to make the text simpler, the non-experts tell us where it's still confusing, and the experts make sure the science stays true. The researchers have released all their data, including the original summaries, the AI drafts, the reader feedback, and the final expert-edited versions, so that others can build better tools for this important job. They conclude that while AI is a powerful helper, it needs a human guide to ensure that when we simplify science, we don't accidentally lose the magic of the discovery.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.