Large Language Models as Supervised Extraction Assistants: Lowering the Barrier to Documentation Standard Adoption in Agent-Based Modelling
This paper investigates the feasibility of using Large Language Models to automate the extraction of Rigour and Transparency Reporting Standard (RAT-RS) documentation from Agent-Based Modelling papers, finding that while LLMs can improve reporting consistency for descriptive tasks, they require human oversight for more complex evaluative work to ensure reliable adoption of documentation standards.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Picture: The "Homework" Problem
Imagine Agent-Based Modelling (ABM) as a complex recipe for a digital simulation. To make sure the recipe works and others can cook it, you need a detailed instruction manual (documentation). There are strict rules for writing these manuals, like the RAT-RS, which asks 39 specific questions about what data was used and why.
The Problem: Writing these manuals is like doing a mountain of extra homework. It takes a lot of time, and researchers often skip it or do a sloppy job because it feels like "supplementary" work rather than the main event. This is a "Tragedy of the Commons": everyone benefits if everyone writes good manuals, but no single person wants to pay the high "time cost" to do it.
The Proposed Solution: The authors suggest using Large Language Models (LLMs)—the smart AI chatbots we know today—as a "Junior Research Assistant." Instead of the senior researcher writing the whole manual from scratch, the AI does the heavy lifting of reading the paper and filling out the forms. The human then acts as the "Editor-in-Chief," checking the work, fixing mistakes, and signing off on it.
The Experiment: A Feasibility Test
The researchers wanted to see if this "AI Assistant" idea actually works. They didn't try to prove it works for every paper in the world; they just wanted to see if it was possible.
- The Test Subject: They picked one specific, published research paper about retail simulation.
- The Task: They asked four different top-tier AI models (like Claude, ChatGPT, and Gemini) to read that paper and fill out the 39-question RAT-RS form.
- The Comparison: They compared the AI's answers against the "Gold Standard"—a form that a human had already manually filled out for that same paper.
What They Found: The "What" vs. The "Why"
The results were a mix of "Great!" and "Be careful." The performance of the AI depended entirely on the type of question being asked.
1. The "Fact-Checker" Mode (High Success)
- The Questions: "What data did you use?" or "What was the publication date?"
- The Analogy: Think of this like a librarian looking for a specific book on a shelf.
- The Result: The AI was excellent at this. It could find the facts, quote the text, and organize the information perfectly. It was often even more detailed than the human who wrote the original report.
2. The "Reasoning" Mode (Low Success)
- The Questions: "Why did you choose this data?" or "Explain the logic behind this decision."
- The Analogy: This is like asking a student to explain why they solved a math problem a certain way, rather than just giving the answer.
- The Result: The AI struggled here. It gave plausible-sounding answers, but they were often inconsistent. If you asked the same AI the same "Why" question twice, it might give two slightly different reasons. It also sometimes missed the subtle nuances a human would catch.
3. The "Style" Gap
- Even when the AI got the facts right, it wrote in a very different style than a human. Humans write short, punchy, analogy-filled answers. The AI wrote long, formal, policy-heavy paragraphs.
- The Takeaway: The difference in wording made the computer's "similarity score" look low, even when the actual information was correct.
The Verdict: The "Supervised Extraction" Strategy
The paper concludes that we shouldn't let the AI work alone. Instead, we should use a Supervised Extraction workflow:
- The AI (The Junior): Reads the paper and drafts the answers. It's great at finding facts and describing what happened.
- The Human (The Senior): Reviews the draft. They don't need to write from scratch; they just need to verify the "Why" questions and fix any weird phrasing.
The "Triage" System:
The authors suggest a smart way to manage this:
- Green Light: For factual questions, trust the AI. It's reliable.
- Red Light: For "Explain why" or "Evaluate quality" questions, the human must step in. The AI is too inconsistent here.
The Call to Action
The authors aren't saying "AI will solve everything." They are saying:
- Don't reinvent the wheel: Use the AI to lower the barrier to entry. It makes doing the "boring" documentation work much faster.
- Be transparent: If you use AI to help write your report, say so. Tell people which AI you used and how you checked it.
- Work together: The community needs to share better "prompts" (instructions for the AI) so everyone gets better results.
In a Nutshell:
Think of the AI as a very fast, very well-read intern. It can scan a document and pull out all the dates, names, and data sources in seconds. But it can't yet fully understand the story behind those numbers. If you let the intern do the legwork and you do the final review, you get a high-quality report in half the time. If you let the intern run the show alone, you might get a report that looks good but misses the point.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.