← Latest papers
💬 NLP

PaperTrail: A Claim-Evidence Interface for Grounding Provenance in LLM-based Scholarly Q&A

The PaperTrail interface introduces a granular claim-evidence mapping mechanism to improve provenance in LLM-based scholarly Q&A, but a user study reveals that while it successfully lowers researchers' trust by exposing unsupported claims, it fails to change their reliance on the system due to the high cognitive cost of verification.

Original authors: Anna Martin-Boyle, Cara A. C. Leckey, Martha C. Brown, Harmanpreet Kaur

Published 2026-02-25
📖 4 min read☕ Coffee break read

Original authors: Anna Martin-Boyle, Cara A. C. Leckey, Martha C. Brown, Harmanpreet Kaur

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a complex case. You have a very fast, very confident assistant (the AI) who writes up a report for you. The assistant says, "The suspect was at the scene because of X, Y, and Z."

In the past, if you asked the assistant, "Where did you get that info?", they might just point to a giant library book and say, "It's in there somewhere." This is like the standard AI tools we have today: they give you a source citation, but it's a big, blurry target. You have to guess if the book actually says what the assistant claims.

PaperTrail is a new tool designed to fix this. Instead of just pointing to the whole book, PaperTrail acts like a super-organized fact-checker. It breaks the assistant's report down sentence by sentence and compares it directly to the source documents.

Here is how the paper explains this, using simple analogies:

1. The Problem: The "Confident but Wrong" Assistant

Scientists and researchers are drowning in too many books (scientific papers). They use AI to help summarize these books. But AI has a bad habit: it sometimes hallucinates. It makes up facts or leaves out important details, but it sounds so confident and smooth that you might believe it.

Current tools try to help by showing you a link to the source paper. But that's like a teacher saying, "The answer is in Chapter 5," without telling you which sentence. You still have to do all the hard work to find the truth.

2. The Solution: The "Claim-Evidence" Map

The researchers built PaperTrail, which works like a highlighting system for truth.

  • The Claim: When the AI makes a statement (e.g., "Mars bases need 50 people"), PaperTrail treats this as a "Claim."
  • The Evidence: It then scans the original scientific papers to find the exact sentences that support that claim.
  • The Match:
    • Green Card (Teal): The AI made a claim, and we found the exact proof in the source paper. ✅
    • Red Card: The AI made a claim, but we couldn't find proof in the source paper. It might be made up! ❌
    • Missing Card: The source paper had an important fact, but the AI forgot to mention it. ⚠️

Think of it like a restaurant menu with a "Chef's Note" next to every dish.

  • Standard AI: "We have a steak." (You have to guess if it's good).
  • PaperTrail: "We have a steak. Source: The cow was raised on organic grass (Page 4, Paragraph 2). Missing: We didn't mention the salt content."

3. The Experiment: Did It Work?

The researchers tested this with 26 real scientists (like NASA engineers). They asked them to use two different tools to edit a draft report:

  1. The Old Way: Just standard citations (pointing to the whole book).
  2. PaperTrail: The detailed Claim-Evidence map.

The Results were surprising:

  • The "Trust" Drop: When scientists used PaperTrail, they trusted the AI less. They became much more skeptical. They saw the red cards (unsupported claims) and realized, "Oh, this AI isn't perfect."
  • The "Action" Gap: Even though they trusted the AI less, they didn't actually change their behavior much. They still mostly kept the AI's text.

Why didn't they fix the errors?
Imagine you are a detective, and your super-fast assistant gives you a report with some red flags. You know the report is shaky. But:

  • Time Pressure: You have a deadline in 20 minutes.
  • Cognitive Load: The PaperTrail interface was a bit cluttered and slow to load.
  • The Result: Even though you knew the AI was wrong, you were too tired and rushed to go back and fix every single sentence. You just accepted the "good enough" version to get the job done.

4. The Big Lesson

The paper teaches us a valuable lesson about human nature and technology:

Knowing the truth isn't enough to change what we do.

Even when we have a tool that clearly shows us where an AI is lying or guessing, we often don't fix it because we are too busy, too tired, or the tool is too hard to use.

The Takeaway:
PaperTrail is a brilliant idea for making AI honest. It's like giving a student a red pen that automatically circles every lie. But if the student is in a rush and the red pen is heavy and hard to hold, they might just ignore the circles and hand in the paper anyway.

To make AI truly useful for serious work, we need not just better truth-telling tools, but also tools that are easy to use when we are in a hurry.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →