← Latest papers
💬 NLP

On-Device Deep Research at 4B: Exposure Bounds Faithfulness, Retrieval Bounds Coverage

This study demonstrates that for on-device 4B research agents, citation faithfulness is primarily determined by the amount of source text exposed to the model rather than source quality, while the coverage of correct sources remains strictly limited by retrieval recall.

Original authors: Vinay Kumar Chaganti

Published 2026-07-15
📖 5 min read🧠 Deep dive

Original authors: Vinay Kumar Chaganti

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a super-smart, pocket-sized robot librarian named "4B." This robot lives on your personal laptop and its job is to read a bunch of research papers, figure out the answers to tricky questions, and write a short report with citations (like "according to page 5 of this book").

But here's the catch: sometimes the robot lies. It might say, "The book says X," when the book actually says Y. Or, it might pick the wrong book entirely. Researchers wanted to know: How do we make this little robot tell the truth, and what does it cost?

They ran a massive experiment on a laptop with 24 GB of memory to find the secret sauce. Here is what they discovered, using two simple levers.

Lever 1: The "Peek-a-Boo" Effect (Exposure)

Think of the robot's reading glasses. In the first test, the robot was only allowed to peek at the first 400 characters (about a short paragraph) of each paper before writing its report. It was like trying to judge a whole movie by only watching the first 30 seconds.

The result? The robot was shaky. It only got about 45% of its claims right when looking at real search results, and even worse (37%) when looking at perfect, "gold standard" papers. It was guessing too much.

Then, the researchers gave the robot better glasses. They let it read 1,500 characters (about a full page) of each paper. Suddenly, the robot's brain lit up!

  • On real search results, its accuracy jumped to 58%.
  • On perfect papers, it also hit 58%.

The Big Surprise: It didn't matter if the paper was perfect or just a random search result. Once the robot could read enough of the text, it supported its claims just as well in both cases. The paper argues that faithfulness is bound by how much the robot sees, not by whether the source is perfect.

However, the paper is careful to say this isn't a magic fix. Even at its best, the robot still got about 42% of its claims wrong. It's better, but it's not perfect yet. Also, this "reading more" trick only worked up to a point; reading even more didn't help much after 1,500 characters.

Lever 2: The "Library Search" Problem (Retrieval)

Now, imagine the robot is reading a perfect book, but the book isn't even on the shelf! The researchers found a second problem: Coverage.

Even when the robot read 1,500 characters, if the search engine didn't find the right book in the first place, the robot couldn't cite it. The "Trustworthy Coverage" (how often the robot cites the right source and gets the facts right) stayed stuck around 0.22 (or 22%) no matter how much text they let the robot read.

Why? Because the search engine was only finding about 40% of the right papers to begin with. The robot can't cite what it can't find. The paper explicitly rules out the idea that "reading more" will fix this. If the robot never sees the right book, reading more of the wrong books won't help.

What the Robot Can't Do (The "Repair" Myth)

You might think, "What if we let the robot write a draft, then check its work and fix the mistakes?" The researchers tested this too. They had the robot try to "gate" (drop bad claims), "revise" (rewrite bad claims), or "reattribute" (move citations).

The result? The robot got slightly better at supporting the claims it did make, but its overall ability to find the right sources didn't improve. In fact, trying to fix things after the fact often made the robot cite fewer sources, which lowered its overall score. The paper suggests that repairing a report after it's written is too late; the damage is done if the right sources weren't found or read enough in the first place.

The "Recipe" for a Better Robot

So, what's the practical takeaway for building a robot that runs on your laptop?

  1. First, give it more to read. Let the robot see 1,500 characters of every source it finds. This is cheap (it only costs about 235 extra tokens of computer memory) and it boosts the robot's truthfulness from 45% to 58%.
  2. Second, fix the search. Once the robot is reading enough, the only thing left to fix is the search engine. You need to make sure the robot actually finds the right books before it starts reading.

The Bottom Line

This study didn't create a perfect robot. It found that a small robot on a laptop can be surprisingly good at telling the truth if it's allowed to read enough of the source material. But it also proved that no amount of reading will help if the robot is searching the wrong library.

The authors are confident about the "reading more" part (it worked consistently across different tests and even with a second, independent judge). They are also confident that "fixing it later" doesn't work well. But they admit that the absolute numbers are still modest—the robot is getting better, but it's not ready to replace a human researcher just yet. It's a map showing us where to look next, not a finished destination.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →