AgentCPM-Report: Interleaving Drafting and Deepening for Open-Ended Deep Research
AgentCPM-Report is a lightweight, local deep research system that leverages a human-like "Writing As Reasoning" policy to dynamically interleave drafting and deepening, enabling small 8B-parameter models to outperform leading closed-source systems in generating insightful reports.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Problem: The "Rigid Blueprint" Trap
Imagine you are asked to write a massive, 50-page encyclopedia entry about a topic you know very little about, like "The Future of AI."
Most current AI systems work like a rigid architect. Before they write a single word, they try to draw a perfect, detailed blueprint of the entire building. They decide exactly what every room will look like before they even lay the first brick.
- The Flaw: If the architect makes a mistake in the blueprint (which is hard to do perfectly without knowing everything), the whole building is doomed. They can't easily change the plan once construction starts.
- The Consequence: To make a good blueprint, you need a genius architect (a massive, expensive, closed-source AI). If you use a smaller, cheaper AI, the blueprint is usually bad, and the final report is shallow and boring.
The Solution: The "Sculptor" Approach
The authors of this paper created a new system called AgentCPM-Report. Instead of an architect, they treat the AI like a sculptor.
A sculptor doesn't plan every curve of the statue in advance. They start with a rough block of stone (a simple outline), chip away some parts, look at the shape, and then realize, "Oh, I didn't think about this detail; I need to dig deeper here." They change the plan while they are working.
This is called WARP (Writing As Reasoning Policy). It means the act of writing is the act of thinking. The AI writes a section, realizes it's too shallow, stops, goes to find more information, updates the plan, and then writes again.
How It Works: The Two-Step Dance
The AI alternates between two modes, like a dancer switching steps:
- Evidence-Based Drafting (The Scribe): The AI picks a section of the outline, searches the internet for facts, and writes a paragraph. It's like a scribe copying down what they found.
- Reasoning-Driven Deepening (The Detective): After writing, the AI steps back and asks, "Is this good enough? Did I miss something?"
- If the answer is No, it acts like a detective. It finds a gap in the story, breaks that section into smaller, more specific questions, and updates the outline.
- If the answer is Yes, it moves on.
This loop allows the AI to discover new ideas during the writing process, rather than being stuck with a bad plan from the start.
The Training: Teaching a Small Dog to Hunt
The paper uses a relatively small AI model (only 8 billion parameters, which is "small" in the world of AI). Usually, small models are bad at complex planning. To fix this, the team used a Multi-Stage Training Strategy:
- Cold Start (The Basics): They taught the AI the alphabet. It learned how to search, how to write a sentence, and how to follow basic rules.
- Atomic Skill RL (The Drills): They practiced specific moves. They taught the AI how to write a good paragraph or how to ask a good search question, rewarding it for getting those small tasks right.
- Holistic Pipeline RL (The Marathon): Finally, they let the AI run the whole race. They didn't just check if the sentences were good; they checked if the entire report was insightful.
- Crucial Trick: They used a "Trajectory Pruning" method. Imagine a teacher writing a long story but stopping at random times. The AI looked at all the drafts the teacher made, picked the best one, and said, "Stop here!" This taught the small AI exactly when to stop digging for more info so it doesn't waste time.
The Results: Small Model, Big Brain
The paper tested this system against the biggest, most expensive AI systems from companies like Google, OpenAI, and Anthropic.
- The Surprise: The small, local AI (AgentCPM-Report) beat the giant, closed-source systems in creating Insightful reports.
- Why? Because the "Sculptor" approach (WARP) allowed the small AI to think on its feet. It didn't need a massive brain to draw a perfect blueprint; it just needed to be good at adjusting its plan as it went.
Why This Matters (According to the Paper)
- Privacy: Because this system is small enough to run on your own computer (or a local server), you don't have to upload your private data to a big company's cloud to get a deep research report.
- Cost: You don't need to pay for expensive, massive AI models to get high-quality research.
- Quality: The reports are deeper and smarter because the AI is allowed to change its mind while it works, just like a human researcher does.
In short: The paper shows that you don't need a giant brain to do deep research; you just need a smart way of working (WARP) that lets the AI learn and adapt as it writes.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.