Towards Verifiable Multimodal Deep Research: A Multi-Agent Harness for Interleaved Report Generation
This paper introduces \textsc{Ptah}, a multi-agent framework that orchestrates planning, evidence collection with a visual working memory, and verifier-enforced report generation to produce reliable, citation-faithful, and visually interleaved multimodal deep research reports.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you ask a smart assistant to write a detailed, professional report on a complex topic, like "How do deep neural networks work?" or "What is the current state of the electric vehicle market?"
In the past, these assistants would just spit out a wall of text. Sometimes, they'd get facts wrong (hallucinate), and if they tried to add pictures, the images would often be random, unrelated, or placed awkwardly, like a collage made by someone who didn't read the instructions.
This paper introduces PTAH, a new system designed to act like a master construction crew for building these reports. Instead of one lone worker trying to do everything at once, PTAH uses a team of specialized agents working together under strict supervision to build a report that mixes text and images perfectly.
Here is how PTAH works, broken down into simple steps:
1. The Blueprint (Planning)
First, a Planner Agent acts like an architect. Before laying a single brick, it draws a detailed blueprint. It decides:
- What sections the report needs.
- What facts need to be found.
- Crucially: Where images should go and what kind of images are needed (e.g., "We need a chart here to show sales growth," or "We need a diagram to explain the engine").
2. The Gathering Crew (Research)
Next, a team of Researcher Agents goes out into the internet to gather materials.
- They find the facts and write them down.
- They also act like a visual scavenger hunt. They don't just grab any picture; they look for specific images that match the blueprint.
- They put these "source-aligned" images (images that actually come from the websites they visited) into a special Visual Working Memory. Think of this as a digital toolbox where every tool (image) is labeled with exactly where it came from and what it's supposed to prove.
3. The Quality Control Inspector (The Verifier)
This is the most important part. Before the report moves to the next stage, a Verifier Agent acts as a strict building inspector.
- It checks: "Did you actually find the facts you claimed?"
- It checks: "Is this picture actually from the website you cited?"
- It checks: "Does this image make sense next to this paragraph?"
- If the answer is "No," the crew has to go back and fix it. This prevents the system from making up fake facts or pasting random, confusing pictures.
4. The Assembly Line (Writing)
Finally, a Writer Agent assembles the final product. Instead of just typing text and hoping an image fits later, it writes the text and the image instructions together, like a conductor leading an orchestra.
- It uses the "toolbox" of verified images.
- If a specific chart is needed that doesn't exist yet, it can generate one.
- It then turns the whole thing into a polished, interactive webpage (like a mini-website) that looks professional and is easy to read.
The "Test-Time Scaling" (The Polish)
Before handing the report to you, PTAH does a final round of "polishing." It looks at the whole report again to fix spacing, make sure the images aren't too crowded, and ensure the layout looks good on a screen. It's like a final walkthrough before the house is sold.
How They Tested It (PTAHEval)
The researchers realized that old tests only checked if the text was good. They created a new test called PTAHEval to judge the whole package:
- Image Content Quality: Is the picture clear? Does it actually help explain the text?
- Presentation Quality: Does the final webpage look good? Is it easy to read, or is it a messy jumble?
The Results
When they tested PTAH against other smart systems:
- Better Facts: PTAH made fewer mistakes and cited real sources much more often.
- Better Pictures: The images were actually relevant and helpful, not just decoration.
- Better Layout: The final reports looked like professional documents, not messy drafts.
In short: PTAH is like hiring a team of architects, builders, and inspectors to build a house, rather than asking one person to guess the blueprint and build it all at once. The result is a report that is not only smart but also visually trustworthy and easy for humans to understand.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.