AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis
AskChem is a new infrastructure that transforms chemistry literature synthesis by shifting retrieval from whole documents to provenance-carrying atomic claims, enabling more accurate, verifiable, and AI-accessible cross-paper search and discovery.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the world of science as a massive, chaotic library where every book is a research paper. For decades, if you wanted to find a specific fact—like "what temperature makes this chemical react?"—you had to hunt through entire books, hoping to stumble upon the right paragraph. This is how traditional search engines work: they hand you a list of documents, and you, the reader, have to do the heavy lifting of opening them, scanning pages, and stitching together answers from different sources. It's like trying to build a puzzle when you're only given a pile of boxes labeled "Puzzle," rather than the individual pieces.
But science is changing. We now have super-smart computer programs called AI agents that can read, plan, and even run experiments. However, these digital helpers are stuck using the old library system. They get lists of papers, not the actual facts inside them. This leads to a frustrating problem: the AI might guess an answer or make up a fake citation (a "hallucination") because it can't easily verify the specific evidence hidden deep inside a document. The big question is: How do we upgrade the library so that both humans and AI can grab the exact, verified fact they need, instantly, without reading the whole book?
Enter AskChem, a new system that flips the script on how we search chemistry literature. Instead of treating a whole research paper as the smallest unit of information, AskChem breaks every paper down into tiny, atomic "claims." Think of a claim as a single, verified fact card—like a specific recipe step or a measured result—that comes with its own ID tag, a direct link to the original page it came from, and a quote proving it's real.
The researchers built a massive database containing 2.4 million of these fact cards, pulled from 147,000 different papers. They didn't just dump them in a pile; they organized them into three clever structures to make them easy to find:
- A Faceted Taxonomy: Imagine a giant, smart filing cabinet where you can sort facts by what they are about (like "catalysts"), how they were measured, or when they were discovered. This lets you browse and group similar findings instantly.
- An Evidence Graph: This is like a web of connections. It links facts together, showing you which findings support each other, which ones contradict, and how one discovery led to another across different papers.
- A Living Taxonomy: This is an exploratory map that places facts under big scientific ideas and principles, helping you see the "why" behind the "what."
The team tested this system with a tool called AskChem-Bench, asking it complex questions that required pulling information from many different papers. The results were striking. When an AI tried to answer these questions without help, it often made up citations, with only 88.3% of the links actually working. But when the same AI was "grounded" in AskChem's database of verified claims, 100% of the links were real and resolvable. The AI also provided answers with much higher "citation density," meaning it could back up its claims with more specific, verified evidence than other systems.
The authors suggest that this claim-centered approach is a practical foundation for the future of chemistry research. It doesn't replace the need to read the original papers for critical decisions, but it acts as a powerful, verified shortcut. By turning the library from a collection of books into a collection of verified, connected fact-cards, AskChem helps both scientists and AI agents find the truth faster and with fewer mistakes. The system is already live, open for anyone to explore, and suggests that the future of scientific search isn't about finding more documents, but about finding the right facts.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.