Symposium: Trust via Auditable Records for Communities of AI Scientist Agents
Symposium is a formal framework and practical implementation that enables small scientific research communities to maintain immutable, auditable records of AI agent-driven research, thereby fostering trust and continuity by separating durable scientific history from the evolving AI systems that utilize it.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Science has always relied on a simple, human rhythm: a researcher makes a claim, gathers proof, and writes it down for others to check. This system works because the record is stable; a paper published today can be read and verified by a colleague ten years from now, who can trace every step back to the original data. But a new kind of scientist is entering the laboratory: the artificial intelligence agent. These programs can read thousands of papers, design experiments, and analyze data far faster than any human. They promise to accelerate discovery, yet they carry a unique risk. Because they are so fast and so capable, they can also make mistakes that are hard to spot, such as inventing facts, misreading sources, or subtly twisting data to fit a desired outcome. When a human researcher makes an error, it is usually obvious. When an AI agent does, the error can be buried inside a mountain of generated text, making it difficult to know what to trust.
The question facing the scientific community is not just how to build better AI, but how to build a system where we can trust what the AI says. If a machine generates a hypothesis or analyzes a dataset, how do we know the reasoning is sound? How do we know it didn't just guess? The answer lies not in changing the AI itself, but in changing how we record its work. We need a way to capture the AI's entire thought process, from its first guess to its final conclusion, in a format that is permanent, transparent, and impossible to alter after the fact. This is the challenge that Dexter Pratt, a researcher at the University of California San Diego, addresses in a new framework called Symposium.
Symposium is not a new type of AI agent, nor is it a software tool that helps agents do their jobs. Instead, it is a digital filing system designed specifically for the era of artificial intelligence. Imagine a library where every book, every note, and every scrap of paper is locked in a glass case that cannot be opened or changed once it is sealed. In this library, the "books" are not just final reports, but the entire trail of reasoning that led to a conclusion. The system records every hypothesis an AI agent considers, every piece of data it examines, and every assumption it makes along the way. Crucially, it forces the AI to declare exactly which pieces of information it is using as proof. If an agent claims a certain drug works, it must point to the specific data table or experiment that supports that claim, and it cannot use any other information as evidence unless it explicitly states that it is doing so.
The core idea behind Symposium is that trust is not a single score or a simple "yes" or "no" judgment. It is a decision made for a specific purpose. A researcher might decide to trust an AI's finding enough to run a small follow-up experiment, but not enough to base a major medical recommendation on it. Symposium captures this nuance by requiring every entry to state its purpose and the stakes involved. When an AI publishes a result, it must explain why it reached that conclusion, what assumptions it had to make, and what the consequences would be if it were wrong. This creates a clear, auditable trail. If a future scientist, or even a different AI agent, wants to build on that work, they can look at the record and see exactly how the original conclusion was reached. They can check the evidence, see if the assumptions hold up, and decide for themselves whether to trust the result.
The system works by treating every piece of research as a permanent record called an "Artifact." These Artifacts are linked together in a history that cannot be rewritten. If an AI makes a mistake or finds new information that changes an earlier conclusion, it does not delete or edit the old record. Instead, it publishes a new record that explains the correction, leaving the original record intact so that the history of the mistake and the fix remains visible. This approach mirrors the way human science handles corrections, but it applies it to the rapid, high-volume output of AI agents. The framework is designed to be flexible, allowing different groups of researchers to use their own tools and methods while still adhering to the same rules for recording their work. It separates the history of the research from the agents that produce it, ensuring that the record survives even as the technology behind the agents changes.
Pratt's paper introduces the rules for this system and provides a working example to show how it functions in practice. The example shows an AI agent generating a scientific argument, breaking it down into claims, evidence, and assumptions, and linking them together in a way that a human reader can follow. The system forces the agent to be explicit about what it knows and what it is guessing. It requires the agent to state clearly which data points are being used as proof and to explain the reasoning behind every step. This level of detail is what makes the system "auditable." It allows anyone, human or machine, to trace a conclusion back to its source and verify that the logic holds up.
The paper also addresses the problem of "hallucination," where AI agents invent facts or citations. In the Symposium framework, an agent cannot simply cite a source; it must declare that source as evidence and explain why it is relevant. If the agent tries to use a piece of data that it has not properly declared, the system rejects it. This creates a barrier against the kind of subtle deception that can occur when an agent tries to make a weak argument look strong. The system does not guarantee that the AI will always be right, but it guarantees that the reasoning will be visible. If the AI is wrong, the error will be obvious because the chain of evidence will be broken or the assumptions will be clearly stated.
One of the most significant aspects of this work is its focus on the "lore of the laboratory." In traditional science, researchers often keep a record of their failed experiments, dead ends, and negative results in personal notebooks that are rarely shared. These records are valuable because they prevent others from wasting time on the same mistakes. Symposium encourages the recording of all work, not just the successful findings. By allowing agents to publish their entire workflow, including the parts that did not work, the system creates a shared memory for the scientific community. This is particularly important for AI, which can generate vast amounts of data and analysis in a single day. Without a system to record and organize this information, much of it would be lost or forgotten.
The framework is designed to be used by small communities of researchers who are deploying AI agents in their labs. It provides the infrastructure for these communities to set up their own shared record, where agents can publish their work and build on the work of others. The system is extensible, meaning that as new types of AI tools emerge, the framework can adapt to include them without breaking the existing record. This flexibility is crucial because the technology is evolving rapidly, and a rigid system would quickly become obsolete.
The paper does not claim that Symposium is a perfect solution or that it will solve all the problems of AI in science. It is a proposal for a new way of recording research, one that prioritizes transparency and accountability. The author acknowledges that the system relies on the agents to follow the rules, and that there is no way to force an agent to be honest. However, by making the reasoning process visible and permanent, the system makes it much harder for an agent to hide its mistakes. It shifts the burden of trust from the agent itself to the record it leaves behind.
In the end, Symposium is about creating a space where AI and humans can work together with confidence. It is a framework that recognizes the power of AI to accelerate discovery while respecting the need for human oversight and verification. By providing a clear, immutable record of how scientific conclusions are reached, it allows the scientific community to assess the trustworthiness of AI-generated work on a case-by-case basis. This is not a magic bullet, but it is a practical step toward a future where AI can be a trusted partner in the scientific process, rather than a black box that produces results we cannot understand. The work lays the foundation for a new kind of scientific collaboration, one that is built on the solid ground of auditable evidence and shared history.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.