SciSchema.org: A Multidisciplinary Collection of Schemas for Structured Scientific Process Descriptions
This paper introduces SciSchema.org, the first multidisciplinary collection of 16 expert-annotated schemas designed to standardize the description of scientific processes across diverse fields, thereby enabling structured annotation, reproducibility, and cross-study comparison through a human-in-the-loop development workflow.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Science has always relied on the written word to share how discoveries were made. A researcher describes an experiment in a journal article, weaving together lists of materials, settings on machines, and step-by-step instructions into paragraphs of text. This approach works well for human readers who can infer context and fill in gaps, but it creates a significant barrier for machines. Computers struggle to understand that "heated to 80 degrees" in one paper means the same thing as "maintained at 353 Kelvin" in another, or that a specific chemical process described in a biology lab follows a similar logic to one in a chemistry lab. As science becomes more automated and interdisciplinary, with artificial intelligence systems and robotic laboratories trying to read and execute these instructions, the lack of a common, structured language for scientific methods has become a bottleneck. Without a way to translate the prose of discovery into a rigid, machine-readable format, the potential for comparing studies, reproducing results, and automating research remains largely untapped.
To address this gap, a large, international team of researchers has introduced SciSchema.org, a new collection of structured templates designed to describe scientific processes in a way that both humans and computers can understand. The team, comprising experts from fields as diverse as biology, physics, psychology, and materials science, did not simply write these templates by hand. Instead, they developed a collaborative workflow where human experts guided artificial intelligence models to mine thousands of scientific papers and extract the essential components of specific experiments. The result is a set of 16 distinct "schemas," or structured outlines, covering processes ranging from the editing of genes using CRISPR-Cas9 to the reconstruction of neutrino events in particle physics. These schemas define exactly what information needs to be recorded for a given type of experiment—such as the specific instruments used, the conditions under which they operated, and the precise steps taken—transforming the messy, unstructured prose of traditional papers into clean, comparable data.
The project began by identifying 16 specific scientific processes that are well-established and frequently described in literature, ensuring there was enough material to work with. The team then recruited domain experts for each process, from molecular biologists to psychologists, to act as guides. These experts provided a basic description of their field's standard procedures and a collection of relevant scientific papers. The core of the work involved a "human-in-the-loop" system, where large language models—advanced AI tools capable of reading and understanding text—were tasked with reading these papers and proposing a structured outline for the process. The AI would scan the text, identify recurring details like temperature settings or chemical concentrations, and draft a schema that organized these details into logical categories.
However, the AI did not work alone. After the models generated their initial drafts, the human experts reviewed them, pointing out missing details, correcting terminology, and suggesting how different pieces of information should be grouped. This feedback was fed back into the AI, which then refined its drafts. This cycle of generation and expert correction was repeated over several stages, with the AI reading more and more papers at each step. The process was designed to be iterative, allowing the AI to learn from the experts' corrections and gradually build a more accurate and comprehensive structure. The team tested this workflow using 12 different AI models, ranging from smaller, faster models to massive, complex ones, to see which approaches produced the best results.
The outcome of this rigorous process was the creation of 16 final, expert-annotated schemas. Each schema serves as a master template for a specific type of scientific process. For instance, the schema for "Polymerase Chain Reaction," a common technique used to copy DNA, defines specific fields for the biological materials used, the design of the genetic primers, the thermal cycling conditions, and the resulting data. Similarly, the schema for "Fatigue Testing of Metallic Materials" outlines fields for the type of metal, the shape of the test sample, the forces applied, and the environmental conditions. These schemas are not filled with data from specific experiments; rather, they are the empty forms that scientists can use to record their own data in a standardized way. By using these templates, a researcher can ensure that every critical detail of their experiment is captured in a format that can be easily searched, compared with other studies, and processed by software.
The researchers found that the quality of the final schemas depended heavily on the collaboration between the AI and the human experts. While the AI models were capable of generating complex structures and identifying patterns across hundreds of papers, they often required human guidance to ensure scientific accuracy and appropriate terminology. The team observed that the AI models became more detailed and structurally complex as they were exposed to more papers and more rounds of expert feedback. The largest and most detailed schemas were generated by the most advanced instruction-following models, which produced outlines with hundreds of specific fields and multiple layers of organization. However, the final decision on what constituted the "gold standard" structure always rested with the human experts, who selected the best elements from the AI's various drafts to create the final version.
One of the key findings of the study was that this human-AI partnership could produce high-quality, domain-specific schemas much faster than traditional manual methods would allow. The team managed to develop and validate all 16 schemas in just two months, a timeframe that would have been impossible if experts had to manually extract and organize every detail from thousands of papers. The process also revealed that different types of scientific processes require different levels of detail. Some processes, like the "Stroop Task" used in psychology to measure reaction times, resulted in schemas with fewer fields and less nesting, reflecting the simpler, more standardized nature of the experiments. Others, like the "Bulk RNA-seq Library Preparation" in biology, resulted in highly complex schemas with deep layers of nested information, mirroring the intricate and multi-step nature of the wet-lab procedures involved.
The released collection includes not just the final schemas, but also the entire history of how they were created. The data archive contains the intermediate drafts generated by the AI, the feedback provided by the experts, and the metadata of the papers used in the process. This transparency allows other researchers to see exactly how the schemas were built, to understand the reasoning behind specific design choices, and to reuse the materials for their own projects. The schemas are available in two formats: one designed for standard computer documents and another for knowledge graphs, which are networks of linked data used to connect information across different sources. This dual format ensures that the schemas can be used in a wide variety of software systems, from simple data entry forms to complex scientific databases.
The significance of this work lies in its potential to change how scientific knowledge is stored and shared. By providing a structured way to describe the "how" of science, SciSchema.org makes it possible to move beyond the limitations of text-based searching. Instead of searching for articles that mention a specific keyword, researchers could eventually search for all experiments that used a specific combination of instruments, conditions, and materials, regardless of how the original authors described them in their prose. This capability is essential for the future of science, where the ability to compare, reproduce, and automate experiments will be critical for solving complex global challenges. The project demonstrates that while artificial intelligence is a powerful tool for organizing information, it is most effective when guided by human expertise, creating a synergy that can capture the nuance and depth of scientific practice in a machine-readable form.
The team made these resources freely available to the public, releasing the schemas, the code used to generate them, and the analysis tools under open licenses. This ensures that the scientific community can build upon this foundation, adding new schemas for other processes and refining the existing ones as needed. The project also includes connections to the Open Research Knowledge Graph, a system that allows these structured descriptions to be linked directly to the original scientific papers, creating a bridge between the raw text of discovery and the structured data of reproducibility. In doing so, SciSchema.org offers a practical step toward a future where the methods of science are as accessible, comparable, and reusable as the results themselves.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.