From Traditional Scientists to Agentic AI Research: Unequivocal Diagram-as-Code in Scientific Publications
This paper advocates for adopting "diagrams-as-code" in scientific publications to eliminate interpretive ambiguity, thereby enabling reliable information handoffs between scientists, large language models, and autonomous AI agents for future scientific discovery.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Science has always relied on two languages to share its discoveries: the written word and the picture. A researcher might describe a complex biological process in a paragraph of text, or illustrate it with a diagram showing how different parts of a cell interact. For decades, these methods have served human readers well, allowing scientists to build upon each other's work. Yet, both language and images carry a hidden flaw: they are open to interpretation. A sentence can be read in slightly different ways depending on a reader's background, and a drawing can be understood differently based on how a viewer interprets the shapes and arrows. This ambiguity is manageable for humans, but it creates a significant problem for the new generation of artificial intelligence systems designed to read scientific literature and make discoveries on their own. These systems, known as AI agents, need information that is precise and unambiguous, leaving no room for guesswork. If an AI misinterprets a diagram or a description, it can make errors that ripple through its work, leading to incorrect conclusions.
In a recent study, researchers Mila Glavaški and Lazar Velicki from the University of Novi Sad in Serbia proposed a solution to this problem. They suggest that scientific papers should include "diagrams-as-code." Instead of just uploading a static image file, authors would provide the actual computer code used to generate that image. This code acts as a strict, mathematical instruction set that tells a computer exactly how to draw every line, shape, and connection. Because code is a language of logic rather than art, it removes the guesswork. When a computer reads the code, it sees the exact same structure that the author intended, with no variation in meaning. The researchers tested this idea by taking three common types of scientific content about heart failure—textual descriptions of causes, explanations of biological mechanisms, and clinical pathways for patient care—and converting them into this code format. They found that the approach works, creating a bridge that allows both human scientists and AI agents to understand scientific processes with total clarity.
The researchers began by looking at how scientific information is currently shared. They noted that while natural language allows for endless variety in expression, it also allows for endless variety in interpretation. A scientist might write that a heart condition is caused by high blood pressure, while another might describe it as a result of valve damage. Both are true, but the specific wording can lead an AI to miss the connection between the two. Similarly, a visual diagram might show arrows connecting different parts of a system, but without a strict definition, an AI might not know if those arrows represent a cause, a side effect, or a simple association. The goal of modern AI development is to create systems that can autonomously make scientific discoveries, acting as independent researchers that can read thousands of papers and synthesize new knowledge. To do this, these AI agents need to pass information to one another with perfect accuracy. If one agent hands off a task to another based on a misunderstood diagram, the entire chain of research can fail.
To test their idea, Glavaški and Velicki selected three distinct examples from the field of cardiology, the study of the heart. The first example was a paragraph describing the various ways heart failure develops, listing causes like ischemic injury, pressure overload, and metabolic disease. The second was a definition of heart failure as a clinical syndrome, detailing how structural problems in the heart lead to symptoms. The third was a step-by-step guide for doctors on how to evaluate a patient suspected of having heart failure, covering everything from medical history to physical exams. Alongside these texts, they used three corresponding diagrams that visually represented these same concepts. These inputs served as the raw material for their experiment.
The team then used a free version of a large language model, a type of AI chatbot, to convert these inputs into code. They asked the AI to generate code for two different diagramming tools, Mermaid and D2, which are programs that turn text instructions into visual charts. For the text descriptions, the prompt was simple: "Create code for a Mermaid diagram for this." For the existing images, the prompt was slightly more specific: "Create code for a Mermaid diagram for this figure. Do not add anything yourself." The researchers wanted to see if the AI could look at a paragraph of text or a static image and translate it into a set of logical instructions that another computer could read.
The results showed that the approach was feasible. For every text description and every image they tested, the AI successfully generated code that, when run through a standard editor, produced a diagram that matched the original meaning. For the text about the shared mechanisms of heart failure, the code created a flowchart showing how different pathways lead to the same outcome. For the text defining the disease, the code produced a diagram linking various causes to the condition. For the clinical pathway, the code mapped out the steps a doctor should take, from the first visit to the physical examination. The researchers also converted the original static images into code by prompting the AI to generate code based on the figures. The AI managed to produce code instructions that, when rendered, recreated the visual structure of the molecular mechanisms, the pathophysiological cascade, and the clinical pathway.
Once the code was generated, the researchers did not just accept it blindly. They manually checked the output in the online editors for Mermaid and D2. They found that while the AI got the general structure right, it sometimes needed minor adjustments. For instance, the direction of an arrow might need to be flipped, or a connection between two blocks might need to be moved to a different spot to accurately reflect the science. These small edits were necessary to ensure the code perfectly mirrored the intended representation. However, the fact that the core structure could be generated automatically was the key finding. The code served as the unchanging, precise record of the information, while the visual diagram was simply a way for humans to confirm that the code was correct.
The researchers emphasize that this method does not replace the need for human oversight. If an AI generates code based on a hallucination—a false fact that the AI invents—the resulting diagram will be wrong. However, the system they propose includes a built-in check. Because the code is human-readable and can be rendered instantly, a scientist can look at the generated diagram and immediately see if it matches the text or the original image. If it does not, they can fix the code. This process ensures that the final representation is accurate. The study suggests that in the future, scientific papers could include these code snippets alongside traditional text and images. This would allow AI agents to read the paper and extract the exact logic of the research without misinterpreting the visual or textual clues.
The implications of this work extend beyond just heart failure. The researchers argue that this approach could be applied to any scientific field where processes are described. Whether it is the flow of chemicals in a lab, the steps in a medical treatment, or the connections in an ecological system, representing these processes as code would make them universally understandable. It would allow a human scientist to hand off a project to an AI agent, and for that agent to pass it to another AI, with the guarantee that the meaning of the work remains unchanged at every step. The code acts as a universal translator, stripping away the ambiguity of natural language and the subjectivity of static images.
In their discussion, the authors acknowledge that this method has limits. It can only represent what is already known. If a scientific process involves unknown elements or relationships that have not yet been discovered, those gaps cannot be filled by code. Furthermore, the quality of the code depends on the quality of the AI generating it. Using more complex prompts might yield better results, but the researchers showed that even with simple instructions, the system works well enough to be useful. They also noted that while they used Mermaid and D2, other tools like PlantUML or Graphviz could serve the same purpose. The specific tool matters less than the concept of having a code-based representation.
The study concludes that the time has come to standardize how scientific information is shared in the age of AI. By adopting diagrams-as-code, the scientific community can ensure that the knowledge they produce is not just readable by humans, but also executable by machines. This shift would transform scientific literature from a collection of static documents into a dynamic, machine-readable resource. It would allow for a seamless transfer of knowledge between human researchers and artificial intelligence, paving the way for a future where AI agents can collaborate with scientists to solve complex problems with a level of precision that was previously impossible. The researchers demonstrated that this is not just a theoretical idea, but a practical reality that can be implemented today with existing tools.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.