← Latest papers
🔬 materials science

SALSA: Semi-Autonomous Literature Summarization Assistant

SALSA is an open-source, human-in-the-loop platform that automates the extraction of structured scientific datasets from multimodal literature by combining document parsing, large language models, and computer vision with user-guided correction tools to support scalable data curation across disciplines.

Original authors: William Schertzer, Sonakshi Gupta, Rampi Ramprasad

Published 2026-09-22
📖 6 min read🧠 Deep dive

Original authors: William Schertzer, Sonakshi Gupta, Rampi Ramprasad

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Science has always relied on the careful reading of past work to build new knowledge. For decades, researchers have published their findings in journals, filling libraries with text, tables, and graphs that describe how materials behave, how chemicals react, and how experiments turn out. In recent years, the amount of this information has exploded. Computers can now run thousands of experiments at once, and scientists can share their results instantly with anyone in the world. This flood of data is a gift, but it comes with a problem: the information is written for human eyes, not for machines. A computer cannot easily understand that a number in a chart belongs to a specific chemical mixture described in a paragraph three pages earlier, or that a graph showing a temperature change is linked to a table of ingredients in a footnote. To use this data for new discoveries, especially in fields like materials science where researchers design new substances for batteries or solar panels, someone must manually read every document, find the relevant numbers, and type them into a structured list. This process is slow, expensive, and limits how much knowledge can be turned into new technology.

To solve this bottleneck, a team of researchers at the Georgia Institute of Technology has developed a new software tool called SALSA, which stands for Semi-Autonomous Literature Summarization Assistant. The tool is designed to act as a bridge between the messy, complex world of scientific papers and the clean, organized data needed for modern analysis. Instead of trying to replace the human expert, SALSA works alongside them. It uses advanced computer programs to read documents, find pictures and charts, and pull out numbers, but it leaves the final decision-making to a person. The system can take a scientific article, whether it is a digital file downloaded from a journal or a scanned image of a lab notebook, and break it down into its parts. It reads the text, reconstructs tables that might have been broken up by page breaks, and even looks at graphs to extract the data points hidden inside them.

The software is built to handle the specific challenges of scientific literature, where information is often scattered. A single experiment might have its chemical recipe in one section, a table of results in another, and a graph showing the performance in a figure with a caption that explains the conditions. SALSA connects these dots. It can identify that a line on a graph represents a specific material mentioned in the text, and it can link that material to the temperature and pressure conditions listed in a table. When the software encounters a chart, it uses a two-step process to read the data. First, it uses optical character recognition to read the numbers on the axes and the labels. Then, it traces the lines or dots on the graph to convert their position into actual numerical values. If the computer gets confused by a blurry image or a complex layout, the system pauses and asks a human user to step in. The user can look at the graph on their screen, correct the axis labels, or adjust the line tracing, ensuring the data is accurate before it is saved.

This approach is different from other tools that try to automate the entire process without human help. Those fully automated systems often make mistakes that are hard to spot, such as pulling the wrong number from a table or misinterpreting a graph because it looks similar to another one. SALSA avoids this by keeping a human in the loop. The software organizes the work into stages, allowing the user to check the results at each step. If the system extracts a list of chemical ingredients, the user can verify that the list matches the experiment described in the paper. If the system digitizes a graph, the user can confirm that the scale is correct. This interaction ensures that the final dataset is not just a collection of numbers, but a reliable record that preserves the context of the original research. The researchers tested the system on a variety of scientific articles, including those from different publishers and with different layouts, and found that it could successfully organize the data into a format ready for further analysis.

The tool is particularly useful for fields like chemistry and materials science, where the details of how a substance was made are just as important as the results it produced. A number like "conductivity" means nothing without knowing what material was tested, how it was processed, and under what conditions. SALSA is designed to capture all of these details together. It can be configured to look for specific types of information, such as the degradation of a membrane over time or the strength of a new alloy. The researchers showed that the system could take a set of articles about anion exchange membranes, find the relevant graphs and tables, and build a dataset that links the conductivity values to the specific aging times and testing conditions. This kind of structured data is essential for training artificial intelligence models to predict new materials, but it has been too difficult to create on a large scale until now.

The software is open-source, meaning anyone can use it and modify it to fit their own needs. It can be run on a local computer, which is important for researchers who work with sensitive or proprietary data that cannot be sent to external servers. The system does not grant permission to access or copy documents that a user does not already have the right to use; it simply helps organize the information that the user is already allowed to see. By automating the repetitive parts of data collection, SALSA frees up scientists to spend more time on interpretation and discovery rather than on typing numbers into spreadsheets. The researchers emphasize that the goal is not to remove human judgment from science, but to make that judgment more effective. The tool handles the heavy lifting of finding and extracting data, while the expert ensures that the data makes sense. This combination of automation and oversight allows for the creation of large, high-quality datasets that can drive the next generation of scientific breakthroughs, turning the vast library of past research into a usable resource for the future.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →