Scrub Data: A Framework for Reproducible Data Curation with AI Coding Agents
This paper introduces Scrub Data, a framework that leverages AI coding agents for data curation while ensuring reproducibility and verifiability through a provenance graph that tracks all transformations as executable steps, demonstrated by reducing 89 million raw GPS coordinates to a curated benchmark dataset.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
Data science is often imagined as a realm of complex algorithms and predictive models, but before a single model can be trained, a massive amount of work must be done to prepare the raw material. This preparatory phase, known as data curation, involves collecting messy information from the real world, cleaning up errors, and organizing it into a format that computers can understand. It is a tedious, time-consuming process where a single mistake—like deleting the wrong row of numbers or misinterpreting a date—can ruin an entire study. For years, scientists have relied on manual effort or rigid scripts to handle this work, ensuring that every change is recorded so that others can repeat the process exactly. However, the rise of artificial intelligence has introduced a new, tempting shortcut: asking a computer program to do the cleaning for you. While these AI agents can write code and perform tasks at incredible speed, they operate with a dangerous lack of transparency. If an AI deletes a file or alters a dataset without leaving a clear trail, the results become impossible to verify or reproduce, turning scientific discovery into a black box where the path from raw data to final conclusion is lost.
To solve this problem, researchers at MIT have developed a new framework called Scrub Data, designed to let scientists use powerful AI agents for data cleaning without sacrificing the ability to check their work. The system acts as a strict supervisor for the AI, ensuring that every single change made to a dataset is recorded as a self-contained, executable step in a permanent history log. Instead of allowing the AI to wander freely through files and make untraceable edits, the framework forces the agent to propose a specific script, run it through a controlled system, and then log the result. This creates a clear, unbroken chain of events, much like a flight recorder on an airplane, where every action is documented from the moment the raw data is collected to the final, polished dataset. The system also includes a visual interface that allows human researchers to see the data in real-time, build custom tools to spot errors, and verify that the AI's changes make sense before they are permanently saved.
The researchers tested this approach by tackling a massive and difficult project in animal movement ecology: creating a standardized benchmark dataset called MOVEBENCH. They began with a chaotic collection of 89 million raw GPS coordinates from wildlife tracking studies, gathered from hundreds of different sources with varying formats and quality. The goal was to clean this data down to a usable set of 2.6 million points from 807 individual animals across 110 species, suitable for training machine learning models to predict animal movement. Using the Scrub Data framework, a team of two researchers worked with an AI agent to navigate this mountain of information. The agent helped them write scripts to filter out bad timestamps, thin out data that was recorded too frequently, and select representative animals from different studies. Crucially, every decision the agent made was captured in a version history graph. If the team wanted to know how a specific animal's location was calculated, or why a certain study was excluded, they could trace the exact code and logic used to make that decision.
Throughout the process, the human researchers remained in control, using the framework to build custom visualization tools that let them see the animal tracks on a map and check for anomalies. When the AI suggested a change, such as removing rows with invalid dates, the system would execute the code, show the results, and ask for confirmation before saving the new version. This loop allowed the team to move quickly, leveraging the speed of the AI to handle repetitive tasks while maintaining the rigorous standards required for scientific research. The final result was a clean, reproducible dataset that was created in just three days, a task that would have taken weeks or months using traditional manual methods. The entire history of the curation process, from the initial raw files to the final benchmark, was preserved as a set of executable steps, ensuring that any other scientist could replay the process and arrive at the exact same result.
The success of this project highlights a shift in how scientific tools are built. Rather than treating AI as a replacement for human judgment, the Scrub Data framework treats it as a powerful assistant that must operate within a transparent, auditable structure. By separating the AI's ability to write code from the direct modification of data files, the system prevents the "vibe curation" approach, where users blindly trust AI outputs without verification. Instead, it enforces a workflow where every modification is a deliberate, recorded step. This approach does not eliminate the need for human expertise; in fact, it relies on it. The researchers still had to decide which animals to keep, how to handle missing data, and what the final dataset should look like. The framework simply ensures that the AI's contribution is reliable and that the path from raw observation to scientific insight remains clear and open to inspection.
This work suggests that the future of data science may not be about choosing between human precision and machine speed, but rather finding a way to integrate them safely. The Scrub Data framework offers a practical way to harness the efficiency of AI agents while keeping the scientific method intact. It proves that it is possible to automate the tedious parts of data cleaning without losing the ability to trace, verify, and reproduce the results. As the volume of data in fields like ecology, medicine, and climate science continues to grow, tools like this will be essential for managing the complexity of modern research. The framework is now available for other scientists to use and adapt, offering a foundation for building their own custom data curation workflows. By making the process of cleaning data transparent and reproducible, the researchers hope to enable a new generation of scientific discovery that is both faster and more trustworthy.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.