DoTime: A Synthetic Benchmark Generator for Interventional and Counterfactual Time Series
The paper introduces DoTime, a scalable and theoretically grounded synthetic benchmark generator for multivariate temporal structural causal models that addresses the lack of interventional and counterfactual time series data by providing open tools, evaluation suites with exact ground truth, and empirical evidence that interventional training improves causal direction accuracy over observational models.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to solve a mystery, but the only clues you have are a pile of old, dusty photos. You can see what happened in the past, but you can never see what would have happened if someone had made a different choice. This is the daily struggle for scientists who study "causality"—the art of figuring out what causes what. In the real world, we often just watch things happen (observational data), like watching a plant grow in the sun. But to truly understand cause and effect, we need to ask "What if?" questions (counterfactuals). What if we watered the plant instead? What if we turned off the lights? To answer these "What if?" questions, we usually need to run experiments, but in fields like medicine or climate science, running experiments can be dangerous, expensive, or even impossible.
So, scientists have started building "simulators"—digital worlds where they can run experiments safely. They create fake data that follows the rules of cause and effect, train their computer brains (AI models) on this fake data, and then hope those brains are smart enough to solve real-world mysteries. The big question is: How do we know if these digital test worlds are good enough? If the simulator is too simple, the AI learns the wrong lessons. If it's too messy, we can't tell if the AI is actually smart or just lucky. We need a perfect, fair, and challenging playground to test these AI detectives before they go out into the real world.
Enter DoTime, a new tool created by researchers Dennis Thumm, Billy Tim Anthony, and Ying Chen. Think of DoTime as the ultimate "flight simulator" for time-traveling detectives. It is a computer program that generates millions of fake, complex stories about how things change over time, complete with a secret "answer key" that tells you exactly what would happen if you changed the plot.
The paper introduces DoTime as a way to fix a major gap in how we test AI. Most existing test worlds only show the AI what did happen, not what could have happened. DoTime is different because it lets the AI practice "interventional" thinking. It creates scenarios where you can say, "Okay, in this story, let's pretend we gave the patient a different medicine," and the simulator instantly rewrites the future to show the result, while keeping everything else exactly the same. It even has a special "shared noise" mode, which is like having two identical twins: one lives their normal life, and the other gets a specific treatment. Because they are twins, any difference in their lives is guaranteed to be caused by that treatment, not by random luck.
The researchers used DoTime to test a specific idea: Does training an AI on these "What if?" stories make it better at solving real problems than just watching normal stories? They set up a fair race between two AI models of the exact same size. One model was trained only on normal, observational stories (watching what happens). The other was trained on the "What if?" intervention stories. The results were clear and measurable: the model trained on the "What if?" stories consistently got the direction of the cause-and-effect right more often than the one that only watched. In their simulations, this "interventional training" gave the AI a measurable advantage, improving its accuracy by about 8 to 9 percentage points compared to its observational twin.
However, the authors are careful not to claim they have solved the world's problems. They explicitly rule out the idea that the AI just got lucky or that it simply memorized the answers. They tested the AI on different types of stories, different lengths of time, and even on real-world data (like wind tunnel experiments and drug reactions) to see if the skills transferred. While the AI showed it could handle real data, the results were mixed; it was good at predicting the general "level" of an outcome but sometimes struggled to track the tiny, fast wiggles in the data. This suggests that while DoTime is a powerful new training ground, the AI is still learning and isn't perfect yet.
In short, DoTime is a massive, open-source playground that lets scientists build better "time-travel" detectives. It proves that teaching an AI to imagine different futures makes it smarter at understanding cause and effect, but it also shows us exactly where these digital brains still stumble. It's a tool for building better models, not a magic wand that fixes everything overnight.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.