From Causal Plausibility to Causal Reliability: Evaluating LLMs as Calibrated Direct Causal-Edge Classifiers
This paper systematically evaluates 12 large language models as direct causal-edge classifiers and finds that while they exhibit strong recall, they suffer from significant overprediction, poor distinction between direct and indirect relationships, and severe overconfidence in incorrect judgments, suggesting they are better suited as sources of soft causal priors rather than reliable direct evidence of causal structure.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Science often seeks to understand not just what happens, but why it happens. When researchers try to map out the causes behind a phenomenon—whether it is a virus spreading through a population or a machine failing in a factory—they build a diagram of connections. In this diagram, a direct line from one thing to another means the first thing immediately causes the second. If a virus leads to an infection, which then leads to a positive test result, the virus causes the test result only indirectly, through the infection. Figuring out which lines are direct and which are indirect is a difficult puzzle known as causal discovery. For decades, scientists have relied on complex mathematical tools and large sets of observations to solve this puzzle, but these tools sometimes struggle when data is scarce or when different patterns look identical.
In recent years, a new kind of tool has emerged: large language models. These are computer systems trained on vast amounts of text that can answer questions, write stories, and reason through problems. Because they have read so much, they seem to hold a deep well of knowledge about how the world works, including how things cause one another. Researchers have begun to ask if these systems can act as expert guides, telling us which connections in a diagram are real and direct. The hope was that these models could provide a shortcut, offering reliable judgments about cause and effect without needing to run expensive experiments. However, a new study suggests that while these models are good at sensing that two things are related, they are often terrible at knowing exactly how they are connected, and they are frequently far too confident when they get it wrong.
A team of researchers set out to test this reliability by treating large language models as judges of direct cause-and-effect relationships. They gathered twelve different versions of these models, ranging from smaller, more efficient systems to massive, powerful ones. They then presented these models with six different real-world scenarios, each represented by a known map of causes and effects. These scenarios covered diverse fields, from medical conditions like hepatitis and COVID-19 to industrial processes and ecological systems. For every pair of variables in these maps, the researchers asked the models a simple question: Does the first item directly cause the second? The models had to answer yes or no and provide a confidence score, a number indicating how sure they felt about their answer.
The results revealed a consistent pattern of over-enthusiasm. The models tended to say "yes" far too often. Instead of drawing a sparse map with only the necessary direct connections, the models produced dense webs where almost everything seemed to be directly connected to everything else. This behavior was so strong that even when the researchers tried different ways of asking the questions, the models could not be easily trained to be more selective. They would correctly identify that two things were related, but they would fail to distinguish between a direct cause and an indirect one. For instance, if a virus caused an infection which caused a fever, the model would often claim the virus directly caused the fever, ignoring the middle step. This error happened in about 40 percent of cases where the relationship was indirect, and in 36 percent of cases where the direction was reversed.
Perhaps the most troubling finding was the models' lack of self-awareness regarding their mistakes. When the models made these structural errors, they did not hesitate. In more than 80 percent of the cases where they incorrectly claimed a direct link existed between indirectly related items, they assigned a confidence score of at least 80 percent. They were not just guessing; they were confidently wrong. The study also looked at whether the models could be trusted to tell the difference between a correct guess and a wrong one based on their confidence scores. The researchers found that the confidence numbers the models gave themselves were often unreliable. A model might say it was 90 percent sure, but that number did not actually correlate with whether the answer was right. The only method that showed a slight improvement was asking the models the same question in different ways and seeing if they agreed with each other, but even this method was not a perfect solution.
The researchers also investigated whether the models were simply memorizing the answers from the test questions they had seen before. They found that for one specific dataset involving a medical condition, several models performed suspiciously well, suggesting they might have seen the data during their training. However, this familiarity did not explain the poor performance on the other five datasets, where the models still struggled with the same types of errors. Even when the researchers forced the models to choose between three options—direct cause, reverse cause, or no direct link—instead of just two, the models still failed to improve significantly. They continued to predict direct links where none existed, proving that the problem was not just a flaw in how the questions were asked, but a fundamental limitation in how these systems understand causality.
Ultimately, the study concludes that while large language models are powerful tools for generating ideas, they should not be treated as definitive proof of how the world works. They are better viewed as sources of soft suggestions, offering a starting point that must be checked and validated by other methods. They can tell us that two things are likely connected, but they cannot be trusted to draw the precise lines of a causal map on their own. The confidence they express is often a reflection of their training data rather than a true measure of truth. For scientists and engineers who need to understand the precise mechanics of cause and effect, relying solely on these models would be like navigating a complex city with a map that connects every street to every other street; it might look comprehensive, but it would lead you astray. The path forward lies in using these models to generate hypotheses that are then rigorously tested, rather than accepting their confident judgments as final answers.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.