← Latest papers
🤖 machine learning

Linear-LLM-SCM: Benchmarking LLMs for Coefficient Elicitation in Linear-Gaussian Causal Models

This paper introduces Linear-LLM-SCM, a framework for benchmarking large language models' ability to estimate quantitative causal coefficients in Linear-Gaussian structural causal models, revealing significant limitations in their accuracy and stability across real-world datasets.

Original authors: Kanta Yamaoka, Sumantrak Mukherjee, Thomas Gärtner, David Antony Selby, Stefan Konigorski, Eyke Hüllermeier, Viktor Bengs, Sebastian Josef Vollmer

Published 2026-07-29
📖 4 min read☕ Coffee break read

Original authors: Kanta Yamaoka, Sumantrak Mukherjee, Thomas Gärtner, David Antony Selby, Stefan Konigorski, Eyke Hüllermeier, Viktor Bengs, Sebastian Josef Vollmer

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a super-smart robot that has read almost every book, article, and website on the internet. You ask it a question, and it answers with the confidence of a seasoned professor. But here's the catch: while this robot is amazing at telling you that two things are connected (like "smoking causes coughs"), it often struggles to tell you how much they are connected (like "smoking increases coughing by exactly 15%"). This is the difference between qualitative reasoning (knowing the story) and quantitative reasoning (knowing the math). In the world of science, we use something called "causal models" to map out these stories. Think of a causal model as a flowchart or a recipe where arrows show how one ingredient changes another. Usually, scientists need to crunch numbers from real experiments to figure out the exact amounts in the recipe. But what if we could just ask our super-smart robot to guess those numbers based on what it has read? That is the big question researchers are asking today: Can an AI that knows everything about the world also do the math to predict the future accurately?

This paper, titled "Linear-LLM-SCM," sets up a playground to test exactly that. The researchers created a game where they gave seven different "Large Language Models" (the fancy name for these super-smart robots) a map of a real-world system—a Directed Acyclic Graph, or DAG for short. You can think of a DAG as a one-way street map of cause and effect, where you can't drive in circles. The map showed which variables (like "Glucose levels" or "Spending money") influenced others. The robots' job was to act like expert statisticians and write down the exact mathematical equations that describe these relationships, specifically guessing the "coefficients" (the numbers that tell you how strong the effect is).

The team tested seven different real-world maps, ranging from how people spend money to how plants grow in algae. They asked the robots to look at the map and the descriptions of the variables, then spit out a regression equation (a math formula) for each part of the map. They then compared the robots' guesses against the "ground truth"—the actual, correct numbers that scientists had already measured in the real world.

The results were a mix of "not bad" and "oops." The researchers found that while the robots could sometimes guess the right direction (positive or negative effect), the actual numbers were all over the place. Even when the researchers told the robots to be as consistent as possible (by setting a "temperature" to zero, which is like telling the robot to stop guessing and just calculate), the answers still varied wildly from one try to the next. For example, on the "Cachexia" map (related to muscle wasting), one robot guessed a coefficient of 1.07, while another guessed 1.04, but the variation was high enough to be worrying. On the "Algal" map, a smaller robot model (Llama 3.1 8B) completely failed to write a readable equation at all.

The paper also put the robots through a stress test to see how tough they were. First, they changed the units of measurement (like switching from micrometers to nanometers). Surprisingly, sometimes this made the robots' answers look better by accident, likely because the numbers became easier to write down, not because the robots understood the physics better. Second, they tricked the robots by adding fake connections to the map (spurious edges). When the map had a fake arrow pointing from one thing to another that didn't actually affect it, the robots' performance dropped. They got confused, and their ability to rank the importance of different factors got worse.

The main takeaway is that while these AI models are great at understanding the story of cause and effect, they are currently unreliable at doing the math to predict the exact size of those effects. The study suggests that using these robots for high-stakes decisions, like in healthcare or clinical settings, is risky right now because their numbers can be inconsistent and sensitive to small changes in how the problem is presented. The authors aren't saying the robots are useless; they are saying we need to be very careful and not trust them with the final numbers until they get much more stable. They have even released their testing framework to the public so other scientists can keep trying to fix this problem, hoping that one day, our digital assistants will be able to not just tell us the story, but also do the math perfectly.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →