Diagnosing Causal Credibility of Machine Learning Explanations: A Dual- SHAP Attribution Framework Applied to Social Survey Data
This paper introduces a dual-SHAP framework that distinguishes between correlational and causal relationships in machine learning explanations by comparing conditional and interventional attributions, demonstrating its effectiveness in diagnosing causal credibility across diverse social survey outcomes.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the modern study of human behavior, researchers increasingly rely on powerful computer programs to find patterns in vast amounts of survey data. These machine learning models are excellent at spotting complex connections that traditional statistics might miss, such as how a person's income, age, and where they live might combine to influence their social habits. However, these models are often opaque; they can predict an outcome with high accuracy but offer no clear explanation of why. To fix this, scientists use a tool called SHAP, which acts like a spotlight, highlighting which factors a model considered most important when making a prediction. The problem is that this spotlight often confuses two very different things: a genuine cause-and-effect relationship and a mere coincidence. Just because two things happen together does not mean one caused the other. For instance, a model might flag "spouse's age" as a key reason for a person's internet usage, not because the spouse's age directly changes how much the person goes online, but simply because spouses tend to be similar ages, and age itself is a strong driver of technology use. Distinguishing between these real causes and misleading correlations is essential for social scientists who want to understand human behavior rather than just predict it.
A new study by Shidong Li from the North University of China tackles this confusion head-on by applying a rigorous diagnostic framework to real-world social survey data. The researcher analyzed responses from the 2023 Chinese General Social Survey, a massive dataset representing nearly 4,800 people across thirty-one provinces. The goal was to test eighteen different aspects of social life, ranging from how often people visit friends to their trust in strangers and their media consumption habits. Instead of accepting the standard explanations provided by a single computer model, the study employed a "dual-SHAP" approach. This method compares two different ways of calculating importance. The first method looks at the data as it naturally exists, preserving all the existing links between variables. The second method breaks those links artificially, asking what would happen to a prediction if a specific factor changed while everything else remained independent. By comparing the results of these two methods, the study could measure a "credibility score" for each factor. A high score means the factor's importance holds up even when correlations are broken, suggesting a direct causal link. A low score indicates the factor was likely just riding along on the coattails of another, more influential variable.
The investigation began by scanning the eighteen social behaviors to see which ones could be reliably predicted by the available data. Out of the twenty-two behaviors originally considered, eighteen were found to be predictable enough to analyze further, while four, such as the frequency of volunteer activities, were too erratic to model effectively. Among the eighteen that remained, the study tested five different types of machine learning algorithms to see which one performed best for each specific behavior. The results challenged a common assumption in the field: that one specific algorithm, known as XGBoost, is the universal champion for all tasks. Instead, the study found that no single model dominated the entire landscape. A model called RandomForest was the best performer for ten of the eighteen outcomes, covering areas like social interaction and media usage. However, other models, including Ridge regression, LightGBM, and GradientBoosting, proved superior for specific outcomes like general social trust or newspaper reading. This finding suggests that the choice of computer model is not a minor technical detail but a substantive decision that fundamentally shapes the conclusions researchers draw about what drives human behavior.
Once the best model was selected for each behavior, the study narrowed down the list of potential causes from thirty-eight predictors to twenty-nine key factors. This selection process used the SHAP tool itself to filter out variables that contributed very little to the overall picture, such as the number of daughters a person has or specific attitudes toward gender roles. The remaining twenty-nine factors included demographics like age and income, as well as attitudes toward social fairness and language abilities. Age emerged as the most dominant predictor across the board, followed closely by personal income and language skills. With these key factors identified, the researchers applied the dual-SHAP framework to calculate the credibility of each factor's influence. They found that the average credibility score across all the relationships was 0.727, indicating that while many factors do have a direct influence, a significant portion of the importance assigned by standard models is inflated by correlations.
The study revealed a wide spectrum of reliability. Some factors, such as a person's perception of social fairness or their political views on migration, showed high credibility scores, meaning their influence on social outcomes was robust and likely causal. In contrast, factors like a spouse's age or the duration of a marriage often showed low or even negative credibility when their connection to the respondent's own age was broken. For example, when the study simulated a scenario where a spouse's age was changed independently of the respondent's age, the apparent influence of the spouse's age on internet usage vanished or even reversed direction. This served as a clear warning that the original model had mistaken a correlation for a cause. The researchers established a threshold to help interpret these scores: they suggested that a score of 0.80 or higher should be considered a strong indicator of a primarily causal relationship. Using this standard, only about forty percent of the factor-behavior pairs qualified as primarily causal. If the threshold was lowered to a more lenient 0.50, nearly ninety-three percent of the pairs met the criteria, but the researchers argued that the stricter standard is necessary to avoid being misled by coincidental patterns.
To ensure these findings were not just statistical artifacts, the study performed a final check by simulating real-world changes. They asked what would happen to a person's predicted behavior if a specific factor, like income or age, increased by a standard amount. The results were consistent with common sense and established social science. For instance, increasing a person's age in the simulation led to a predicted decrease in internet usage, reflecting the well-known digital divide where older adults use technology less frequently. Similarly, increasing personal income led to a predicted increase in newspaper reading, aligning with the idea that higher socioeconomic status is linked to print media consumption. These simulations confirmed that the models were capturing genuine, interpretable relationships rather than random noise. Furthermore, when the researchers tested the framework on different subgroups, such as younger versus older adults or urban versus rural residents, the overall patterns held steady, though specific credibility scores varied slightly depending on the population. This demonstrated that while the method is robust, the strength of a causal link can depend on the specific group being studied.
Ultimately, this work provides a transparent and reproducible tool for social scientists to separate signal from noise in machine learning explanations. It does not claim to prove causality in the absolute sense, as unmeasured factors could still be at play, but it offers a practical way to diagnose when an explanation is likely to be misleading. By comparing how a model behaves when it sees data as it is versus how it behaves when it breaks the natural links between variables, researchers can now identify which factors are truly driving social outcomes and which are merely shadows cast by other variables. The study concludes that while machine learning is a powerful lens for viewing human behavior, it requires this kind of careful calibration to ensure that the insights it provides are not just accurate predictions, but trustworthy explanations of why people act the way they do.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.