← Latest papers
🧬 biology

An Explainable XGBoost Framework for Anxiety and Depression Risk Prediction with SHAP and Counterfactual Reasoning

This study presents an explainable XGBoost framework for predicting anxiety and depression risk that addresses class imbalance and evaluation rigor while integrating SHAP and counterfactual reasoning to provide transparent, interpretable insights into model decisions without replacing clinical diagnosis.

Original authors: Anima Dahal, Bibash Basnet

Published 2026-09-25
📖 5 min read🧠 Deep dive

Original authors: Anima Dahal, Bibash Basnet

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

In the modern landscape of healthcare, artificial intelligence has emerged as a powerful tool for sifting through vast amounts of human data to find patterns that might otherwise remain hidden. When applied to mental health, these systems aim to look at the answers people give to questions about their lives, their habits, and their feelings to spot early signs of anxiety or depression. The challenge, however, is that these conditions are complex and often subtle, and the data used to study them is frequently unbalanced, with far more people reporting good mental health than those reporting struggles. Furthermore, a computer model that makes a prediction is often a black box; it gives an answer without explaining why, which makes it difficult for doctors to trust or use that information responsibly. To be truly useful, a system must not only predict risk accurately but also show its work, revealing which factors drove the decision and ensuring it treats different groups of people fairly.

A team of researchers at Nilai University has developed a new framework designed to meet these exact challenges. They built a sophisticated computer program capable of assessing the risk of anxiety and depression using a large collection of survey responses and personal details. Rather than simply training a model to guess correctly as often as possible, the researchers focused on creating a system that is transparent, reliable, and carefully tested. They started with a dataset containing information from more than 24,000 participants, including their answers to standard questionnaires about stress, sleep, and mood, alongside details like age, gender, and education. Because the number of people with anxiety or depression in the group was much smaller than those without, the team had to teach the computer to pay extra attention to the minority cases, ensuring it did not simply ignore them to achieve a high overall score.

The researchers split their data into three distinct groups to ensure their results were honest and not just a lucky guess. They used one group to teach the computer, a second group to fine-tune how the system makes decisions, and a third, completely isolated group of 3,644 people to test the final result. This strict separation prevented the system from memorizing the test answers, a common pitfall that leads to overly optimistic results. They also adjusted the system's decision-making process. Instead of using a standard rule where a score above fifty percent counts as a positive result, they raised the bar significantly. They found that a much higher threshold was necessary to make the predictions reliable, effectively telling the system to be very sure before flagging someone as being at risk.

When the final test was run on the isolated group of 3,644 participants, the system performed with remarkable precision. For anxiety, it achieved an AUC of 0.98, and for depression, it achieved an AUC of 0.95. More importantly, the system did not just give a yes or no answer; it provided a clear explanation for its choices. Using a method that breaks down the influence of each question, the researchers found that answers related to stress were the most powerful drivers for both anxiety and depression predictions. Sleep-related questions also played a major role, particularly for depression. The system could show exactly how a person's specific answers pushed the final score up or down, allowing a clinician to see the logic behind the risk assessment.

To further test the system's logic, the researchers asked a different kind of question: what would have to change for the computer to change its mind? They created hypothetical scenarios where they slightly altered a person's answers to see if the risk score would drop. They discovered that changing just a few key answers, such as reducing reported stress levels or improving sleep quality in the data, was often enough to move a person from a high-risk category to a lower one. This demonstrated that the model was sensitive to specific changes and not just making random guesses. The team also checked to ensure the system worked equally well for men and women, finding that while the exact numbers varied slightly, the ability to distinguish between high and low risk remained consistent across genders.

The researchers emphasize that this system is not a replacement for a doctor's diagnosis. Mental health involves deep personal and contextual factors that a computer cannot fully capture. Instead, this framework is designed as a decision-support tool, a way to help practitioners identify individuals who might benefit from a closer look. By combining high accuracy with clear explanations and fair testing across different groups, the study offers a more trustworthy path forward for using artificial intelligence in mental health screening. It shows that when technology is built with care and transparency, it can illuminate the path to better care without losing sight of the human complexity at its core.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →