← Latest papers
📈 economics

Transparent Environment, Social, and Governance (ESG) Risk Score Prediction Using Machine Learning and Explainable AI

This study demonstrates that an explainable machine learning pipeline, specifically utilizing XGBoost combined with SHAP and LIME techniques on public data, can effectively predict transparent and reproducible Environmental, Social, and Governance (ESG) risk scores, offering a viable alternative to proprietary rating systems.

Original authors: Avuzwa Lerotholi, Ibidun Christiana Obagbuwa, Olaperi Okuboyejo

Published 2026-08-26
📖 5 min read🧠 Deep dive

Original authors: Avuzwa Lerotholi, Ibidun Christiana Obagbuwa, Olaperi Okuboyejo

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the modern financial world, a company's value is no longer measured solely by its profits. Investors, regulators, and the public increasingly demand to know how a business treats the planet, its workers, and its own leadership. This broader assessment is known by the acronym ESG, standing for Environmental, Social, and Governance. The environmental pillar looks at a company's ecological footprint, such as its carbon emissions and waste. The social pillar examines how the organization interacts with society, including its treatment of employees and local communities. The governance pillar focuses on how the company is run, considering its leadership structure and board diversity. While these ratings are meant to guide ethical investment, the methods used to calculate them are often secret. Major rating agencies keep their formulas private, treating their scores as proprietary products. This lack of transparency means that even the people relying on these scores cannot easily see how the numbers were derived, turning the ratings into black boxes that are difficult to trust or verify.

A team of researchers from South Africa set out to open these black boxes. They wanted to see if they could build a clear, transparent system to predict these ESG risk scores using only public information. Instead of relying on expensive, private databases, the team gathered data from a free financial portal that aggregates information from a service called Sustainalytics. They collected information on thousands of companies across twenty-four different countries, covering a wide range of industries from technology to mining. Their goal was to test two different types of computer models. The first type, known as white-box models, are simple and easy to understand, like a straightforward list of rules. The second type, called black-box models, are far more complex and powerful but usually impossible to interpret because they work like a tangled web of calculations. The researchers trained these models to guess the ESG risk scores based on financial data and company characteristics, then used special tools to force the complex models to explain their own reasoning.

The results showed that the more complex models were significantly better at making accurate predictions. The most successful tool was a sophisticated algorithm called XGBoost, which learned to spot subtle patterns in the data that simpler models missed. For the environmental pillar, this model was able to explain 92 percent of the variations in the scores, and for the governance pillar, it explained 85 percent. The social pillar proved the most difficult to predict, with the model explaining 78 percent of the variations, suggesting that social risks are influenced by a wider and more chaotic set of factors. The researchers found that the simpler, transparent models could not match this level of accuracy. While they were easier to read, they failed to capture the intricate relationships between a company's finances and its sustainability performance. This suggests that to truly understand ESG risks, one must use powerful tools that can handle complex, non-linear connections, even if those tools require extra effort to interpret.

To solve the problem of these powerful models being too complex to understand, the researchers applied a technique called Explainable AI. This method acts like a spotlight, illuminating exactly which factors drove a specific prediction. When they turned this spotlight on their best-performing model, a clear picture emerged. The most influential factor for predicting a company's risk was not a single internal metric, but rather the average risk level of its industry peers. In other words, a company's score was heavily influenced by the typical behavior of other companies in the same sector. If a company operated in an industry where high environmental risk was the norm, the model predicted a higher risk for that company as well. Financial health also played a major role; companies with strong profits and stable debt structures tended to have lower governance and social risk scores. The researchers also found that specific controversies, such as scandals involving labor or safety, pushed risk scores sharply upward.

The study further revealed that these predictions were not equally accurate for every type of business. The models worked exceptionally well for companies in the finance and consumer sectors, where data was consistent and patterns were clear. However, the models struggled more with companies in the energy, materials, and industrial transportation sectors. These industries showed higher errors in prediction, likely because their ESG data is more complex, inconsistent, or difficult to standardize across different countries. The researchers also looked at specific companies in South Africa to see how the model behaved in real-world scenarios. They found that a financial company with the lowest environmental risk score was correctly identified as such because the sector naturally has a small physical footprint. Conversely, a company in the energy sector with the highest risk score was flagged due to its heavy reliance on resource-intensive operations. Interestingly, the model also identified a financial company with a surprisingly high governance risk, suggesting that even in highly regulated sectors, specific internal irregularities can drive up risk scores.

Ultimately, this research demonstrates that it is possible to create a transparent and reproducible system for assessing corporate sustainability without relying on secret, proprietary formulas. By using publicly available data and advanced machine learning, the researchers showed that they could predict ESG risk scores with high accuracy. They proved that while complex models are necessary to capture the full picture of corporate risk, tools exist to make their decisions understandable. The findings suggest that a company's ESG standing is deeply tied to the norms of its industry and its own financial stability. For investors and regulators, this work offers a new way to look at corporate responsibility, moving away from opaque ratings toward a system where the reasons behind a score are clear, logical, and open to scrutiny.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →