← Latest papers
💻 computer science

Extracting and Coding Digital Sustainability Data using Text Mining and AI: A Design Science Approach

This study employs a design science approach to develop and validate semi-automatic artifacts combining text mining, a specialized dictionary, and generative AI prompts, which significantly enhance the efficiency and scalability of extracting and coding digital sustainability data from textual sources while maintaining high agreement with expert manual coding.

Original authors: Thomas Abraham, Viet Dao, Nesreen El-Rayes

Published 2026-08-08
📖 5 min read🧠 Deep dive

Original authors: Thomas Abraham, Viet Dao, Nesreen El-Rayes

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a massive mystery: how do companies use technology to help the planet and people? The clues aren't hidden in a dusty attic; they are buried inside thousands of thick, boring reports that companies write every year. These reports are like giant libraries of text, filled with words about "green data centers," "smart sensors," and "digital equity." The problem is, reading every single page of every report by hand is like trying to drink the ocean with a straw—it takes forever, and you might miss the most important drops. This is the world of "Digital Sustainability," a field where researchers study how computers and software can make the world a better place. But to understand the big picture, they need to find and sort these clues quickly. That's where the magic of "text mining" (teaching computers to read and find patterns) and "Generative AI" (super-smart chatbots that can write and think) comes in. Instead of drowning in paperwork, researchers want a robot assistant to do the heavy lifting, so they can focus on solving the mystery.

This paper is about building that very robot assistant. The researchers, Thomas Abraham, Viet Dao, and Nesreen El-Rayes, decided to use a method called "Design Science," which is basically a fancy way of saying, "Let's build a tool to fix a problem, test it, and make it better." Their goal was to create a semi-automatic system that could scan sustainability reports, pull out the relevant digital sustainability stories, and then sort those stories into neat, logical categories. They didn't just guess; they built three specific "artifacts" (tools) to do the job: a special dictionary of words to look for, a set of instructions (prompts) to teach the AI how to think, and the whole process of how to use them together.

First, they tackled the "finding" part. Imagine you are looking for specific types of treasure in a cave. If you just shout "Treasure!" you'll get a lot of rocks and dirt. You need a specific map. The researchers built a "Digital Sustainability Dictionary," which is like a highly tuned metal detector. They fed it thousands of words from past research to teach it what "digital sustainability" actually sounds like. They tested this dictionary on reports from companies like Biogen and Coca-Cola. The results were impressive: the computer found 94% of the digital sustainability stories that human experts had found manually, but it did it in a fraction of the time. It even found 54 extra stories that the humans had missed! However, it wasn't perfect; the computer sometimes got distracted by words like "website" that sounded important but weren't actually about sustainability initiatives. This taught the researchers that while the computer is a great scanner, it still needs a human to double-check the results.

Next, they tackled the "sorting" part. Once the computer found the stories, it needed to put them in the right boxes. The researchers used two different "filing systems" based on established theories. One system sorted stories by what the technology does (like "Automating" a process or "Transforming" a business model), and the other sorted them by who benefits (like "Internal" vs. "External" or "Today" vs. "Tomorrow"). To teach the AI how to use these filing systems, they didn't just give it a simple command. They used a technique called "few-shot prompting," which is like showing the AI a few examples of a completed puzzle and saying, "See how this piece fits here? Now you do the rest." They trained the AI on data from three companies, letting it make mistakes, correcting it, and refining its instructions until it got the hang of it.

When they tested this trained AI on new companies (like Johnson & Johnson and Volvo), the results were strong. For the "what it does" categories, the AI matched the human experts' sorting 78% of the time. For the "who benefits" categories, it matched 85.3% of the time. The researchers calculated a score called "Cohen's Kappa," which measures how much two people (or a person and a robot) agree. The score was over 0.7, which is considered a high level of agreement. This suggests that the AI isn't just guessing; it's actually learning the logic of the researchers.

However, the paper is careful not to say this is a "magic wand" that solves everything. The researchers explicitly rule out the idea that this process can be fully automatic. They found that the AI sometimes needed help understanding the context of a business. For example, if a company talked about a safety system over several years, the AI might count it as multiple different projects, while a human would know it's just one long project. The dictionary also isn't perfect; it's a living thing that needs to be updated as new words enter the world. The authors conclude that the best approach is a partnership: the computer acts as a super-fast research assistant that does the heavy lifting of finding and sorting, but a human expert must always be in the loop to train the AI, fix its mistakes, and make the final call.

In the end, this paper suggests that we don't have to choose between human wisdom and machine speed. By combining text mining and Generative AI with a careful, step-by-step training process, researchers can build massive datasets of digital sustainability data much faster than before. It's not a solved problem, but it's a powerful new tool that makes the job of saving the world a little less lonely and a lot more efficient.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →