Scope3Trace: Evidence-Based Identification and Extraction of Scope 3 GHG Emissions from Sustainability Reports
The paper introduces Scope3Trace, an evidence-grounded framework that combines document processing, LLM-assisted localization, and hybrid rule-LLM extraction to reliably and transparently identify and extract Scope 3 GHG emissions from heterogeneous sustainability reports, accompanied by a new dual-level multimodal dataset.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to solve a giant, global puzzle: figuring out exactly how much carbon pollution every single company on Earth is responsible for. This isn't just about the smoke coming out of their own chimneys; it's about the invisible cloud of pollution created by everything they buy, sell, and ship. In the world of climate science, this massive cloud is called "Scope 3" emissions. Think of Scope 1 and 2 as the smoke you can see right in front of you—like a factory's own smokestack or the electricity bill for its lights. But Scope 3 is the "ghost in the machine." It's the pollution from the steel used to build the factory, the trucks delivering the products, and even the waste customers throw away.
The problem is that while companies are getting better at reporting their own smoke, they are terrible at reporting this ghostly Scope 3 cloud. Their reports are messy, scattered across thousands of different PDF documents, written in confusing language, and often hidden inside complex tables. It's like trying to find a specific needle in a haystack made of other haystacks, where the needles are made of invisible ink. Scientists and regulators care deeply about this because you can't fix a problem if you can't measure it accurately. If we can't track these hidden emissions, we can't know if companies are actually helping the planet or just pretending to.
Enter Scope3Trace, a new digital detective team designed to hunt down these hidden numbers. The researchers built a smart system that acts like a super-powered librarian combined with a forensic accountant. Instead of hiring a team of humans to read millions of pages (which would take forever and cost a fortune), they created an automated pipeline that uses Artificial Intelligence (AI) to do the heavy lifting. But here is the clever part: they didn't just let the AI guess. They built a "rule-based" safety net. The AI reads the document, finds the numbers, and then a set of strict rules checks the math and the context to make sure the AI didn't get confused.
The system works in three main steps, like a well-oiled machine. First, it grabs the PDF reports and turns them into readable text, even if they are scanned images. Second, it uses a smart AI to find the exact pages and tables where the emissions are mentioned, ignoring the boring fluff. Third, it reconstructs the messy tables and pulls out the specific numbers for Scope 1, 2, and 3. Crucially, for every number it extracts, the system grabs a "receipt"—a tiny snippet of the original text proving where that number came from. This means you can always trace a number back to its source, just like checking a receipt at a store.
The results are impressive. When the researchers tested their system against a "gold standard" set of data that humans had carefully checked, Scope3Trace got it right almost every time, scoring a near-perfect 0.99 out of 1.0 on accuracy. This is significantly better than using just a raw AI model or simple rule-based tools alone. The team also used this system to build a massive new dataset called Scope3Trace, which includes over 54,000 records of building-level data and organization-level reports from countries like Australia and across Europe. They found that while companies report their total emissions, they rarely break it down by specific building. In fact, only about 3% of the building-level data in their new dataset was directly reported by the companies; the rest had to be carefully estimated using public energy data and smart math, but the system kept track of which numbers were "reported" and which were "estimated" so no one gets confused.
The paper also tested if this new data could be used to predict future emissions. They fed the data into a model that looked at building details and even satellite photos of the buildings. The model learned that combining the text data with satellite images gave the best predictions, suggesting that what a building looks like from space holds clues about how much pollution it creates. However, the authors are careful to note that this isn't a magic wand. The system works best where public data exists (like in Australia and parts of Europe) and struggles where data is missing. They explicitly state that their estimates are not a replacement for official, audited reports but are a transparent way to fill in the gaps so we can see the bigger picture.
In short, Scope3Trace doesn't solve the mystery of corporate pollution overnight, but it provides the best map we've ever had to navigate the fog. It proves that by combining smart AI with strict rules and clear evidence, we can turn a chaotic mountain of PDFs into a clear, trustworthy dataset. This allows researchers, investors, and policymakers to finally see the "ghost" of Scope 3 emissions, understand where it comes from, and start making decisions based on real evidence rather than guesswork.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.