HiFi-KPI: A Dataset for Hierarchical KPI Extraction from Earnings Filings
This paper introduces HiFi-KPI, a large-scale dataset containing 1.65 million paragraphs and 198,000 hierarchically organized labels linked to iXBRL taxonomies to facilitate the extraction and classification of Key Performance Indicators from earnings filings, alongside a curated Lite subset for rapid evaluation and open-source code.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to read a financial report from a giant corporation like Apple or Tesla. These reports are like massive, dense forests of numbers and text. To make sense of them, computers need a map.
For years, the financial world has used a digital map called iXBRL. It's like a giant filing cabinet where every single number (like "Revenue" or "Profit") is tagged with a specific label. But here's the problem: the labels are way too specific.
Think of it like this:
- The Old Way: Instead of a label saying "Fruit," the computer sees "Red Apple, 2023, Washington State, Gala Variety, 150g." If you want to compare "Apples" across different companies, the computer gets confused because one company calls it "Gala Apple" and another calls it "Red Delicious." It's like trying to compare two libraries where one books are sorted by "Author's Middle Name" and the other by "Book Cover Color." It's a mess.
- The Result: Computers are great at reading, but they are terrible at understanding the context (like what currency the money is in, or what dates the numbers cover) when the labels are this hyper-specific.
Enter: HiFi-KPI (The "High-Fidelity" Map)
The authors of this paper built a new tool called HiFi-KPI. Think of it as a universal translator and a smart librarian rolled into one.
Here is what they did, broken down simply:
1. The Massive Collection (The Library)
They scraped 1.65 million paragraphs from thousands of official financial reports (like the 10-K and 10-Q forms companies file with the government).
- The Analogy: Imagine they read every single page of every financial report filed in the US for the last seven years. That's a lot of reading!
2. The "Smart Hierarchy" (The Organizing System)
The old system had 198,000 different, confusing labels. HiFi-KPI organized these into a hierarchy (a family tree).
- The Analogy: Instead of forcing the computer to memorize 198,000 specific names, they taught it to understand the family.
- Grandparent: "Money Made" (Revenue)
- Parent: "Money from Selling Stuff"
- Child: "Money from Renting out a specific building in 2023"
- This allows the computer to zoom in or zoom out. If it can't find the exact "2023 Building Rent," it can still confidently say, "Ah, this is just Revenue." This makes the system much smarter and more flexible.
3. The "Context Clues" (The Detective Work)
A number like "$50 million" is useless without context. Is it US Dollars? Euros? Is it for last year or next year?
- The Analogy: HiFi-KPI doesn't just grab the number; it grabs the entire scene. It links the "$50 million" to the specific dates (Jan 1 to Dec 31) and the currency (USD). It's like a detective who doesn't just find a clue but also writes down where they found it and when.
4. The "Lite" Version (The Training Wheels)
The full dataset is huge and complex. So, the authors also created HiFi-KPI-Lite.
- The Analogy: This is like a "Study Guide" or a "Flashcard set." They asked a human financial expert to pick the 4 most important things to look for: Earnings, EBIT (profit before interest), EPS (earnings per share), and Revenue.
- They mapped thousands of confusing computer tags to these 4 simple human concepts. This allows researchers to test new AI models quickly without needing a supercomputer.
What Did They Test? (The Race)
The authors put different types of AI to the test to see who could read these reports best:
- The "Old School" AI (Encoder Models): These are like diligent students who have read the textbook many times.
- Result: They were excellent at identifying the labels (getting 90%+ accuracy). They are great at saying, "This sentence is about Revenue."
- The "New Genie" AI (Large Language Models like GPT, Qwen, etc.): These are the flashy, creative AIs that can write stories and chat.
- Result: They were okay at the basics but struggled with the details. They often got the dates wrong or mixed up currencies.
- The Problem: If you ask a Genie to extract a specific number with its date and currency, it sometimes guesses. It's like a smart friend who knows the general idea but might get the specific time of your meeting wrong.
Why Does This Matter?
- For Investors: Imagine being able to instantly compare the "Revenue" of 500 companies without manually reading 500 PDFs. This tool automates that, helping investors spot trends faster.
- For Companies: It helps them check their own reports for errors before they are filed.
- For Regulators: It makes it easier to spot if a company is hiding something or tagging things incorrectly.
The Bottom Line
The paper says: "We built a massive, organized library of financial data that teaches computers how to understand the context of numbers, not just the numbers themselves."
While the newest, flashiest AI models (LLMs) are impressive, the paper shows that for this specific, high-stakes job, specialized, fine-tuned models (the "diligent students") are currently more reliable than the "flashy genies." However, this new dataset gives everyone a better foundation to build even smarter tools in the future.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.