Earnings25: A Comprehensive 500-Hour Speech Benchmark for Finance
The paper introduces Earnings25, a comprehensive 500-hour benchmark comprising full earnings calls and industry-balanced segments with aligned transcripts and structured metadata, designed to evaluate and establish reproducible baselines for automatic speech recognition systems on English-language finance earnings calls under realistic conditions.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a robot to understand human conversation. You might start by having it listen to clear, slow stories read from a book. But real life is messy. People talk over each other, use slang, speak with different accents, and throw in complicated words that don't appear in storybooks. This is the world of Automatic Speech Recognition (ASR)—the technology that turns spoken words into text. While these robots are getting pretty good at listening to clear voices, they often stumble when the conversation gets chaotic, technical, or full of jargon. The big question for scientists is: How do we know if a robot is truly ready for the real world, or if it's just good at reading scripts? To answer this, we need a "final exam" that isn't just a test of memory, but a test of survival in a noisy, complex environment.
This is where a team from Bloomberg steps in with a new, massive challenge called Earnings25. Think of financial earnings calls as the "Olympics" of corporate speech. These are hour-long phone conferences where company bosses and financial analysts talk about money, stocks, and future plans. It's a perfect storm of difficulty: people speak fast, use heavy financial slang, talk over one another, and switch between reading prepared scripts and answering spontaneous, tricky questions. The authors realized that existing tests were like practicing for a marathon by running on a treadmill in a quiet gym; they didn't capture the hills, the wind, or the crowd. So, they built a new, super-detailed benchmark to see how well current speech-recognition robots can handle the real, messy chaos of the financial world.
The Great Financial Listening Challenge
The paper introduces Earnings25, a massive new dataset designed to be the ultimate stress test for speech-recognition AI. The authors didn't just grab a few random phone calls; they built two distinct "arenas" to test the robots in different ways.
The first arena is testset-full, a colossal collection of 498 hours of complete, unedited earnings calls from the S&P 500 companies in the fourth quarter of 2025. Imagine listening to nearly 500 hours of back-to-back financial meetings without skipping a beat. This set preserves the full, natural flow of conversation, including the awkward silences, the overlapping voices, and the long, winding stories that happen in real life. It's like giving the robot a whole season of a complex TV show to watch, rather than just a few clips.
The second arena is testset-segmented, a more curated set of 46 hours made up of 290 specific chunks of audio. Here, the researchers played a clever game of "fairness." In the real world, some industries (like banks or oil companies) talk way more often than others (like niche manufacturing). If you just listen to the most common ones, you might think the robot is great, but it might fail miserably when it hears about something rare. To fix this, the team picked exactly one segment per industry from over 2,000 calls, ensuring that every single type of business—from "3D Printers" to "Wind Turbines"—got an equal voice. These segments are short, about 5 to 10 minutes each, making them perfect for testing how well the robot handles specific, tricky vocabulary without getting overwhelmed by hours of content.
The Secret Sauce: Metadata and "Who Said What"
What makes Earnings25 truly special isn't just the audio; it's the "cheat sheet" that comes with it. Most old datasets just gave you the sound and the text. Earnings25 gives you the sound, the text, and a detailed map of who said what, when, and why.
The researchers tagged every speaker with their role: was it the operator reading the boring rules at the start? Was it the CEO giving a polished speech? Or was it an analyst asking a tough, unscripted question? They also labeled the industry and the structure of the call. This allows scientists to ask questions like, "Does the robot get confused more when the CEO is talking, or when the analyst is interrupting?" or "Does it fail more often in the biotech industry than in the banking industry?" It turns a simple "how many words did it get right?" test into a deep investigation of where and why the robot is struggling.
The Results: Robots Are Good, But Not Perfect
The team put two popular speech-recognition systems, Whisper and Parakeet-TDT, through the wringer. They didn't give the robots any special training on this specific data; they just asked them to listen and transcribe, simulating a real-world scenario where the robot has to figure things out on the fly.
The results showed that while these robots are impressive, they still have a long way to go. On the massive 498-hour set, the best model (Parakeet-TDT) got about 10.8% of the words wrong. On the smaller, industry-balanced set, the error rate jumped slightly to 11.1%. This small jump is actually a big deal. It suggests that when you force the robot to listen to rare, difficult industries (the "long tail" of business), it stumbles more often.
The paper highlights a fascinating discovery: aggregate scores can be misleading. If you just look at the average error rate, you might think the robot is doing great because it's good at the common industries like banking. But when you look at specific, complex fields like Biotech or Pharma, the error rate spikes to over 15%. It's like a student who gets an A on a math test because they aced the easy questions, but fails the advanced calculus section. The "average" grade hides the fact that they are still struggling with the hardest parts.
Why This Matters
The authors are careful not to claim they have "solved" the problem. Instead, they provide a reproducible benchmark—a standardized way for everyone to measure progress. Before Earnings25, it was hard to tell if a new speech AI was actually better or just lucky with the data it was tested on. Now, researchers have a clear, fair playing field with rich details to understand exactly where their models are failing.
The paper also points out its own limits. Earnings25 focuses on English-language calls, mostly from U.S.-based companies. While this makes the test controlled and high-quality, it doesn't yet capture the full diversity of global accents and languages. The authors suggest that the next step is to expand this "financial listening challenge" to include more countries and languages.
In short, Earnings25 is a giant, meticulously organized library of financial chaos. It doesn't just tell us if a robot can hear; it tells us if a robot can understand the messy, jargon-filled, overlapping, high-stakes world of business. And for now, the robots are listening, but they're still learning how to keep up with the conversation.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.