HALAS: A Human-Annotated Dataset of Hallucinations of Modern ASR Systems
This paper introduces HALAS, the first human-annotated dataset of naturally occurring hallucinations in modern ASR systems derived from real earnings calls, which reveals that current detection methods underperform and establishes a rigorous benchmark for evaluating hallucination mitigation.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a very smart, high-tech scribe named "ASR" (Automatic Speech Recognition). This scribe is hired to listen to people talking and write down exactly what they say. Usually, this scribe is amazing. But sometimes, when the audio gets a bit tricky or the scribe is tired, it starts to hallucinate.
In this context, a hallucination isn't seeing ghosts; it's the scribe confidently writing down words that were never spoken. It might add a "thank you" when no one said it, repeat a sentence three times, or invent a whole new sentence that changes the meaning of the conversation.
Here is a simple breakdown of what the researchers in this paper did, using everyday analogies:
1. The Problem: The "Silent" Test
Previously, scientists tried to fix these hallucinations by testing the scribes on fake problems. They would feed the scribe static noise or silence and see if it made up words.
- The Analogy: It's like testing a chef's ability to cook a steak by asking them to cook a rock. If they fail, you know they can't cook rocks, but you don't know if they can handle a real, tough piece of meat.
- The Reality: The researchers realized that real-world speech (like people talking over each other in a busy office) is different from fake noise. They needed to see how the scribes behaved in the real world.
2. The Solution: HALAS (The "Hallucination Hall of Fame")
The team created a new dataset called HALAS. Think of this as a "Hall of Shame" or a "Hall of Fame" for mistakes, but specifically for things the scribes made up.
- How they built it: They took 119 hours of real recordings from corporate earnings calls (people talking about money and business).
- The "Tricky" Selection: They didn't just pick random clips. They looked for the moments where seven different top-tier scribes all gave different answers.
- Analogy: Imagine asking seven different people to describe a blurry photo. If they all say different things, that photo is likely very hard to see. The researchers picked those "blurry" moments because that's where the scribes were most likely to start making things up.
- The Human Touch: They hired 10 human experts to listen to the audio and the scribe's output. If the scribe wrote a word that wasn't in the audio, the human marked it as a "Hallucination."
3. What They Found: The "Ghost Words"
When they analyzed the data, they found some surprising patterns:
- Everyone Does It: All seven of the most advanced AI models they tested made these mistakes. No one was perfect.
- The "Favorite" Mistakes: The scribes didn't make random mistakes. They kept making the same mistakes over and over.
- Analogy: It's like a student who keeps forgetting to put a period at the end of a sentence. They might write "you," "okay," "thank you," or "yeah" even when no one said those words. About 55% of all the hallucinations were just these few common phrases.
- Low Error, High Risk: Sometimes the scribe got 95% of the words right (a low "Word Error Rate"), but the 5% it got wrong were total lies (hallucinations). This is dangerous because the text looks mostly correct, so you might not notice the lie.
- Severity: Some mistakes were minor (adding a "um" or "uh"), but others were severe (changing the meaning of a sentence or contradicting what was said).
4. The Test: Can We Catch the Liars?
The researchers used HALAS to test if current methods could catch these hallucinations.
- The "Proxy" Tests: They tried using simple math tools (like checking how long the text is or how similar it is to the original audio).
- Result: These tools were okay, but not great. They could catch the obvious, crazy mistakes, but they missed the subtle ones.
- The "AI vs. AI" Test: They tried using other AIs (like Large Language Models) to check the work.
- Result: Surprisingly, the "reference-free" methods (which look at the scribe's internal brain signals rather than comparing it to a "correct" answer) worked better than the ones that had the "correct answer" in front of them.
- The Score: The best detection method they found only got about 53% of the hallucinations right. This means that even with the best tools, we are still missing almost half of the lies the scribes tell.
The Bottom Line
This paper introduces HALAS, the first real-world "hallucination dataset" where humans have carefully marked exactly where AI speech-to-text systems lie.
They found that:
- Even the smartest AI scribes make up words in real conversations.
- They tend to make the same specific mistakes repeatedly.
- Current tools are not very good at catching these lies, especially when the rest of the text looks correct.
The paper concludes that HALAS is now the new "gold standard" for testing how well we can detect these AI hallucinations, moving us away from testing on fake noise and toward testing on real, messy human speech.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.