IndicContextEval: A Benchmark for Evaluating Context Utilisation in Audio Large Language Models Across 8 Indic Languages
This paper introduces IndicContextEval, a comprehensive multilingual benchmark spanning 8 Indian languages and 23 domains, designed to rigorously evaluate whether Audio Large Language Models genuinely utilize explicit contextual prompts or rely on pretraining knowledge through a novel 7-level prompting framework.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a very smart, multilingual assistant who can listen to people speaking and write down exactly what they say. This is what AudioLLMs (Audio Large Language Models) do. They are like super-powered dictation tools that can also listen to a little note you give them beforehand, like "This is a medical meeting" or "Here is a list of names to watch for."
The big question the authors of this paper asked is: Does this assistant actually read your note and use it to help, or does it just pretend to listen while relying on what it already memorized during its training?
To find the answer, they built a new testing ground called IndicContextEval. Here is a simple breakdown of what they did and what they found, using some everyday analogies.
1. The Test Kitchen: IndicContextEval
Think of this benchmark as a massive, organized kitchen where they cooked up 56 hours of natural speech from 555 different people speaking 8 different Indian languages (like Hindi, Bengali, and Tamil).
- The Ingredients: They didn't just record random chatter. They focused on 23 different professional fields, ranging from "Robotics and Automation" to "Culinary Arts" and "Law." This ensures the speakers use tricky, technical words that are hard to guess just by hearing them.
- The Goal: They wanted to see if the AI could correctly transcribe these tricky words when given a "hint" (context) versus when it had to guess on its own.
2. The Seven Levels of Hints (The "Prompt" Ladder)
The researchers designed a clever ladder of 7 levels to test the AI. Imagine you are asking a friend to write down a story about a specific topic. You give them hints that get more and more specific:
- Level 0 (The Blank Slate): "Just write down what you hear." (No hints at all).
- Level 1 (The Language Hint): "Write down what you hear in Hindi." (Just tells the language).
- Level 2 (The Structured Note): "This is a medical talk in Hindi, spoken by a doctor." (Facts only).
- Level 3 (The Story Note): "This is a doctor explaining a surgery in Hindi." (A natural sentence description).
- Level 4 (The English List): "Here is a list of medical terms in English: Heart, Scalpel, Antibiotic." (But the AI must still write the answer in Hindi).
- Level 5 (The Native List): "Here is the same list, but written in Hindi script." (The hint matches the answer language).
- Level 6 (The Trap): "Here is a list of cooking terms (like Flour, Oven, Spice) for a medical talk." (This is a trick to see if the AI blindly follows the list even when it's wrong).
3. The Results: Who Actually Listens?
They tested five different AI models. Here is what happened, using the "Friend" analogy:
The "Smart & Skeptical" Friend (GPT-4o Transcribe):
This friend is good. When you give them the right list of words (Level 5), they do a great job. But when you give them the wrong list (Level 6, the cooking trap), they ignore it and stick to what they actually hear. They know when to use the hint and when to ignore it.The "Eager Learner" (Gemini 3 Flash):
This friend loves the hints. When you give them the right list, they improve their writing significantly. They seem to genuinely use the context to get the tricky words right.The "Blind Follower" (Gemma-3N):
This friend is too trusting. When you give them the wrong list (Level 6), they get confused and start writing nonsense based on that wrong list. They rely too much on the note and forget to listen to the actual voice.The "Deaf to Notes" Friend (Sarvam Audio):
This friend is actually the best at just listening and writing down words (lowest errors overall). However, they barely change their behavior whether you give them a hint or not. It's like they are so good at hearing that they don't really need the notes, or they simply ignore them.The "Confused" Friend (GPT-4o and others at Level 0):
When no language was specified (Level 0), many models got confused about which language script to use, making lots of mistakes. Once you told them "Speak Hindi," their performance jumped up immediately.
4. Key Takeaways
- The "Hint" Format Matters: Telling the AI "Here is a list of words in English" (Level 4) wasn't as helpful as giving the list in the local language script (Level 5). It's like trying to read a recipe written in a language you don't speak; it doesn't help as much as having it in your own language.
- Natural vs. Robot: Describing the audio in a natural sentence (Level 3) worked better than giving a dry list of facts (Level 2).
- The Trap Test: The "Level 6" trap was crucial. It proved that some models are smart enough to ignore bad hints, while others are so eager to please that they follow bad instructions and make mistakes.
Summary
The paper concludes that while these AI models are powerful, they don't all use context in the same way. Some are smart enough to know when to use a hint and when to ignore it. Others are either too blind to the hints or too easily tricked by them.
The authors built this test (IndicContextEval) to help developers understand these differences so they can build better, more reliable voice assistants for Indian languages in the future. They made all their data and tests public so others can continue this research.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.