BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali
This paper introduces BenHalluEval, the first comprehensive hallucination evaluation framework for Bengali large language models, which utilizes a dual-track protocol and a new calibration metric to assess performance across four tasks and reveal significant variations in hallucination detection capabilities.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a very smart, multilingual robot that can speak Bengali. You ask it questions, ask it to summarize news, or solve math problems. Sometimes, this robot is brilliant. Other times, it confidently makes things up—like saying a famous author's mother was named "Kathleen" when her real name was "Anne." In the world of AI, we call these confident lies hallucinations.
Until now, no one had built a proper "lie detector" specifically for Bengali. The authors of this paper, BenHalluEval, decided to fix that. Here is how they did it, explained simply:
1. The Big Problem: The "Silent Lie"
Most tests for AI only check if the robot gets the right answer. But what if the robot gets the right answer by accident, or lies so smoothly you don't notice?
- The Analogy: Imagine a student taking a test. If you only check the final grade, you might miss that they cheated on half the questions. You need to check how they answered, not just what they answered.
2. The Solution: A "Two-Track" Test
The researchers built a special exam with two separate lanes (tracks) to catch different types of mistakes:
- Track A (The Truth Check): They gave the robot questions with correct answers. They wanted to see if the robot would falsely accuse a correct answer of being a lie. (Did it cry wolf when there was no wolf?)
- Track B (The Lie Detector): They gave the robot questions with fake, made-up answers. They wanted to see if the robot could spot the lie. (Did it catch the wolf?)
The Score (BenHalluScore):
They created a new score called BenHalluScore. Think of it like a "Trust Meter."
- If a robot says "Yes, that's a lie!" to everything, it catches all the fake answers but also wrongly accuses all the real answers.
- If it says "No, that's fine" to everything, it misses all the lies but never accuses the truth.
- The Goal: A perfect robot needs to be smart enough to catch the lies without accusing the truth. The lower the score, the better the robot is at telling truth from fiction.
3. How They Built the Exam
They didn't just write questions; they built a massive factory to create "fake" answers.
- The Factory: They used a super-smart AI (GPT-5.4) to generate 12,000 fake answers across four different jobs:
- Answering Questions: (e.g., "Who is the mayor?")
- Code-Mixed Questions: (Mixing Bengali and English, like how people text on social media: "Kothay jabe?" instead of "কোথায় যাবে?")
- Summarizing: (Condensing a long article into a short paragraph).
- Reasoning: (Solving math problems step-by-step).
- The Types of Lies: They made the AI lie in specific ways, like changing a date, inventing a person's name, or doing math wrong.
4. The Results: Who Passed the Test?
They tested 7 different AI robots (some are huge and general, some are smaller and specialized in Bengali).
- The Winner: A small, specialized robot called TigerLLM-9B (built specifically for Bengali) did surprisingly well, beating much larger, general robots. It's like a local guide knowing the city better than a giant, expensive GPS that tries to know the whole world.
- The Losers: Some very famous, large robots failed badly. One robot (Mistral-nemo) was so confused it accused every single correct answer of being a lie. Another (LLaMA) was so lazy it said nothing was a lie, even when it was obvious.
- The Script Surprise: When the questions were written in "Banglish" (Bengali written with English letters), some robots actually performed better than when the questions were in standard Bengali script. It's like some robots feel more comfortable reading a text message than a formal letter.
5. The "Think Aloud" Trick Didn't Always Work
People often tell AI to "think step-by-step" (Chain-of-Thought) to stop it from lying. The researchers tried this trick.
- The Result: It didn't work consistently. Sometimes it helped, but often it just made the robot change its mind without actually getting better at spotting lies. It's like telling a liar to "think harder"—sometimes they just come up with a better lie, not the truth.
The Bottom Line
This paper is the first time anyone has built a dedicated "lie detector" for the Bengali language. They found that:
- Size isn't everything: A smaller, Bengali-focused robot can be more trustworthy than a giant, generic one.
- One-sided tests are dangerous: If you only test if a robot can find lies, you might miss that it also hates the truth. You need to test both.
- Low-resource languages need love: Just because a language has fewer digital resources doesn't mean we can't build better tools for it.
The authors have made their "exam questions" and "fake answers" available for everyone to use, so we can keep testing and improving these robots to make sure they tell the truth in Bengali.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.