← Latest papers
💻 computer science

Evaluating and Mitigating Hallucinations in Retrieval-Augmented Question Answering over Uzbek School Textbooks

This study introduces the UzHistQA dataset and demonstrates that while retrieval-augmented generation significantly reduces hallucinations in Uzbek history question answering, enforcing citation and abstention mechanisms is necessary to completely eliminate fabricated answers, highlighting the need to separately evaluate retrieval quality and generation faithfulness in under-resourced languages.

Original authors: Ruzimboy Azimjonovich Kholmurotov

Published 2026-09-23
📖 5 min read🧠 Deep dive

Original authors: Ruzimboy Azimjonovich Kholmurotov

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the modern classroom, students increasingly turn to artificial intelligence to explain difficult concepts, summarize chapters, or prepare for tests. These digital assistants are built on large language models, systems trained on vast amounts of text that can generate fluent, confident, and often convincing answers. However, a distinct risk emerges when a student asks a question they do not yet know the answer to: they lack the knowledge to verify the reply. If the system invents a fact, the error goes unnoticed, and the student learns something false. To fix this, researchers have developed a method called retrieval-augmented generation. Instead of relying solely on the model's internal memory, the system first searches a specific, trusted document—like a school textbook—to find the relevant passage before answering. This keeps the response grounded in reality, but it does not guarantee perfection. The system might still fail to find the right passage, or it might find the right passage and then ignore it, or it might misinterpret what it found.

This uncertainty is particularly acute in Uzbek, a language spoken by millions but considered "low-resource" in the world of artificial intelligence. Unlike English, which dominates the data used to train these models, Uzbek has fewer digital texts available. Furthermore, Uzbek is an agglutinative language, meaning that a single root word can be modified with many different endings to create a vast array of surface forms. A student asking about "independence" might use a word form that looks completely different from the one used in the textbook, causing the search system to miss the answer entirely. Until now, there has been no standard way to test how well these systems work for Uzbek history or to measure how often they make things up.

A recent study by independent researcher Ruzimboy Kholmurotov addresses this gap by building a specific test set called UzHistQA, based on an official tenth-grade history textbook covering the years 1917 to 1991. The researcher analyzed the textbook and found a significant linguistic hurdle: out of nearly 31,000 words in the book, more than half of the unique word forms appeared only once. This extreme variety means that a simple search for exact word matches is likely to fail. To test how to overcome this, the study created 56 questions, including some that the textbook simply could not answer, to see if the system would admit its ignorance or invent a response. The researcher then compared four different ways of running the system: one that relied only on its internal memory, one that searched the textbook using a standard semantic search, one that combined that with a keyword search, and a final version that was strictly instructed to cite its source and refuse to answer if the evidence was missing.

The results were stark. When the system relied only on its internal memory without looking at the textbook, it made unsupported or invented claims in 91 percent of its responses. It confidently fabricated dates, names, and lists for questions the textbook did not cover. When the system was forced to look at the textbook first, the rate of these unsupported claims dropped dramatically to just 9 percent. This confirmed that grounding the answers in the actual text is essential for reliability. However, the study also revealed a surprising complication. The version that used the most sophisticated search method, which combined semantic understanding with keyword matching, actually found the correct passages more often than the simpler search. Yet, despite finding better evidence, this advanced system produced more unsupported claims than the simpler one. It appears that when the system retrieved a larger, more diverse set of passages, the extra information confused the model, leading it to make incorrect inferences or calculations even when the facts were right there.

The most effective configuration for safety was the one that enforced strict rules: the system had to cite the specific page for every claim and was instructed to say "I do not know" if the textbook did not contain the answer. This approach eliminated all fabricated answers for questions the textbook could not answer. The trade-off was that the system refused to answer 14 percent of the questions it could have answered, sometimes even when the answer was present in the retrieved text. The researcher concludes that for an educational setting, this refusal is a necessary feature. A student who receives a "I do not know" can ask a teacher for help, but a student who receives a confident, fabricated answer learns a falsehood and has no way to know it is wrong. The study establishes that while retrieval-augmented generation makes these systems viable for Uzbek education, it requires careful design to prevent the system from hallucinating, and it highlights that finding the right information and writing a correct answer are two separate challenges that must be measured independently.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →