Khondo: A Multimodal Benchmark for Document Packet Splitting of Bangla Forms
The paper introduces Khondo, the first vision-native, bilingual benchmark for splitting and reordering concatenated Bangla government forms, revealing that while multimodal large language models can effectively cluster pages by document, they struggle significantly with reconstructing the original page order, particularly in low-resource languages.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a librarian in a chaotic, magical library where books don't just sit on shelves; they sometimes get glued together into giant, messy scrolls, and then someone shakes the table, mixing the pages of different stories until they are completely jumbled. Your job is to look at this tangled mess, figure out which pages belong to which story, and then re-stack them in the correct order so the story makes sense again. This is the real-world problem of "document packet splitting." In the digital world, governments and offices often scan hundreds of papers into one big file. Sometimes, the scanner gets confused, or a worker accidentally shuffles the pages, mixing a birth certificate with a tax form or a license with a medical report. For computers, this is a nightmare. While we have built smart AI that can read text, teaching it to look at a picture of a page, understand the visual clues, and then untangle a shuffled pile of documents is a much harder challenge, especially when the documents are written in languages that computers don't see very often.
This is where a new project called Khondo (which means "split" or "segment" in Bangla) comes in. The researchers behind this project created a special test set, like a giant puzzle box, specifically designed to see how well modern AI can untangle these messy document piles. They didn't just use English; they used real government forms from Bangladesh, which are often written in Bangla, English, or a mix of both. They took thousands of pages, glued them together into fake "packets," and then deliberately shuffled the pages in five different ways: from keeping them in order, to mixing them up within a single document, to completely scrambling pages from different documents together. They then asked the world's most advanced AI models to look at the pictures of these pages and try to sort them out.
The results were a mix of "not bad" and "oh no." The AI models were surprisingly good at the first part of the job: they could look at a jumbled pile and correctly guess which pages belonged together. It was like the AI could tell, "Hey, these four pages are all about farming, and these three are about taxes," even when the pages were mixed up. However, the AI hit a massive wall when it came to the second part: putting the pages back in the right order. When the pages were shuffled, the AI struggled to figure out which page came first, second, or last. It was as if the AI could identify the characters in a story but couldn't remember the plot sequence.
The researchers dug deeper to find out why this was happening. They tried two main tricks. First, they gave the AI very specific instructions, telling it, "Hey, these pages are mixed up; please fix the order!" This helped a lot, but it didn't solve the problem completely. The AI still made mistakes, suggesting that the difficulty wasn't just about the AI not listening, but about the task itself being genuinely hard. Second, they tested the AI with English-only packets and Bangla-only packets. They found that the AI was much better at ordering the English pages than the Bangla ones. It seems that reading the visual clues in the Bangla script was an extra layer of difficulty that slowed the AI down, even though it could still group the pages correctly.
In short, this paper shows that while AI is getting great at recognizing what a document is about, it is still struggling to act like a human librarian who can look at a shuffled stack of papers and instantly know the correct reading order, especially when the text is in a language like Bangla. The researchers have now made this "messy library" puzzle available for everyone to use, hoping that by studying these failures, future AI will learn to untangle the knots and restore order to our digital documents.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.