Chitrakshara: A Large Multilingual Multimodal Dataset for Indian languages
This paper introduces Chitrakshara, a large-scale multilingual multimodal dataset series comprising interleaved image-text data and image-caption pairs for 11 Indian languages, designed to address the lack of representation in Vision-Language Models and enable the development of more culturally inclusive AI systems.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a robot how to understand the world. Right now, most robots are like students who have only read books written in English and only seen pictures from Western magazines. They are smart, but they don't really understand the culture, the slang, or the daily life of people in India, where over a billion people speak dozens of different languages.
This paper introduces Chitrakshara (which roughly translates to "Image-Text" in Sanskrit), a massive new "textbook" designed specifically to teach AI about India.
Here is the story of how they built it, explained simply:
1. The Problem: The "English-Only" Robot
Think of current AI models as chefs who only know how to cook with ingredients from one specific supermarket. They can make a great English-style burger, but if you ask them to cook a spicy Biryani or a sweet Mysore Pak, they are clueless because they've never seen those ingredients before.
Most AI training data is in English. Even when researchers try to add other languages, they often just translate English articles. This is like giving a robot a recipe for a burger but calling it "Biryani." The robot learns the words, but it misses the flavor and the culture.
2. The Solution: A Massive Digital Library
The team at Krutrim AI went out and built a giant digital library called Chitrakshara. Instead of just text, this library is multimodal, meaning it contains both words and pictures mixed together, just like a real newspaper or a blog post.
They created two main "books" from this library:
- Chitrakshara-IL (The Storybook): This is a huge collection of 50 million web pages. Imagine opening a page and seeing a paragraph of text, then a picture, then another paragraph, then a chart. This is how humans actually read. It covers 11 major Indian languages (like Hindi, Tamil, Bengali, etc.).
- Chitrakshara-Cap (The Flashcards): This is a set of 44 million picture-and-description pairs. Think of it like a flashcard where you see a picture of a cow and the text says "A cow grazing in a field." This helps the AI learn to describe what it sees.
3. The Recipe: How They Made It
Gathering this data wasn't as simple as copying and pasting. The internet is messy, like a giant attic full of old boxes, broken toys, and random papers. The team had to clean it up.
- The Scavenger Hunt: They went through the "Common Crawl" (a massive archive of the entire internet) and looked for 230 million web links.
- The Filter (The Sieve): They built a sophisticated sieve to separate the gold from the dirt.
- The "Trash" Filter: They threw away ads, "Subscribe to our newsletter" pop-ups, and broken links.
- The "Quality" Filter: They checked if the pictures were blurry or if the text was just gibberish.
- The "Safety" Filter: They made sure to remove anything inappropriate or offensive, ensuring the AI learns from safe, useful content.
- The Assembly: Once cleaned, they stitched the text and images back together in the order they appeared on the original websites. This is crucial because it teaches the AI that a picture of a temple usually comes with a story about a festival, not just a random caption.
4. Why This Matters: Giving AI a "Cultural Soul"
Before this, if you asked an AI in Hindi to explain a local festival, it might give a generic, translated answer. With Chitrakshara, the AI has "read" millions of real Indian stories and seen millions of real Indian photos.
- It understands context: It knows that a picture of a "Diwali" celebration isn't just "lights"; it's about family, sweets, and specific cultural vibes.
- It speaks the language: It learns the nuances of 11 different languages, not just English.
- It's inclusive: It helps build AI that works for the "next billion" people, not just the tech-savvy few in the West.
The Bottom Line
Think of Chitrakshara as a bridge. On one side is the raw, chaotic internet of India. On the other side is the future of Artificial Intelligence. This dataset is the bridge that allows AI to finally cross over, understand Indian culture deeply, and speak to people in their own languages with genuine understanding, rather than just translating words.
It's a giant leap toward making AI that feels less like a foreign robot and more like a helpful neighbor who actually knows your community.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.