Multilingual TinyStories: A Synthetic Combinatorial Corpus of Indic Children's Stories for Training Small Language Models
This paper introduces Multilingual TinyStories, a large-scale synthetic corpus of over 132,000 children's stories in 17 Indic languages, generated via a hybrid pipeline of native model generation and cross-lingual translation to address data scarcity for training Small Language Models.
Original paper dedicated to the public domain under CC0 1.0 (http://creativecommons.org/publicdomain/zero/1.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a tiny, eager puppy how to speak. If you throw a massive, chaotic encyclopedia at it, the puppy gets confused. But if you give it simple, short, and clear stories about rabbits, forests, and lost toys, the puppy learns the rules of language quickly and happily.
This paper is about building a massive library of those simple stories, but with a twist: it's written in 17 different Indian languages, many of which don't have enough good books for computers to learn from.
Here is the story of how they built this library, explained simply:
1. The Problem: The "Empty Bookshelf"
For a long time, the smartest computer brains (AI models) have been trained mostly on English. They are like students who only read English textbooks. If you try to teach them Hindi, Tamil, or Urdu, they struggle because there aren't enough clean, simple, high-quality stories available in those languages. The existing data is often messy, like a pile of old newspapers mixed with random internet comments.
The researchers wanted to fix this for Small Language Models (SLMs). Think of these models as "puppy-sized" AIs. They are smaller, cheaper, and faster than the giant AIs, but they need very specific, clean training data to learn properly.
2. The Solution: A "Story Factory"
Instead of trying to find existing stories (which are rare), the team decided to manufacture them. They built a "Story Factory" with two main machines:
The Master Chef (Sarvam-M): For 7 of the languages, they used a powerful AI called Sarvam-M. But they didn't just ask it to "write a story." That would be boring and repetitive.
The Slot Machine (Combinatorial Prompts): Instead, they built a "slot machine" for the AI. They created a template with empty slots, like a Mad Libs game:
- Character: [A curious rabbit]
- Setting: [A magical forest]
- Problem: [A lost lantern]
- Lesson: [Bravery]
They mixed and matched these slots millions of times. This forced the AI to write millions of unique stories that were simple enough for a 5-year-old to understand, ensuring the computer learned the basics of grammar and storytelling without getting confused by complex words.
3. The Expansion: The "Translator Bridge"
They had great stories in 7 languages, but they needed 17. To get the other 10 languages, they used a Translator Bridge.
- They picked Gujarati as their "base camp" because the AI wrote the cleanest stories in that language.
- They used Google Translate to carefully translate those Gujarati stories into the other 10 languages (like Urdu, Sanskrit, and Manipuri).
- The Safety Net: They added a strict filter to make sure no English words or Latin letters "leaked" into the stories. If a story had a stray English word, it got thrown in the trash. They wanted the stories to be pure and native.
4. The Result: A Massive Library
The result is a dataset called Multilingual TinyStories.
- Size: It contains over 132,000 stories.
- Scale: That's roughly 94 million words (tokens).
- Variety: It covers 17 languages, including some that are rarely seen in AI research.
5. Why This Matters
Think of this dataset as a gym for small AI models.
- Before this, small AI models trying to learn Indian languages were trying to lift heavy, rusty weights (messy data).
- Now, they have a set of perfectly balanced, lightweight dumbbells (clean, simple stories).
- This allows researchers to train smaller, more efficient AI models that can understand and speak these languages fluently, making technology accessible to millions more people across India.
The Catch (Limitations)
The authors are honest about the flaws. Since the stories were made by a computer, sometimes the logic might be slightly weird (like a rabbit flying to the moon without a rocket). Also, the translated stories might not sound exactly like how a human grandmother would tell a story; they might feel a bit "robotic" in their phrasing. But for teaching a computer the basics of language, it's a fantastic starting point.
In a nutshell: They built a factory to print millions of simple, clean children's stories in 17 Indian languages, giving small computer brains the perfect training ground to learn how to speak and understand their local tongues.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.