← Latest papers
💬 NLP

TFD: A Comprehensive Structured Tibetan Foundation Dataset for Low-Resource Language Processing and Large-Scale Modeling

This paper introduces TFD, the first comprehensive, structured, and expert-curated dataset covering all stages of Tibetan large language model development, which enables the creation of the Sun-Shine family of models that significantly outperform existing baselines in understanding, safety, reasoning, and generation.

Original authors: Cheng Huang, Fan Gao, Nyima Tashi, Yutong Liu, Yadi Liu, Wenbin Wei, Xiangxiang Wang, Yongbin Yu

Published 2026-05-27
📖 4 min read☕ Coffee break read

Original authors: Cheng Huang, Fan Gao, Nyima Tashi, Yutong Liu, Yadi Liu, Wenbin Wei, Xiangxiang Wang, Yongbin Yu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the world of Artificial Intelligence (AI) as a massive library. For languages like English or Chinese, this library is overflowing with books, magazines, and handwritten notes. AI models can read millions of these pages to learn how to speak, write, and think.

But for the Tibetan language, this library has been almost empty. While researchers have recently started bringing in some raw books (text data) to fill the shelves, there was a critical missing piece: a complete curriculum.

Think of building a smart AI like training a student. You can't just hand them a pile of encyclopedias (pre-training) and expect them to be ready for the real world. They also need:

  1. Instruction manuals (how to follow specific commands).
  2. Safety training (how to avoid saying harmful things).
  3. Ethics classes (how to choose the "best" answer when there are two good ones).
  4. Logic puzzles (how to think step-by-step to solve hard problems).

Until now, Tibetan AI had the encyclopedias but lacked the rest of the curriculum.

The Solution: TFD (The Tibetan Foundation Dataset)

The authors of this paper introduced TFD, which is like a complete, expert-designed "school kit" for Tibetan AI. It's the first resource that covers every single stage of training an AI from scratch.

The kit has two main parts:

1. TIBSTC: The "Textbook and Workbook" Collection
This is a massive library of over 11 billion words (tokens) covering everything from ancient philosophy and medicine to daily chat and law. But it's not just a pile of text; it's organized into specific "workbooks":

  • Instruction Tuning (Alpaca-Ti): Like a teacher giving specific homework ("Write a poem," "Translate this sentence").
  • Safety Alignment (Safety-Prompts-Ti): Like a "Do Not Do" list, teaching the AI what topics are dangerous or offensive and how to politely refuse them.
  • Preference Optimization (CValues-Ti & hh-rlhf-Ti): Like a "Best Answer" guide. If the AI gives two possible answers, this data teaches it which one humans prefer (e.g., the one that is helpful and harmless).

2. TIBSTC-CoT: The "Logic Puzzle" Book
This is the first large collection of Tibetan "Chain-of-Thought" data. Imagine a math problem where the student doesn't just write the answer, but writes out every step of their thinking process. This dataset teaches Tibetan AI how to think through complex problems step-by-step, rather than just guessing the result.

The Result: The "Sun-Shine" Family

To prove this "school kit" works, the researchers built a new family of AI models called Sun-Shine.

  • Sun-Shine 1.0 is a general-purpose model trained on the whole kit.
  • Sun-Shine 2.0 is a specialized "thinking" model trained heavily on the logic puzzles.

The Big Win:
The paper shows that these models, even when they are relatively small (like an 8-billion-parameter model), can outperform massive, expensive AI giants (some with hundreds of billions of parameters) on Tibetan tasks.

Why? Because they weren't just fed random text; they were given a structured education.

  • They understand the language better.
  • They are safer and less likely to say offensive things.
  • They can solve complex reasoning problems.
  • They even handle classical Tibetan (ancient texts) much better than other models, preserving the culture's traditional style rather than making it sound modern or robotic.

The Bottom Line

The paper argues that for low-resource languages like Tibetan, size isn't everything. You don't just need more data; you need the right kind of data organized in a complete pipeline. By providing this first-ever "full-stack" dataset, the authors hope to help build AI that is not only smart but also culturally respectful and safe for Tibetan speakers.

They have released this "school kit" (the dataset) and the "graduates" (the models) to the public so others can build upon this foundation.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →