← Latest papers
💬 NLP

Polyglot Teachers: Evaluating Language Models for Multilingual Synthetic Data Generation

This paper systematically evaluates multilingual language models as teachers for generating synthetic supervised finetuning data, introducing the "Polyglot Score" to demonstrate that specific data quality attributes like prompt diversity and fluency, rather than model scale alone, determine student performance and offering practical guidelines for selecting effective teacher-student pairs.

Original authors: Lester James V. Miranda, Ivan Vulić, Anna Korhonen

Published 2026-04-14
📖 5 min read🧠 Deep dive

Original authors: Lester James V. Miranda, Ivan Vulić, Anna Korhonen

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a group of small, eager students (the Student Models) how to speak and understand six different languages. You want them to be smart, culturally aware, and good at math and chatting.

The problem? There aren't enough human teachers or textbooks for these specific languages. So, you decide to hire a Super-Teacher (a large Language Model) to write the textbooks (the Synthetic Data) for your students to study.

This paper is essentially a massive report card on which Super-Teachers are actually good at writing textbooks for these languages.

Here is the breakdown of their findings, using some everyday analogies:

1. The Old Way: "Bigger is Better" (The Wrong Assumption)

For a long time, people thought the best teacher was simply the biggest, most expensive one. They assumed, "If a teacher is a genius in English and has a massive brain, they must be a genius in every language."

The Paper's Finding: This is like hiring a world-famous French chef to teach a class on how to cook authentic Thai street food. Even though the chef is amazing, they might not know the specific spices or techniques for Thai cuisine.

  • Result: The biggest models (like Llama 3.1 70B) often wrote "textbooks" that were confusing or full of errors in non-English languages. Their students didn't learn well.

2. The New Metric: The "Polyglot Score"

The researchers created a new way to grade teachers called the Polyglot Score. Instead of just asking, "How smart is the teacher?" they asked two questions:

  1. Intrinsic Quality: How good is the textbook the teacher wrote? (Is the grammar correct? Is the vocabulary diverse? Is it natural?)
  2. Extrinsic Performance: How well did the student do on the final exam after studying that specific textbook?

They combined these into one score. It's like grading a teacher not just on their own knowledge, but on how much their students actually learned.

3. The Winners: It's About "Fit," Not Just Size

The study tested 10 different "Super-Teachers" across 6 different languages (like Arabic, German, Japanese, etc.).

  • The Surprise Winners: The Gemma 3 27B and Aya Expanse 32B models were the best teachers.
  • The Losers: Some of the giant models (like the 70B Llama) actually ranked near the bottom.
  • The Lesson: A medium-sized teacher who specializes in multilingual tasks is often better than a giant teacher who is just "okay" at them.

4. The Secret Sauce: What Makes a Good Teacher?

The researchers dug deep to find why some teachers were better. They found that brain size (parameters) didn't matter much. Instead, the quality of the "textbook" depended on:

  • Variety: Did the teacher ask different types of questions, or just repeat the same thing?
  • Fluency: Did the sentences sound like a native speaker, or like a robot?
  • Length: Was the answer detailed enough?

Analogy: Imagine a teacher who writes a 10-page essay that is boring and repetitive vs. a teacher who writes a 3-page essay that is lively, diverse, and perfectly clear. The 3-page essay is the better "textbook" for the student.

5. Three Ways to Write the Textbooks

The paper tested three methods for generating data:

  1. Generate: The teacher invents a question and answer from scratch. (Works best for languages with lots of existing data, like German).
  2. Translate: The teacher takes an English question, translates it, and answers it. (Great for languages with fewer resources).
  3. Respond: The teacher is given a question in the target language and just answers it. (Also great for less common languages).

The Takeaway: You shouldn't just pick one method for all languages. If you are teaching a rare language, "Translating" or "Responding" to existing prompts works better than trying to invent new ones from scratch.

6. The "Family Match" Rule

One of the most practical tips they found is the Family Match.

  • Analogy: If you are teaching a student who speaks "Gemma-ese," it helps if the teacher also speaks "Gemma-ese."
  • Result: When the Teacher and Student are from the same "family" of models (e.g., both are Gemma models), the student learns much faster and performs better. It's like a teacher and student sharing the same dialect; the instructions just click better.

Summary: The Recipe for Success

If you want to build a smart AI for a specific language, don't just grab the biggest, most expensive model available. Instead:

  1. Pick a teacher that is known for being good at that specific language (like Gemma 3 or Aya).
  2. Match the family: Try to use a student model from the same family as the teacher.
  3. Check the data quality: Make sure the teacher is generating diverse, fluent, and natural-sounding text, not just long, robotic paragraphs.
  4. Use the right method: For rare languages, translate existing English prompts or have the teacher respond to local prompts, rather than asking them to invent everything from scratch.

The Bottom Line: In the world of AI, quality of data matters more than the size of the model. A well-chosen, medium-sized teacher with a good "recipe" for data can teach a small student to outperform a giant student taught by a confused giant teacher.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →