Mi:dm 2.0 Korea-centric Bilingual Language Models
KT introduces Mi:dm 2.0, a bilingual large language model family featuring 11.5B and 2.3B parameter variants that leverage a comprehensive data pipeline and cultural alignment to achieve state-of-the-art performance on Korea-specific benchmarks while supporting both research and commercial applications under an MIT license.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a robot to understand not just the Korean language, but the soul of Korean culture. That is exactly what the team at KT (Korea Telecom) did with their new creation, Mi:dm 2.0.
Think of most existing AI models as tourists who have memorized a phrasebook. They can say "Hello" and "Thank you" in Korean, but they might not understand why you bow, how to use the right level of politeness for your boss versus your friend, or the deep historical references in a joke. Mi:dm 2.0, however, is designed to be a local resident. It doesn't just speak the language; it thinks, reasons, and feels like it belongs in Korean society.
Here is a breakdown of how they built it, using simple analogies:
1. The Problem: The "Tourist" vs. The "Local"
The authors noticed that while many AI models can speak Korean, they often sound unnatural or miss the cultural nuances. It's like a tourist trying to order food in a local market but accidentally offending the chef because they didn't know the local customs. Existing models were trained on data that was either too small, too low-quality, or just didn't capture the specific "vibe" of Korean life.
2. The Solution: A "Cultural Diet"
To fix this, the team didn't just feed the AI more data; they fed it better data. They treated the training data like a chef preparing a gourmet meal:
- The Ingredients (Data): They didn't just grab everything off the internet. They built a strict "quality control" pipeline. Imagine a sieve that filters out broken sentences, rude comments, and confusing text, keeping only the clear, high-quality, and culturally relevant stories, news, and books.
- The Synthetic Supplement: Since high-quality Korean data is rare (like finding a specific rare spice), they created "synthetic" data. They took existing knowledge and rewrote it into new formats, like turning a dry encyclopedia entry into a textbook lesson or a news story into a conversation, ensuring the AI learned from diverse sources.
- The Balanced Menu: They noticed the AI was eating too much "humanities" (history, literature) and not enough "STEM" (science, math). So, they deliberately cooked up extra "STEM dishes" to ensure the AI was well-rounded and didn't get sick from a one-sided diet.
3. The Two Models: The "Master Chef" and the "Pocket Knife"
They released two versions of this AI to fit different needs:
- Mi:dm 2.0 Base (11.5 Billion Parameters): Think of this as the Master Chef. It's large, powerful, and handles complex tasks like deep reasoning, writing long stories, or solving difficult math problems. It's built by taking a smaller, 8-billion-parameter model and "stretching" it deeper (adding more layers) to make it smarter without changing its basic structure.
- Mi:dm 2.0 Mini (2.3 Billion Parameters): This is the Pocket Knife. It's small, lightweight, and designed to run on devices with limited power (like a phone or a laptop). They took the Master Chef's knowledge and "pruned" it down, keeping the most important skills so it can still do great work without needing a massive kitchen.
4. The Training: From "Reading" to "Conversation"
- Pre-training (Reading the Library): First, the AI read millions of documents to learn the basics of language, facts, and logic.
- Post-training (Apprenticeship): After reading, the AI went through a specialized apprenticeship. They taught it specific skills:
- Following Instructions: Learning to say "Yes, sir" and do exactly what you ask, not just what it thinks you want.
- Safety: Learning what not to say. They taught it to recognize harmful topics (like hate speech or illegal acts) and politely refuse to engage, acting like a responsible adult.
- Tool Use: Teaching it how to use external tools, like a calculator or a search engine, to get real-time answers.
5. The Results: Passing the "Local Test"
The team tested Mi:dm 2.0 against other famous AI models using special Korean tests (like a cultural IQ test).
- The Verdict: Mi:dm 2.0 didn't just pass; it aced the tests. It scored higher than other models on understanding Korean history, culture, and complex social situations.
- The "Needle in a Haystack": They tested if the AI could find a tiny piece of information in a massive document (like finding a needle in a haystack). Mi:dm 2.0 was excellent at this, proving it can handle long, detailed conversations without getting lost.
- Safety: When tricked with tricky questions designed to make the AI say something mean or dangerous, Mi:dm 2.0 held its ground better than its competitors, showing it has a strong "moral compass."
Summary
Mi:dm 2.0 is a bilingual (Korean-English) AI that was specifically engineered to be Korea-centric. Instead of being a generic robot that happens to speak Korean, it was raised on a diet of high-quality, culturally rich data to understand the nuances, humor, and values of Korean society. Whether you need a powerful brain for complex tasks (Base) or a nimble helper for everyday tasks (Mini), this AI is designed to feel less like a machine and more like a knowledgeable local friend.
Note: The paper states the models are released under the MIT license for research and commercial use, and the developers acknowledge that while they strive for safety, no AI is perfect and can still make mistakes.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.