ModernBERT-TR: A Modern Encoder Foundation Model for Turkish
The paper introduces ModernBERT-TR, a 150M-parameter Turkish encoder pretrained from scratch on 144.4 billion tokens using the ModernBERT architecture and a custom tokenizer, which significantly outperforms existing Turkish baselines and matches larger concurrent models despite being trained on substantially fewer resources.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the internet as a giant, chaotic library where every book is written in a different language. For a long time, the librarians (computer scientists) built one massive, super-dense encyclopedia to help computers understand all the languages at once. But when you try to find a specific fact about a tiny, specific language inside that giant book, the information gets a bit blurry. It's like trying to hear a whisper in a crowded stadium; the signal gets lost in the noise.
To fix this, researchers started building smaller, specialized dictionaries just for single languages. The most famous of these for Turkish was built back in 2020. It was a good start, but it was built with old tools, like a bicycle when everyone else was driving sports cars. Meanwhile, the rest of the world moved on to "Modern" engines that could read faster, remember longer sentences, and understand the twists and turns of grammar much better. The big question was: Could we build a brand-new, super-fast sports car specifically for Turkish, using these modern tools, without needing a billion-dollar budget or a library the size of the moon?
This paper introduces ModernBERT-TR, a new artificial intelligence model designed to be the ultimate "Turkish brain" for computers. The researchers took the latest, most advanced architectural designs (the "ModernBERT" engine) and trained it from scratch using only Turkish text. They didn't just throw a massive amount of data at it; they carefully curated a balanced mix of high-quality Turkish writing and used a custom-made dictionary (tokenizer) specifically tuned for the unique way Turkish words are built.
The results are surprisingly powerful. Even though this new model is relatively small—containing about 150 million "neurons" (parameters)—it outperformed much larger models, including some with nearly four times as many neurons. When tested on 11 different Turkish tasks, like spotting hate speech or understanding movie reviews, ModernBERT-TR scored an average of 60.2%, beating the next-best model by a significant margin. It even matched a competitor that was trained on roughly 7 times more data (1 trillion tokens vs. the model's 144.4 billion tokens) in a full fine-tuning test, proving that smart training and better architecture can beat raw data volume.
The authors suggest that the secret sauce wasn't just the size of the model, but the quality of the data mix. They found that repeating a specific, high-quality Turkish dataset five times during training was far more important than tweaking other settings like batch size. In short, they built a lean, mean, Turkish-speaking machine that proves you don't need to be the biggest to be the best; you just need to be the most modern and the most focused. They have released all their code and the model itself to the public, hoping to give a huge boost to anyone trying to build technology that speaks Turkish.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.