Language Models for Portuguese: A Systematic Mapping Study
This paper presents a systematic mapping study of 46 language models developed for Portuguese, offering a comprehensive overview of their technical characteristics, evolutionary relationships, and current research gaps to guide future advancements in the field.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the internet as a giant, bustling library where every book is written in a different language. For a long time, the smartest computers in this library—the ones that can read, write, and chat like humans—were mostly trained on English books. They became incredibly fluent in English, but when you asked them about stories from other lands, they often stumbled, confused, or just made things up. This is because the "brain" of these computers is built from the words they've read; if they haven't read enough books in a specific language, they can't truly understand the culture, jokes, or history behind those words. Recently, scientists have started building special computers for Portuguese, the language spoken by hundreds of millions of people across Brazil, Portugal, and parts of Africa and Asia. But just like a library with scattered books, the information about these new Portuguese computers is messy, hidden in different places, and hard to find all at once.
This paper is like a super-organized librarian who decides to clean up the whole Portuguese section. The authors, a team from a university in Brazil, went on a massive hunt to find every single computer model designed to speak Portuguese that was published between 2020 and 2025. They didn't just count them; they opened the boxes, checked the manuals, and tried to figure out how they were built, what they were taught, and whether anyone could actually use them. They found 46 different models, ranging from general chatbots to specialists that only talk about law, medicine, or oil and gas. Their big discovery? While there is a lot of exciting progress, the "library" is still a bit chaotic. Many of these models are like secret recipes that only the chef knows (the code isn't shared), some were taught using translated books that don't quite capture the local flavor, and very few of them have been checked to see if they are fair or safe to use. The authors suggest that to make these computers truly helpful for everyone who speaks Portuguese, we need to stop relying on translations, share our blueprints more openly, and make sure the computers understand the rich diversity of the language, not just the most popular version of it.
The Great Portuguese Model Hunt
Think of the world of Artificial Intelligence (AI) as a massive construction site. For years, the biggest, most impressive skyscrapers were built using English bricks. These are the "Large Language Models" (LLMs), the smart computers that can write essays, answer questions, and even write code. But the workers on this site realized that if you only build with English bricks, you can't build a house that feels like home to someone who speaks Portuguese. So, between 2020 and 2025, a new wave of builders started constructing houses specifically for Portuguese speakers.
The problem was that these new houses were scattered all over the map. Some were described in fancy scientific journals, others were just notes on a website, and some were hidden in technical reports that no one read. It was like trying to find a specific toy in a room where the toys were thrown into different piles, some labeled, some not. The authors of this paper decided to tidy up the room. They performed a "systematic mapping study," which is a fancy way of saying they created a master list of every Portuguese language model they could find.
They found 46 models in total. That's a lot of different computers! They sorted them out like a collector sorting stamps. They looked at what "base model" each one started with (like whether it was built on a foundation of BERT, Llama, or T5), what it was designed to do (general chatting, legal advice, medical diagnosis, or analyzing tweets), and how big it was (ranging from tiny models with 12 million parameters to massive ones with 72 billion).
The "Family Tree" of Portuguese AI
One of the coolest things the authors did was draw a "phylogenetic tree." Imagine a family tree for these computers. Instead of showing how humans are related to their grandparents, this tree shows how these AI models are related to their "parents."
They found that most of the Portuguese models are descendants of a few famous "ancestor" models. The biggest family branch comes from BERT, which is like the great-grandfather of many of these models; 20 of the 46 models they found were built directly on top of BERT or its cousins. The second biggest family comes from the Llama family, which is the new, trendy ancestor that 10 models are based on.
The tree also shows how these models evolved. Some started as general-purpose computers (good at everything) and then got specialized training to become experts in specific fields. For example, there's a branch where a model started as a general reader, then studied medical textbooks to become a "CardioBERTpt" (a heart doctor AI), and then another version studied legal documents to become a "JurisBERT" (a lawyer AI). It's like a student who starts in a general school and then goes to specialized colleges to become a doctor or a lawyer.
The "Open Door" Problem
The authors also checked the doors of these 46 houses to see if anyone could walk in. They looked for three things:
- The Code: Can you see the blueprints?
- The Model: Can you download the finished computer brain?
- The Data: Can you see the books the computer was trained on?
The news here is a bit mixed. Out of the 46 models, only 11 had their source code (the blueprints) available for everyone to see. About 5 models were completely locked up—you couldn't use them at all. And for the training data (the books), only 17 models used data that was freely available to download. The rest used data that was either hidden behind a paywall, required special permission to ask for, or was kept secret by the company that built it.
The authors also checked if these models came with a "Model Card," which is like a nutrition label for AI. It tells you what the model is good at, what it's bad at, and if it has any biases. Shockingly, only 3 of the 46 models had a complete, perfect nutrition label. Most had half-finished labels, and 10 had no label at all. This means if you use these models, you might not know exactly what you're getting into.
The "Translation Trap" and Missing Voices
One of the most important findings is about how these models learned to speak. The authors noticed that many models were taught using books that were translated from English into Portuguese. Imagine trying to learn about Brazilian culture by reading a book that was originally written in New York and then translated word-for-word. You might get the grammar right, but you'd miss the slang, the jokes, and the local feelings.
The paper suggests that relying on translated data is a problem. It means these models might not understand the unique cultural nuances of Portuguese speakers in Brazil, Portugal, or Africa. Furthermore, the authors point out that almost all the models they found only focus on two versions of Portuguese: Brazilian Portuguese and European Portuguese. They completely ignored the other seven countries where Portuguese is an official language, like Angola, Mozambique, and Cape Verde. The authors argue that this is a form of "linguistic prejudice"—assuming that just because a model speaks Brazilian and European Portuguese, it can speak for everyone who speaks the language. They suggest that future models need to be trained on data that truly represents the diversity of all Portuguese-speaking nations.
What's Next?
The authors conclude that while we have made a lot of progress, we still have a long way to go. They suggest that the next generation of Portuguese AI needs to:
- Stop using translated data and start using real, original Portuguese books and conversations from all over the world.
- Share their blueprints (code) so other scientists can check their work and improve it.
- Include more voices from African and Asian Portuguese speakers, not just Brazil and Portugal.
- Add more senses. Right now, most of these models only understand text. The authors see a big opportunity to build models that can also understand images and sounds (multimodal models) specifically for Portuguese, like a computer that can look at a picture of a Brazilian street and describe it in local slang.
In short, the library of Portuguese AI is growing fast, but it needs better organization, more open doors, and a wider variety of voices to truly serve the millions of people who speak the language.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.