Diversidade linguística e inclusão digital: desafios para uma ia brasileira
Drawing on sociolinguistic insights, this paper argues that the dominance of generative AI models threatens linguistic diversity by creating a self-reinforcing cycle where only well-documented language varieties are preserved, while underrepresented dialects face marginalization due to a lack of training data.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine that Artificial Intelligence (AI) is like a super-learner who is trying to understand how humans speak. To learn, this student needs to read millions of books, listen to thousands of conversations, and watch countless videos. This collection of information is called a "training dataset."
This paper, written by Raquel Meister Ko. Freitag, argues that if we want to build a truly Brazilian AI, we are currently feeding this student a very limited diet. Here is the breakdown in simple terms:
1. The Problem: The "One-Size-Fits-All" Myth
The Brazilian government has a plan to build a national AI that represents the country's diversity. However, there is a big misunderstanding: Brazil does not just speak one language.
- The Myth: Many people think everyone in Brazil speaks the exact same "Portuguese."
- The Reality: Brazil is a linguistic melting pot. Besides Portuguese, there are:
- Indigenous languages (spoken by native peoples).
- Sign Language (Libras) for the deaf community.
- Immigrant languages (like Italian or German dialects in some regions).
- Creole languages and many different dialects of Portuguese.
The Analogy: Imagine a chef trying to cook a "Brazilian Feast." If the chef only buys tomatoes and ignores the peppers, corn, cassava, and spices that make up the real cuisine, the dish won't taste like Brazil. It will just taste like a generic tomato soup. Currently, most AI models are only eating "tomatoes" (standard Portuguese), ignoring the rest of the ingredients.
2. The Danger: The "Vicious Circle" of Bias
When AI is trained only on "standard" or "prestigious" Portuguese (the kind spoken in newsrooms and universities), it learns that this is the only correct way to speak.
- The Bias: The AI starts thinking that other ways of speaking (like regional slang, rural dialects, or indigenous languages) are "wrong" or "broken."
- The Consequence: The AI starts ignoring or insulting people who speak these other ways.
- The Vicious Circle: Because the AI only understands the "standard" language, it only generates data in that language. This makes the standard language even more dominant, while the other languages disappear from the digital world. It's like a library that only buys books written in one style; eventually, all other styles of writing vanish because no one can find them in the library.
3. The Solution: Building a "Linguistic Pantry"
The author suggests that to fix this, we need to stop treating language data as a side project and start treating it as a national treasure.
- The Current Mess: Right now, linguists in Brazil have spent 50 years recording these diverse languages. They have thousands of hours of audio and text. But this data is scattered in different university drawers, locked in private files, or stored in messy formats that computers can't easily read.
- The Fix: We need to build a National Repository (a giant, organized digital pantry).
- Think of it as a centralized library where every Brazilian dialect, indigenous language, and sign language is carefully cataloged, cleaned, and made available for AI developers to use.
- This ensures that when the AI "student" learns, it gets a balanced diet of all Brazilian voices, not just the loud ones.
4. Why This Matters for "Sovereignty"
The paper argues that for Brazil to have true digital sovereignty (control over its own technology), the AI must reflect the real Brazil.
If the AI only speaks the language of the elite, it excludes millions of citizens. An ethical AI must be able to understand a grandmother in the Amazon, a fisherman in the Northeast, a deaf person in São Paulo, and a speaker of an indigenous language with equal respect.
In a nutshell:
To build a Brazilian AI that works for everyone, we must stop pretending Brazil is a monolingual country. We need to gather all the scattered pieces of our linguistic puzzle and feed them to the AI, so it learns to speak the language of the entire nation, not just a small, privileged part of it.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.