← Latest papers
💬 NLP

P3B3: A Multi-Turn Conversational Benchmark for Measuring European and Brazilian Portuguese Variety Bias in LLMs

The paper introduces P3B3, an expert-curated benchmark and evaluation framework that reveals a significant bias toward Brazilian Portuguese over European Portuguese in current Large Language Models, highlighting the urgent need for more balanced multilingual representation.

Original authors: Rafael Ferreira, Inês Vieira, Inês Calvo, James Furtado, Iago Paulo, Diogo Tavares, Diogo Glória-Silva, David Semedo, João Magalhães

Published 2026-06-16
📖 4 min read☕ Coffee break read

Original authors: Rafael Ferreira, Inês Vieira, Inês Calvo, James Furtado, Iago Paulo, Diogo Tavares, Diogo Glória-Silva, David Semedo, João Magalhães

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a group of very smart, well-read robots (Large Language Models, or LLMs) that speak Portuguese. The problem is that Portuguese isn't just one single voice; it's like a family with two very distinct cousins: European Portuguese (spoken in Portugal) and Brazilian Portuguese (spoken in Brazil).

While they share the same family name, they speak differently. They use different words for the same things (like "bus" vs. "autocarro"), talk to each other differently, and even arrange their sentences in unique ways.

The Problem: The "Brazilian Bias"

The researchers in this paper noticed something odd. Because the internet is flooded with text from Brazil, these robots have read way more Brazilian Portuguese than European Portuguese. It's like if a student only read textbooks written in one dialect; when they try to speak, they accidentally sound like they're from that region, even if they are talking to someone from the other side of the ocean.

The team wanted to know: Do these robots have a favorite cousin? And if we tell them, "Please speak like a European," can they actually do it?

The Solution: P3B3 (The "Accent Test")

To find out, the researchers created a special test called P3B3. Think of this as a "blind taste test" for language.

  • The Setup: They didn't just ask the robots, "What is a bus?" (which might trigger a specific answer). Instead, they created natural, multi-turn conversations about everyday topics like transportation and beauty products.
  • The Trick: The questions were designed to be "variety-agnostic." This means the questions didn't give any hints about which accent to use. It was like asking, "How do I get to the store?" without saying "in Lisbon" or "in São Paulo."
  • The Goal: They wanted to see what the robots would naturally say. Would they default to the Brazilian accent because that's what they've read the most? Or could they switch gears if asked?

How They Measured It

The researchers used two methods to grade the robots' answers:

  1. The "Grammar Police" (Classifiers): Computer programs trained to spot specific words and grammar rules that belong to either Portugal or Brazil.
  2. The "Expert Judge" (LLM-as-Judge): They used a super-smart AI (Gemini) to read the answers and act like a human linguist, giving a score from 0 (Pure Brazilian) to 10 (Pure European) and explaining why (e.g., "They used the word 'ônibus' instead of 'autocarro'").

What They Found

The results were a bit like a reality TV show where the contestants have a hard time hiding their true origins:

  • The Default Mode: When left alone (no instructions), almost every robot naturally spoke in Brazilian Portuguese. It was their "comfort zone." Even robots that were specifically trained to be European (like AMALIA) were the only ones that consistently stuck to the European accent.
  • The "Switch" Test: When the researchers explicitly told the robots, "Please speak Brazilian," most of them did a great job. But when they said, "Please speak European," it was much harder. Many robots struggled to stay consistent. They would start speaking European but then accidentally slip back into Brazilian words or grammar after a few sentences.
  • Size Matters: The newer, bigger, and more powerful robots were better at following instructions and switching accents than the older, smaller ones. However, even the best ones sometimes drifted back to their Brazilian roots during long conversations.

The Big Picture

The paper concludes that while these AI models are getting better, they still have a strong "Brazilian bias" because of the data they were trained on. They are like actors who are great at playing a Brazilian character but struggle to stay in character as a European one for a whole play, even when the director asks them to.

The researchers built this test (P3B3) so that in the future, developers can check if their robots are becoming more balanced and fair to all Portuguese speakers, ensuring that a user in Lisbon gets the same high-quality, culturally accurate experience as a user in Rio de Janeiro.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →