Prosa: Rubric-Based Evaluation of LLMs on Real User Chats in Brazilian Portuguese
The paper introduces Prosa, the first real-user multi-turn Brazilian Portuguese chat benchmark, and demonstrates that using binary rubric scoring with multi-judge filtering significantly improves evaluation stability and discriminative power compared to traditional holistic LLM-as-a-judge methods.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to rank the best chefs in a city. In the past, you might have asked a food critic to taste a dish and give it a single score from 1 to 10. This is like how most AI models are currently evaluated: a "judge" AI reads a response and gives it one holistic score.
The problem, as this paper explains, is that different critics have different tastes. One might love spicy food, while another hates it. If you ask three different critics to rank the same 16 chefs, they might all agree on who is the worst, but they might completely disagree on who is the best. This makes the ranking unreliable.
The Paper's Solution: The "Checklist" Method
The researchers behind Prosa (a new benchmark for Brazilian Portuguese AI chats) decided to change the game. Instead of asking the judge, "How good is this dish overall?", they gave the judges a specific checklist of binary (Yes/No) questions.
- Old Way (Holistic): "Rate this conversation 1 to 10." (Subjective, prone to bias).
- New Way (Rubric-Based): "Did the AI answer the question? Yes/No. Did it stay polite? Yes/No. Did it avoid making things up? Yes/No."
Think of it like a driving test. Instead of an instructor saying, "You drove pretty well, I'll give you an 8," they check specific boxes: "Did you use your turn signal? Yes/No. Did you stop at the red light? Yes/No."
The "Filtering" Magic
The researchers realized that sometimes the checklist items themselves are broken. Maybe one question is so easy that every chef passes it, or so hard that everyone fails. These "broken questions" mess up the ranking.
So, they added a filtering step. Before finalizing the scores, they ran the checklist through a "quality control" team. They threw out the questions that were too easy, too hard, or confusing. This left them with a sharp, high-quality set of questions that could actually tell the difference between a good chef and a great one.
The Big Discovery
Here is the most surprising part of their findings:
- The Judge Matters Less Than the Method: When they used the old "1-to-10" method, three different AI judges (from different companies) gave completely different rankings for the top chefs. But when they switched to the "Checklist + Filtering" method, all three judges agreed on the exact same ranking for all 16 chefs.
- It's Cheaper: Because the checklist breaks the task down into simple Yes/No questions, you don't need a super-expensive, "super-smart" AI to do the judging. You can use a cheaper, faster AI (like Gemini 3 Flash) and still get the same perfect ranking. It costs about $2.10 to evaluate a new model with this method, compared to much higher costs for other methods.
What is Prosa?
Prosa is the first benchmark of its kind for Brazilian Portuguese.
- Real Life: Instead of using made-up test questions (like a school exam), they used 1,000 real conversations that real people had with AI in Brazil. This captures how people actually talk, the slang they use, and the real problems they try to solve.
- The Result: They ranked 16 different AI models. The top spot went to OpenAI's GPT-5.2, followed closely by Google and Alibaba models. Interestingly, an open-source model (Qwen3-235B) performed so well it beat several expensive, proprietary models.
In Summary
The paper argues that to fairly rank AI models, we should stop asking judges for a vague "feeling" of quality and start using a strict, filtered checklist of simple facts. This method removes the bias of the specific judge, makes the ranking much more stable, and saves money, all while using real-world conversations from Brazil as the test ground.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.