← Latest papers
💬 NLP

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams

This paper introduces BLUEX v2, a new benchmark comprising 395 open-ended questions from Brazil's leading university entrance exams (2022–2025) to evaluate 21 state-of-the-art LLMs on Portuguese discursive tasks, revealing significant performance gaps and highlighting challenges in mathematical reasoning and image understanding.

Original authors: João Guilherme Alves Santos, Giovana Kerche Bonás, Thiago Laitz, Thales Sales Almeida, Helio Pedrini

Published 2026-06-23
📖 5 min read🧠 Deep dive

Original authors: João Guilherme Alves Santos, Giovana Kerche Bonás, Thiago Laitz, Thales Sales Almeida, Helio Pedrini

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you've been testing how smart a group of robots are by asking them simple multiple-choice questions, like "What color is the sky? A) Blue, B) Red." They've been doing great at picking the right letter. But the researchers behind this paper, BLUEX v2, realized that picking a letter is easy. The real test of intelligence is writing a full essay, explaining why the sky is blue, and drawing a diagram to prove it.

This paper introduces a new, much harder "final exam" for AI models, specifically designed to test their ability to speak and write in Portuguese.

Here is the breakdown of their work using simple analogies:

1. The Problem: The "Multiple-Choice" Trap

Previously, researchers tested AI on Brazilian university entrance exams, but only the first round. Think of the first round as a trivia game where you just circle the right answer. It's easy for AI to guess or recognize patterns.

The second round of these exams (the one BLUEX v2 focuses on) is like a final thesis defense. Students must write long, detailed answers, solve complex math problems step-by-step, and interpret maps or graphs. The old tests couldn't measure if an AI could actually think and write like a human; they only measured if it could recognize the right answer.

2. The Solution: A New "Final Exam" for AI

The team created BLUEX v2, a massive dataset of 395 real questions from Brazil's top universities (UNICAMP and USP) between 2022 and 2025.

  • The Questions: These aren't just text. About 56% of them include images, charts, or maps that you must understand to answer correctly.
  • The Format: The AI has to write free-form answers, not pick A, B, C, or D.
  • The Subjects: It covers 9 different subjects, from History and Sociology to Physics and Math.

3. How They Graded the Robots: The "AI Teacher"

Grading a 5-page essay is hard for humans to do quickly for thousands of tests. So, the researchers built a system where an AI acts as the teacher.

  • The Rubric: First, they used an AI to read the "official correct answer" and break it down into a checklist (a rubric). For example, if the answer needs to mention "ATP" and "muscle contraction," the checklist has those two items.
  • The Judge: Another AI (the "Judge") reads the student robot's answer and checks off the items on the list. Did it mention ATP? Yes. Did it mention muscle contraction? No.
  • The Score: The robot gets a score from 0 to 10 based on how many checklist items it hit.
  • The Proof: To make sure the "AI Teacher" wasn't hallucinating, they compared its grading against real human teachers. The AI agreed with the humans 89.5% of the time. That's a very high score, proving the automated grading is reliable.

4. The Results: Who Passed and Who Failed?

They tested 21 different AI models (the "robots") on this exam.

  • The Spread: The scores ranged from a low of 4.18 to a high of 9.10. This shows the test is good at telling the difference between a smart robot and a not-so-smart one.
  • The Winners: The top performers were giant, expensive models like Gemini 3.1 Pro and GPT-5.
  • The Losers: Smaller, open-source models struggled, with the lowest scores around 4.18.
  • The Brazilian Stars: Interestingly, Brazilian-made models (Sabiá-4) performed very well, almost as good as the global giants, showing that local training helps.

5. The "Hard Parts": Where Robots Stumble

Even the best robots had trouble with two specific things:

  1. Math Reasoning: This was the hardest subject. Just like a human who can write a great story but can't do long division, the AIs could write beautiful Portuguese essays but often messed up the math steps.
  2. Image Understanding: When a question included a graph or a map, the robots' scores dropped by about 0.5 points on average. It's like giving a student a math problem but drawing the numbers on a crumpled piece of paper; the AI struggled to "see" the details clearly enough to solve the problem.

6. The Cost of the Test

The researchers were transparent about the price tag. Running this entire experiment—asking the questions, generating the image descriptions, and grading the answers—cost about $127. This is surprisingly cheap for a study of this size, showing that we can test AI rigorously without spending millions.

The Bottom Line

BLUEX v2 is a new, tougher standard for testing AI in Portuguese. It moves beyond simple trivia and forces the AI to write, reason, and look at pictures. The results show that while AI is getting very good at writing and arguing in Portuguese, it still struggles with complex math and interpreting visual data. This benchmark gives researchers a clear map of where to focus their next improvements.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →