← Latest papers
💬 NLP

MATH-PT: A Math Reasoning Benchmark for European and Brazilian Portuguese

This paper introduces Math-PT, a new benchmark dataset of 1,729 mathematical problems in European and Brazilian Portuguese sourced from native competitions and exams, which reveals that while state-of-the-art LLMs perform well on multiple-choice questions, their capabilities decline on visual and open-ended tasks.

Original authors: Tiago Teixeira, Ana Carolina Erthal, Juan Belieni, Beatriz Canaverde, Diego Mesquita, Miguel Faria, Eliezer de Souza da Silva, André F. T. Martins

Published 2026-04-30
📖 4 min read☕ Coffee break read

Original authors: Tiago Teixeira, Ana Carolina Erthal, Juan Belieni, Beatriz Canaverde, Diego Mesquita, Miguel Faria, Eliezer de Souza da Silva, André F. T. Martins

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine Large Language Models (LLMs) as brilliant students who have read almost every book in the world library. For years, we have tested their mathematical abilities using exams written exclusively in English. This is like giving a student a driving test only in English and then assuming they can drive perfectly in any language just because they passed that single test.

The paper MATH-PT argues that this approach has a huge blind spot. It introduces a new "driving test" specifically written in Portuguese (both European and Brazilian variants) to check whether these AI students can truly handle mathematics when the language changes.

Here is the breakdown of their findings with simple analogies:

1. The Problem: The "English-Only" Library

Most mathematical benchmarks (such as MATH or MathVista) are like a library where every book is written in English. Even when researchers attempt to test other languages, they merely translate the English books. The authors say this is unfair, as it does not tell us whether the AI truly understands mathematical concepts or has simply memorized English patterns. They wanted to build a library of math problems born in Portuguese, drawn from real Portuguese and Brazilian math olympiads and school exams.

2. The New Test: MATH-PT

The team created a dataset named MATH-PT containing 1,729 problems.

  • The Source: They did not simply translate; they dug into the original archives of competitions such as the Olimpíadas Portuguesas da Matemática and Brazil's OBMEP.
  • The Variety: The test covers everything from 5th-grade puzzles to challenges before university studies.
  • The Format: It includes multiple-choice questions (like an answer sheet) and open-ended questions (where you must write the answer from scratch). Crucially, many questions contain diagrams and illustrations (such as geometric shapes), which the team carefully preserved using code so the AI could "see" them.

3. The Results: The "Smart One" vs. The "Struggling One"

They tested 13 different AI models, from the most powerful "Frontier" models (the super-geniuses) to smaller, open-source models (the average students).

  • The Multiple-Choice Advantage: Think of multiple-choice questions as a game of "Find the Difference." The AI models were very good at this. Even the smaller models could often guess the correct letter (A, B, C, D, or E).
  • The Struggle with Open Questions: When the test switched to open questions (where the AI must generate the answer itself), scores dropped significantly. It is the difference between recognizing a face in a lineup and drawing that face from memory. The AI had much more difficulty when it had to "think" and write the answer itself, rather than just selecting one.
  • The "Image" Problem: This was the biggest hurdle. When a question contained a diagram (an image), the AI's performance plummeted. Even the smartest models became confused when they had to interpret a visual shape alongside the text. It is like asking a student to solve a geometry problem while wearing a blindfold; they can read the words, but they cannot "see" the shape.

4. The Winners and Losers

  • The Champions: The "Frontier" models (such as GPT-5 and Qwen3-235B) were the clear winners. They handled Portuguese math well, especially the multiple-choice questions.
  • The Struggling Ones: The smaller, open-source models (such as Llama-3 and Gemma-3) found the math much more difficult. Their accuracy dropped drastically as soon as the questions became harder or contained images.
  • The Scaling Effect: The work found that for the Qwen family of models, enlarging the model (adding more "brainpower") helped very much. It is like switching from a bicycle to a sports car; the larger models handled complex reasoning tasks much better than the smaller ones.

The Conclusion

The work concludes that while AI is becoming very good at mathematics, it still exhibits a "language bias" and struggles when the test becomes harder (open-ended) or visual (diagrams). By releasing this Portuguese dataset, the authors hand researchers a new, fair ruler to measure how well AI can actually think in languages other than English. They are not saying that AI is already perfect; they are saying: "Here is a new test that shows us exactly where the AI is still learning."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →