← Latest papers
💬 NLP

Assessment of Evolving Large Language Models in Upper Secondary Mathematics

This study demonstrates that evolving large language models have rapidly advanced in mathematical proficiency, progressing from moderate performance to achieving near-perfect scores on the Finnish upper secondary matriculation examination, thereby highlighting their potential as powerful tools to support mathematics education.

Original authors: Mika Setälä, Pieta Sikström, Ville Heilala, Tommi Kärkkäinen

Published 2026-06-18
📖 4 min read☕ Coffee break read

Original authors: Mika Setälä, Pieta Sikström, Ville Heilala, Tommi Kärkkäinen

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are watching a race where the runners are not humans, but different versions of "super-brains" (Large Language Models or LLMs). The track they are running on is the Finnish High School Math Exam, a notoriously difficult test that determines if students can get into university.

Here is the story of how these digital runners evolved from stumbling beginners to perfect champions, based on the research paper.

The Starting Line: The "Baby" Phase (August 2023)

At the beginning of the race, the AI runners were like students who had just started learning algebra.

  • The Runner: An early version of OpenAI's GPT-4.
  • The Result: It scored a 64 out of 120. In the Finnish grading system, this is a passing grade, but a low one (Grade "E"). It was like a student who passed the class but barely scraped by.
  • The Competitor: Google's "Bard" was even further behind, scoring a 34 (Grade "C"), which is like a student who failed the exam.

At this stage, the paper notes that these AI models were good at simple things but struggled with the complex logic needed for high school math. They were like calculators that sometimes forgot how to multiply.

The Training Montage: Getting Smarter (Late 2023 – Early 2024)

The researchers checked back in a few months later. The AI runners had been "training" hard.

  • The Upgrade: The AI learned to use tools, specifically a programming language called Python. Think of this as the runner suddenly being allowed to use a high-tech bicycle instead of just running on foot.
  • The Result: By late 2023, the GPT-4 runner sprinted to a 93 out of 120, earning the top grade ("L"). It could now solve tricky geometry problems that involve 3D space, which it previously found impossible.
  • The Gap: However, not all runners improved at the same speed. In early 2024, a free version of Google's AI was still stuck in the middle of the pack, while the paid versions of OpenAI and Microsoft's AI were leading the race.

The Finish Line: The "Superhuman" Era (January 2025)

By the time the researchers ran the final test in early 2025, the race had changed completely. The AI runners weren't just competing; they were breaking the scoreboard.

  • The New Champions: Two new runners, DeepSeek R1 and OpenAI's o3, crossed the finish line with a perfect score of 120 out of 120.
  • The Meaning: These models didn't just pass the exam; they scored higher than almost any human student ever has. In fact, if these AI models were real students applying to Finnish universities today, they would be accepted into the most competitive programs (like medicine or engineering) and would likely be ranked at the very top of the applicant pool.

How Did They Get So Fast? (The Secret Sauce)

The paper explains that these models didn't just "memorize" more math. They changed how they think.

  • Chain of Thought: Imagine a human solving a hard puzzle. They don't just guess the answer; they pause, think step-by-step, check their work, and fix mistakes if they get stuck. The new AI models (like OpenAI's "o1" and "o3") do the same thing. They simulate a "thinking process" before answering, allowing them to catch their own errors.
  • Self-Correction: Older models were like a student who writes an answer and hands it in immediately. The new models are like a student who writes a draft, realizes a mistake, rewrites the solution, and then hands it in. This "deliberative" thinking is what allowed them to get perfect scores.

The Bottom Line

The paper concludes that the "mathematical brain" of these AI tools has evolved incredibly fast in just a couple of years.

  • What they can do: They can now solve high-stakes, difficult math exams better than the average human student, often achieving perfect scores.
  • What the paper asks next: The paper does not claim these AIs are ready to be teachers yet. It asks: Just because they can solve the problem perfectly, can they actually explain it to a confused student in a helpful way? The paper suggests that while the "solving" part is mastered, the "teaching" part is the next big challenge.

In short: The AI went from a struggling student to a math genius in record time, but the question now is whether it can be a good tutor.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →