← Latest papers
🤖 AI

MIRA-Math: A Benchmark for Minimal Information Requesting and Mathematical Reasoning

The paper introduces MIRA-Math, a deterministic benchmark comprising 2,310 mathematical problems where solvers must identify and request a single missing atomic fact under a strict budget before computing the final answer, thereby enabling the separate evaluation of information-seeking and reasoning capabilities in language models.

Original authors: Charbel Al Bateh, Samer Saab Jr

Published 2026-07-09
📖 5 min read🧠 Deep dive

Original authors: Charbel Al Bateh, Samer Saab Jr

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to solve a complex puzzle, like a Sudoku or a math problem. Usually, when you take a test or use a smart computer program, you are given all the pieces you need to finish the picture. You just have to put them together correctly.

But in the real world, information is often missing. You might know the rules of a game, but you don't know the score. Or you know a recipe, but you don't know how much salt to add.

MIRA-Math is a new "test" designed to see if AI models (like the smart chatbots we use today) can handle this specific situation: knowing what they don't know, and asking for exactly the right piece of information to finish the job.

Here is how the paper explains it, using simple analogies:

1. The "Missing Ingredient" Game

Most math tests for AI are like giving a chef a full recipe and all the ingredients, then asking them to bake a cake. The test is: "Can you bake the cake?"

MIRA-Math is different. It gives the AI a recipe that is deliberately incomplete.

  • The Setup: The AI sees a math problem, but one crucial number or fact is hidden (like a missing ingredient).
  • The Goal: The AI must realize something is missing, ask for it in plain English, and then use that new fact to solve the problem.
  • The Catch: The AI has a strict "budget." It can only ask a few questions. If it asks the wrong thing, or asks too vaguely, it fails.

2. The "Strict Librarian"

To make this test fair and precise, the paper introduces a special character: a Fixed Information Holder (think of them as a very strict, literal-minded librarian).

  • The Librarian's Job: This librarian holds the one missing fact. They do not help you solve the problem. They do not give you hints. They do not explain things.
  • The Rules:
    • If you ask, "What is the missing number?" (too vague), the librarian says, "I don't have that."
    • If you ask, "What is the value of variable X?" (precise), and the librarian has that specific fact, they hand it over.
    • If you ask for something they don't have, they say, "No."

This setup tests if the AI can be precise in its questions, not just smart at math.

3. Two Types of Challenges

The paper created 2,310 different puzzles, split into two categories:

  • Type A (The "Fixed Slot"): The AI knows where the missing piece is.
    • Analogy: Imagine a form with a blank line that says "Date of Birth: ______". The AI knows it needs to ask for the date. The test is: Can it ask for the date correctly, and then use it to solve the math?
  • Type B (The "Hidden Slot"): The AI has to find where the missing piece is.
    • Analogy: Imagine a form where one of the ten lines is blank, but it doesn't say which one. The AI has to look at the whole form, figure out, "Oh, the 'Phone Number' line is empty," and then ask for that specific thing. This is harder because the AI has to diagnose the problem first.

4. What the Tests Revealed

The researchers tested many different AI models (some very smart, some smaller) and found some surprising things:

  • Asking and Solving are Different Skills: Just because an AI is good at asking the right question doesn't mean it's good at doing the math afterward.
    • Example: One AI asked for the perfect missing fact (like a great detective), but then failed to do the calculation (like a bad accountant).
    • Example: Another AI did the math perfectly but asked for the wrong fact, so it never got the answer.
  • Practice Doesn't Always Help: The researchers tried showing the AI examples of how to ask questions (like a "cheat sheet" or "4-shot prompting"). Sometimes this helped, but often it made the AI worse.
    • Why? The AI started copying the examples too rigidly. Instead of looking at the specific problem at hand, it just repeated the pattern it saw in the examples, even when the problem needed a different question.

5. Why This Matters

The authors say this benchmark is like a diagnostic tool for doctors.

  • Current tests tell you if a student can solve a problem if they have all the notes.
  • MIRA-Math tells you if a student can identify a gap in their knowledge and communicate effectively to fill it.

The paper concludes that for AI to be truly useful in the real world, where information is often incomplete, it needs to master this specific skill: knowing exactly what to ask for, and then using that answer correctly.

In short: MIRA-Math isn't just testing if AI can do math; it's testing if AI can admit what it doesn't know and ask the right question to fix it, without getting confused or guessing.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →