Math Takes Two: A test for emergent mathematical reasoning in communication
The paper introduces "Math Takes Two," a new benchmark designed to evaluate whether language models can demonstrate true emergent mathematical reasoning by tasking two agents with co-developing a shared symbolic protocol to solve visually grounded problems from first principles.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Idea: Can Robots "Invent" Math?
Imagine you and a friend are stranded on a deserted island. You can’t speak the same language, and you don’t have any numbers written down anywhere. However, you need to trade coconuts. You see a pile of 12 coconuts, and your friend needs to know exactly how many are there so they can trade you some fish.
Since you don't have the word "twelve," you might start pointing, or making specific hand gestures, or even tapping stones in a pattern. Eventually, you might realize, "Hey, every time I tap this specific rock three times, it means a group of four."
You have just invented math through the need to communicate.
This paper, called "Math Takes Two," asks a profound question: Can Artificial Intelligence do the same thing? Can a computer "invent" math from scratch just because it needs to talk to another computer to solve a problem?
The Problem: Pattern Matchers vs. Real Thinkers
Right now, AI (like ChatGPT) is incredibly good at math, but it might be "cheating."
Think of it like a student who has memorized every answer in the back of the textbook. If you ask them a question from the book, they look like a genius. But if you change the numbers or use a weird new symbol they’ve never seen, they often crumble. They aren't reasoning; they are just recognizing patterns they saw during training.
The researchers want to know if AI can move from being a "Pattern Matcher" (someone who memorizes) to a "Reasoner" (someone who understands the underlying rules).
The Experiment: The "Blind Trading" Game
To test this, the researchers created a game. Imagine two AI agents: The Speaker and The Listener.
- The Speaker looks at a picture (for example, a grid of 3x4 apples).
- The Speaker cannot show the picture. They can only send a short "text message" using a tiny, random alphabet (like
A, B, C, 0, 1, 2, +, *). - The Listener receives the message and has to pick the correct picture from a lineup of options.
The Twist: The researchers don't tell the AI what the symbols mean. They don't say "1 means one object." The AI has to figure out, through trial and error, that "Symbol A" might represent a certain shape, or that "Symbol +" might mean "add these groups together."
To make it even harder, they throw a "Curveball" (Out-of-Distribution testing). Once the AI thinks it has the hang of it, the researchers introduce brand-new shapes and much larger numbers that the AI has never seen before. If the AI truly "understands" math, it should be able to use its new "invented language" to solve these new problems.
What Did They Find?
The researchers compared the AI to Humans.
- Humans are Math Wizards: Even when given a weird, made-up alphabet, humans were incredibly good at adapting. We quickly realized, "Oh, I see! This symbol represents a row, and that symbol represents a column," and we could solve problems with huge numbers we'd never seen before.
- AI is Still a Toddler: The current AI models struggled. They could handle simple tasks, but as soon as the "Curveball" arrived (new shapes or bigger numbers), they got confused. They were good at mimicking what they had seen, but they weren't "inventing" the rules of math the way humans do.
Why Does This Matter?
If we want to build AI that can solve brand-new scientific mysteries or navigate the unpredictable real world, we can't rely on them just memorizing the internet.
We need AI that can build its own tools. This paper provides a "training ground" to see if we can teach machines to do the most human thing of all: take a messy, visual world and turn it into elegant, logical math.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.