CrowdMath: A Dataset of Crowdsourced Mathematical Research Discussions
The paper introduces CrowdMath, a dataset of 164 expert-annotated collaborative research discussions from the MIT PRIMES–AoPS program, to benchmark large language models on understanding dynamic mathematical progress and reveals that while models can follow local discussion flows, they struggle to identify the functional significance of individual contributions in open-problem solving.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are watching a group of brilliant friends try to solve a massive, unsolved puzzle together. They aren't working in a quiet library; they are shouting ideas across a digital forum. One person suggests a path, another spots a crack in the logic, a third fixes the crack, and eventually, after dozens of back-and-forth messages, they finally build the whole picture.
This paper introduces CROWDMATH, a new dataset that captures exactly this kind of messy, collaborative puzzle-solving.
Here is the breakdown of what the researchers did, using simple analogies:
1. The Problem with Current "Math AI"
Right now, most tests for Artificial Intelligence (AI) in math are like multiple-choice exams. You give the AI a clear question (like "What is 2+2?") and a single correct answer. If the AI gets the answer right, it passes.
The authors argue this is like testing a chef by only asking them to bake a cake when the recipe is already written down. It doesn't tell us if the chef can handle a situation where the recipe is missing, the ingredients are wrong, or the oven breaks. Real math research isn't a straight line; it's a winding path full of dead ends, wrong turns, and "aha!" moments that happen over time.
2. The New Dataset: A "Reality Show" for Math
The researchers collected 164 real-life "story arcs" from a program called CrowdMath, where high schoolers and undergraduates work together online to solve hard math problems.
Think of these story arcs as episodes of a detective show:
- The Mystery: A math problem with no known answer.
- The Clues: People posting partial ideas, asking questions, or spotting errors.
- The Red Herrings: Wrong theories that lead nowhere.
- The Solution: The final post that ties everything together into a proof.
The team didn't just save the text; they labeled every single post with its "role" in the story. Did this post start the investigation? Did it make progress? Did it find an error in someone else's logic? Or did it finish the case?
3. Testing the AI: The "Next Episode" Game
The researchers asked six of the smartest AI models available to play two games using this dataset:
Game A: "What happens next?" (Next-Post Prediction)
They showed the AI the beginning of a conversation and asked it to guess the very next message.
- The Result: The AI was surprisingly good at this (about 83–88% accuracy). It was like a TV viewer who knows the characters well and can guess, "Oh, the detective is about to ask a question." The AI understood the flow of the conversation.
Game B: "What is this person actually doing?" (Role Classification)
They showed the AI a specific post and asked: "Is this person solving the problem, making a mistake, or just asking a question?"
- The Result: The AI struggled badly (only about 42% accuracy). It couldn't tell the difference between a partial step and a final solution.
- The Analogy: It's like a movie critic who can predict the plot twists but can't tell the difference between a character trying to open a door and a character actually opening it. The AI sees the words, but it doesn't understand the weight or significance of the contribution.
4. The Big Takeaway
The paper concludes that while AI is getting very good at solving math problems that have a clear "start" and "finish" (like a textbook problem), it is still bad at understanding how math is actually discovered.
Current AI models are like excellent students who can memorize the solution to a problem but are terrible at being research partners. They can't yet tell if a new idea is a breakthrough or a dead end, or if a comment is a crucial correction or just a casual question.
In short: The AI can follow the conversation, but it doesn't yet understand the journey of discovery. It knows the destination, but it gets lost in the steps to get there.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.