COMPOSITE-Stem
The paper introduces COMPOSITE-STEM, a new expert-curated benchmark of 70 open-source tasks across physics, biology, chemistry, and mathematics that utilizes flexible rubric-based grading to reveal that current frontier AI agents still struggle with complex scientific reasoning, achieving a top score of only 21%.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a group of the world's smartest scientists (PhDs, professors, and researchers) getting together to build a giant, high-stakes obstacle course for Artificial Intelligence.
This paper introduces that obstacle course, called COMPOSITE-STEM. Its goal is simple but tough: to see if AI agents (robots that can think and act) can actually do real scientific work, or if they are just good at guessing answers on multiple-choice tests.
Here is the breakdown of what they did, using some everyday analogies:
1. The Problem: The "Textbook" Trap
For a long time, we tested AI like we test students in school: Multiple Choice Questions.
- The Issue: AI got really good at memorizing the answers to these questions. It was like a student who studied the answer key but couldn't actually do the math or perform the experiment.
- The Result: The tests became "saturated." The AI passed with flying colors, but we didn't know if it could actually help a scientist in a lab.
2. The Solution: The "Real-World" Obstacle Course
The team built COMPOSITE-STEM, which is like a survival challenge instead of a written exam.
- The Setup: They created 70 difficult tasks in Physics, Biology, Chemistry, and Math.
- The Twist: The AI isn't just asked to "tell me the answer." It has to go into a digital lab, open a computer terminal, install software, look at images (like X-rays or microscope photos), run code, and build the solution itself.
- The Analogy: Instead of asking the AI, "What is the chemical formula for water?" (which it can guess), they say, "Here is a messy lab bench and a broken machine. Fix it, analyze this sample, and tell me what you found."
3. The Judges: A "Panel of Experts"
How do you grade a robot's messy, open-ended work? You can't just look for a single "correct" number.
- The Old Way: Exact Match. If the robot says "350" and the answer is "350," it gets a point. If it says "350.1," it gets zero. This is too strict for real science.
- The New Way (LLM-as-a-Jury): They used a panel of five other super-smart AIs to act as judges.
- Think of it like a science fair. You don't just check the final number; you look at the process. Did the robot use the right tools? Did it understand the image? Did it explain its logic?
- The judges use a "rubric" (a checklist) to give partial credit. If the robot gets 7 out of 10 steps right, it passes, even if the final number is slightly off.
4. The Results: The "Reality Check"
They tested four of the most famous AI models (like Claude, GPT, and Gemini) on this course.
- The Score: The best AI only passed 21% of the tasks.
- The Analogy: Imagine a group of geniuses trying to fix a car engine blindfolded. Even the smartest one only managed to fix the engine correctly about 1 out of every 5 times.
- Why did they fail?
- Giving Up Too Soon: Some AIs tried to guess the answer immediately without doing the work.
- Tool Confusion: Some AIs didn't know how to install the right software (like a mechanic trying to fix a car with a hammer instead of a wrench).
- The "Strong" vs. "Weak" Gap: The top-performing AI (Claude) was much better at persisting. It kept trying different tools and checking its work, whereas the others gave up or got stuck in loops.
5. Why This Matters
This paper is a wake-up call.
- The Good News: We finally have a way to tell if an AI is actually "smart" or just "good at memorizing."
- The Bad News: Current AI is not ready to replace scientists yet. It's like a very bright intern who needs constant supervision.
- The Future: By making these tasks public, the team hopes other researchers will use them to train better AIs. They want to build a future where AI can truly help humans discover new medicines, solve climate change, and unlock the secrets of the universe.
In short: COMPOSITE-STEM is the "final exam" that proves AI still has a lot of homework to do before it can be trusted with real scientific discovery.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.