IsoSci: A Benchmark of Isomorphic Cross-Domain Science Problems for Evaluating Reasoning versus Knowledge Retrieval in LLMs
The paper introduces ISOSCI, a benchmark of isomorphic cross-domain science problems demonstrating that most reasoning improvements in large language models are actually driven by domain knowledge retrieval rather than structural reasoning capabilities, thereby challenging the assumption that chain-of-thought reasoning inherently enhances short-horizon scientific problem-solving.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to figure out if a student is smart because they are good at logic, or just because they have a really good memory.
Usually, when we test AI models (like the ones that power chatbots) on science questions, we can't tell the difference. If the AI gets a chemistry problem right, did it use clever reasoning to solve it? Or did it just remember the answer from its training data?
The authors of this paper, ISOSCI, created a special test to solve this mystery. Here is how they did it, explained simply:
The "Twin" Test (Isomorphic Pairs)
The researchers created pairs of "twin" problems. These twins are identical in their structure (the steps you have to take to solve them) but completely different in their content (the specific facts you need to know).
Think of it like two different recipes:
- Recipe A (Physics): "Take 2 cups of flour, add 1 egg, and bake at 350 degrees."
- Recipe B (Chemistry): "Take 2 moles of gas, add 1 constant, and calculate the pressure."
Both recipes require the exact same logic: Take a number, multiply it by a constant, and get a result. But Recipe A needs you to know about flour and eggs, while Recipe B needs you to know about gas laws.
If an AI is truly "reasoning," it should be able to solve both twins equally well because the logic is the same. If it only solves one, it's just relying on its memory of that specific subject.
The Big Discovery: Memory vs. Logic
The researchers tested several top-tier AI models using these twin pairs. They turned a "reasoning mode" switch on and off to see if it helped.
Here is what they found, using a simple analogy:
The "Flashlight" Analogy
Imagine the AI is in a dark room trying to find a specific tool to fix a machine.
- Reasoning Mode is like turning on a bright flashlight.
- Knowledge is the tool sitting on the shelf.
The paper found that for short, standard science problems, turning on the flashlight (reasoning) didn't help the AI find the tool any faster. The AI only got better at solving the problem if it already knew where the tool was (the specific science fact).
In fact, 91.3% of the time the AI got better at solving a problem, it was because it remembered the specific fact, not because it got better at the logical steps.
The "O3-Mini" Surprise
One of the most interesting parts of the paper involves a specific AI model called o3-mini.
- On a famous test called GPQA (which has very hard, deep conceptual questions), o3-mini was a superstar, beating other models by a huge margin. People thought, "Wow, this model is a reasoning genius!"
- But when the researchers put o3-mini on their new ISOSCI test (the twin logic test), it actually performed worse than a standard model.
The Lesson: The "genius" label depends entirely on which test you give the AI. If the test requires deep, abstract thinking, the reasoning model shines. If the test requires recalling specific facts and plugging them into a formula, the reasoning model doesn't help much—and might even get in the way.
The "Toggle" Experiment
The researchers also tested models where they could literally flip a switch to turn "reasoning" on or off without changing the model itself.
- Result: Flipping the switch made almost no difference (less than 5% improvement) on these short science problems.
- It's like giving a calculator a "super-calculation" mode, but the problems are so simple that the calculator doesn't need the extra power. It just needs to know the numbers.
Summary
The paper concludes that for short, step-by-step science problems:
- Reasoning isn't magic: It doesn't automatically make an AI smarter at logic.
- It's mostly about memory: When AI seems to "reason" better, it's usually just because it successfully retrieved a specific fact it needed.
- One size doesn't fit all: A model that looks like a reasoning expert on one test might look average on another, depending on whether the test needs deep thinking or just good memory.
The authors released their "Twin Test" (ISOSCI) so others can use it to stop guessing whether an AI is actually reasoning or just memorizing.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.