← Latest papers
💬 NLP

Improving Data and Reward Design for Scientific Reasoning in Large Language Models

The paper introduces **Dr. SCI**, a systematic post-training pipeline featuring a large-scale STEM dataset and a redesigned training workflow—comprising exploration-expanding SFT, dynamic difficulty curriculum, and rubric-guided RL—that significantly improves the scientific reasoning and open-ended problem-solving capabilities of large language models.

Original authors: Zijie Chen, Zhenghao Lin, Xiao Liu, Zhenzhong Lan, Yeyun Gong, Peng Cheng

Published 2026-02-11
📖 3 min read☕ Coffee break read

Original authors: Zijie Chen, Zhenghao Lin, Xiao Liu, Zhenzhong Lan, Yeyun Gong, Peng Cheng

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are training a world-class scientist, but instead of a human professor, you are training an AI.

Most AI training is like teaching a student using a multiple-choice textbook. It’s easy to grade: if they pick "C," they get a gold star; if they pick "B," they fail. This works great for math or simple facts, but science is much messier than that. Real science involves writing long, complex explanations, connecting different ideas, and arguing a point.

The problem is that if you give an AI an open-ended question like, "Explain how a star is born," how do you grade it? If the AI writes a beautiful, three-page essay that is mostly right but misses one tiny detail about gravity, does it get an A or an F? If the grading is too strict, the AI gets discouraged and stops trying. If it’s too easy, the AI becomes a "slacker" that learns to write long, flowery sentences that sound smart but actually say nothing (we call this "reward hacking").

This paper, "Dr. SCI," is a new "training school" designed to fix this. Here is how they did it using three clever strategies:

1. The "Diverse Library" Strategy (Exploration-Expanding SFT)

Imagine if you only taught a student by having them read the same three biology textbooks. They would become experts in those books, but they’d be clueless about chemistry or physics.

The researchers created a massive, diverse library of 1 million scientific questions. But instead of just giving the AI everything at once, they used a special filter. They looked for data that introduced new ways of thinking—new vocabulary, new logic patterns, and new ways of explaining things. It’s like making sure a student doesn't just learn what to think, but learns a hundred different ways to think.

2. The "Climbing the Mountain" Strategy (Dynamic Difficulty Curriculum)

If you try to teach a toddler calculus, they’ll give up. If you teach a PhD student basic addition, they’ll get bored. Both fail to learn.

The researchers created a "smart ladder." They started by giving the AI easy questions to build its confidence. As the AI got better, the "ladder" automatically moved up, presenting harder and harder challenges. They essentially kept the AI in the "Goldilocks Zone"—not too easy, not too hard, but just right for constant growth.

3. The "Expert Professor" Strategy (SciRubric-Guided RL)

This is the most important part. Instead of a simple "Right/Wrong" grade, they gave the AI a Rubric—just like a real professor uses.

Instead of saying, "Your answer is a 7/10," the "Professor" (a specialized AI) looks at the answer and checks specific boxes:

  • ✅ Did they define the key term? (Essential)
  • ✅ Did they explain the why behind the process? (Important)
  • ✅ Did they mention a real-world example? (Optional)
  • ❌ Did they make that common mistake about gravity? (Pitfall)

Most importantly, they added a "Truth Guard." Even if the AI writes a beautiful, poetic essay, if the final answer is factually wrong, the AI gets a failing grade. This prevents the AI from becoming a "smooth talker" who is wrong but sounds confident.

The Result

By using this "Dr. SCI" method, even a relatively small AI (a "4B" model, which is like a compact, efficient student) ended up performing better at complex science than much larger, more expensive AI models like GPT-4o in certain open-ended tests.

In short: They stopped treating the AI like a test-taker and started treating it like a true scientist.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →