← Latest papers
💬 NLP

MUSE: A Full-Text Cross-Domain Knowledge Base of Scientific Problems, Solutions, and Rationales

This paper introduces MUSE, a full-text, cross-domain knowledge base containing 37,000 expert-curated Problem-Solution-Rationale triplets derived from scientific literature, which demonstrates that training large language models with rationale supervision enhances performance on complex, multi-constraint problems while potentially hindering it on simpler ones.

Original authors: Tsofia Cohen, Tom Hope

Published 2026-08-12
📖 3 min read☕ Coffee break read

Original authors: Tsofia Cohen, Tom Hope

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the world of science as a massive, endless library where every book is a research paper. For a long time, computers trying to read these books have been like students who only skim the table of contents or the very first page. They can tell you the title of the book or the main character, but they often miss the messy, brilliant, and specific details of how the characters solved their problems. Scientists don't just state a problem and a solution; they write long, winding stories explaining why they chose that specific solution over another. This is the "why" behind the "what." If we want computers to truly understand science—like a smart apprentice rather than a basic librarian—we need to teach them to read these hidden stories of reasoning. This is where a new project called MUSE comes in, aiming to turn those hidden stories into a structured map that anyone, or any computer, can follow.

The paper introduces MUSE (Mining Underlying Scientific Explanations), a giant new digital treasure chest filled with 36,960 carefully extracted "Problem–Solution–Rationale" triplets. Think of it like a recipe book, but instead of just listing ingredients (the solution) and the final dish (the result), MUSE captures the chef's secret notes: the specific problem they were trying to fix (like "the sauce is too watery"), the exact fix they applied (adding a thickener), and the chef's reasoning for why that specific thickener was the best choice (because it wouldn't change the flavor). The researchers built this by first hand-annotating 579 paragraphs from scientific papers to create a "gold standard" guide. Then, they trained a team of AI tools to act like a high-tech mining crew, scanning millions of full-text papers across many different fields to find these specific three-part stories.

The team found that while a single, all-powerful AI model tried to read a whole paragraph and guess the problem, solution, and reason all at once, it often got confused, mixing up the "why" with the "how." So, they built a modular pipeline instead: a step-by-step assembly line where one tool finds the problem, another finds the solution, and a third explains the connection. This approach worked much better, successfully turning raw text into a clean, organized database. When they tested this new knowledge base by teaching a large language model (an AI) to solve scientific problems using these "rationales" as a guide, they discovered something interesting. The AI got significantly better at solving complex problems that had many tricky constraints, suggesting that understanding the "why" helps when the path is foggy. However, for simple problems, adding the extra "why" information actually made the AI slightly worse, as if it was overthinking a task that didn't need it. This suggests that while MUSE is a powerful new tool for helping computers learn from expert reasoning, it's not a magic wand that fixes everything; it shines brightest when the scientific puzzle is truly difficult.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →