← Latest papers
💻 computer science

SciCore-Mol: Augmenting Large Language Models with Pluggable Molecular Cognition Modules

SciCore-Mol is a modular framework that bridges the gap between linguistic symbols and molecular data by integrating pluggable perception, generation, and reasoning modules into an 8B-parameter LLM, achieving performance competitive with proprietary models across diverse chemical tasks.

Original authors: Yuxuan Chen, Changwei Lv, Yunduo Xiao, Zhongjing Du, Daquan Zhou, Yukun Yan, Zheni Zeng, Zhiyuan Liu

Published 2026-05-22
📖 4 min read☕ Coffee break read

Original authors: Yuxuan Chen, Changwei Lv, Yunduo Xiao, Zhongjing Du, Daquan Zhou, Yukun Yan, Zheni Zeng, Zhiyuan Liu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a brilliant, all-knowing librarian (a Large Language Model, or LLM) who has read every book in the world. This librarian is amazing at answering questions, telling stories, and solving logic puzzles using words. However, if you ask this librarian to design a new molecule for a medicine, they hit a wall.

Why? Because molecules aren't just words; they are 3D shapes, like tiny, complex Lego structures that can twist and turn. The librarian is used to reading flat text (like a list of ingredients), but they can't "see" the 3D shape or the invisible forces holding the Lego bricks together. When they try to guess the shape based only on a text description, they often get it wrong, missing crucial details about how the molecule actually behaves.

SciCore-Mol is a new system designed to fix this. Think of it as giving the librarian a set of specialized, plug-in tools that they can snap onto their brain whenever they need them. Instead of just guessing based on text, the librarian can now "feel" the shape, "build" the structure, and "calculate" the reaction.

Here is how the three main tools work, using simple analogies:

1. The "3D X-Ray Glasses" (Topological Perception Module)

  • The Problem: If you describe a molecule as a string of letters (like "C-C-O"), the librarian can't tell if it's a straight line or a twisted knot.
  • The Solution: This module is like a pair of 3D X-ray glasses. When the librarian sees a molecule, this tool instantly scans its 3D shape, its twists, and its angles. It translates this visual geometry into a "secret code" (a virtual token) that the librarian can understand and use in their thoughts.
  • The Result: The librarian stops guessing the shape and actually knows the geometry, leading to much more accurate predictions about how the molecule will act.

2. The "Master Sculptor" (Molecular Generation Module)

  • The Problem: If you ask the librarian to "invent a new molecule," they usually try to write it out letter by letter (like typing a sentence). This is risky; if they make one typo, the whole molecule breaks and becomes useless.
  • The Solution: This module acts like a Master Sculptor working in a clay studio. Instead of typing letters, the librarian gives the sculptor a vague idea (e.g., "make something that fights cancer"), and the sculptor slowly shapes a perfect, valid 3D molecule from a cloud of possibilities.
  • The Result: The librarian can now design new molecules that are structurally perfect and chemically valid, rather than just random strings of text.

3. The "Chemical Accountant" (Reaction Sensing Module)

  • The Problem: Chemistry isn't just about single objects; it's about how things mix and react. A librarian reading a recipe might miss the exact amounts needed or the specific conditions required for a reaction to work.
  • The Solution: This module is a Chemical Accountant. It looks at a reaction and counts the ingredients, checks the roles (who is the reactant? who is the catalyst?), and calculates the numbers (how much product will we get?). It handles the math and the logic of the reaction separately from the text.
  • The Result: The librarian can now predict the outcome of a chemical reaction with high precision, including exactly how much product will be made, without getting confused by the numbers.

How They Work Together

The genius of SciCore-Mol is that these tools are pluggable. The librarian (the main AI) stays in charge. It decides when to put on the X-ray glasses, when to call the Sculptor, or when to ask the Accountant for help. They all share a "backstage" communication channel (hidden states) so they don't have to shout instructions back and forth in plain English, which would lose information. They talk directly in their specialized languages.

The Results

The paper tested this system on many chemistry tasks:

  • Understanding: It understood 3D shapes better than other AI models.
  • Creating: It designed new molecules that were more valid and useful.
  • Predicting: It guessed reaction outcomes and yields (how much product) more accurately.
  • Knowledge: It didn't forget its general knowledge; it just got better at science.

Remarkably, this system, which has about 8 billion parameters (a medium-sized AI), performed as well as or better than massive, expensive, closed-source models (like GPT-4o or GPT-5) on these specific chemistry tasks.

In Summary

SciCore-Mol doesn't just make the librarian read more chemistry books. It gives the librarian specialized senses and tools to perceive, create, and calculate in the physical world of molecules. It bridges the gap between "words" and "reality," allowing AI to truly understand and work with the complex, 3D nature of chemistry.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →