← Latest papers
💻 computer science

G-Mamba: A Mechanism-Guided Graph State Space Network for Multi-Objective Prediction in Tyrosine Fermentation Process

This paper proposes G-Mamba, a gray-box multi-objective prediction network that integrates mechanistic knowledge with a bidirectional Mamba architecture and graph convolutional networks to achieve real-time, computationally efficient, and accurate forecasting of biomass and product concentrations in tyrosine fermentation processes by effectively capturing both long-sequence temporal dynamics and complex spatial couplings.

Original authors: Lihui Wang, Chunyuan Wang, Wenjing Li, Min Chen, Kun Han, Jianye Xia, Haixuan Sun, Zhenying Zhao

Published 2026-09-17
📖 6 min read🧠 Deep dive

Original authors: Lihui Wang, Chunyuan Wang, Wenjing Li, Min Chen, Kun Han, Jianye Xia, Haixuan Sun, Zhenying Zhao

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the vast, humming world of industrial biomanufacturing, factories do not churn out plastic or steel; they cultivate life. Inside massive stainless steel tanks, microscopic organisms like bacteria are coaxed into producing valuable substances, from life-saving medicines to essential amino acids used in food and chemicals. The success of these operations depends entirely on keeping the microscopic inhabitants in a state of perfect balance. If the environment shifts even slightly—becoming too hot, too acidic, or lacking oxygen—the tiny workers may stop producing or, worse, die off. To manage this, engineers rely on a constant stream of data. Some measurements, like temperature or the flow of air, are taken instantly by sensors. Others, like the concentration of the final product or the number of cells growing, require taking a physical sample and sending it to a lab for chemical analysis. This delay, often lasting hours, creates a dangerous blind spot. By the time the lab reports the results, the fermentation process may have already drifted off course, making real-time control nearly impossible.

For decades, scientists have tried to bridge this gap using computers. They have built models that try to guess the hidden values based on the visible ones, hoping to predict the future state of the culture before the lab results arrive. Early attempts relied on rigid mathematical formulas based on known laws of physics and chemistry, but these often failed because biological systems are messy and full of unknown variables. Later, researchers turned to artificial intelligence, teaching computers to find patterns in historical data. While these "black box" systems could learn from experience, they often struggled with two major problems: they were too slow to handle the long, complex sequences of data generated by a full fermentation run, and they lacked an understanding of the biological rules that govern how different variables influence one another. Without this internal logic, the computer might find a pattern that looks correct but makes no biological sense, leading to unreliable predictions.

A team of researchers from the University of Science and Technology of China and the Suzhou Institute of Biomedical Engineering and Technology has proposed a new approach to solve these specific difficulties. They developed a system called G-Mamba, designed specifically to predict the outcomes of tyrosine fermentation, a process where engineered bacteria produce an amino acid used in pharmaceuticals and food. The researchers focused on a dataset containing thirty historical batches of this fermentation process, tracking thirty different variables ranging from the speed of the stirring blades to the amount of oxygen entering the tank. Their goal was to create a model that could accurately forecast two critical outcomes that are usually measured only in the lab: the density of the bacterial cells and the concentration of the tyrosine they produce.

The core innovation of G-Mamba lies in how it combines two different ways of thinking about data. First, it uses a specialized architecture designed to handle long sequences of time. Unlike older systems that struggle to remember events from the distant past or require massive computing power to do so, this new structure can track the evolution of the fermentation process over many hours with high efficiency. It learns how the state of the tank at one moment influences the state hours later, capturing the slow, steady drift of a biological culture. Second, and perhaps more importantly, the researchers did not let the computer guess how the variables relate to each other. Instead, they fed the system a map of known biological relationships. They told the model that certain actions, like increasing the flow of air, directly affect the amount of dissolved oxygen, and that the production of waste gases like carbon dioxide is linked to the acidity of the broth. This "prior knowledge" acts as a guide, ensuring the model respects the actual rules of biology.

To make this work, the system splits the data into two parallel paths. One path focuses on the timeline, watching how each variable changes second by second. The other path looks at the web of connections between the variables, using the pre-programmed map of biological rules to understand how a change in one area ripples through the rest of the system. The model also has the ability to learn new, hidden connections that the scientists did not explicitly define, allowing it to adapt to the specific quirks of each batch. Finally, a lightweight mechanism merges these two streams of information, deciding how much weight to give to the timeline versus the biological connections at any given moment. This allows the system to understand not just what happened, but why it happened, and what is likely to happen next.

When the researchers tested this new system against eight other leading models, the results were clear. In tasks where the model had to predict the current state of the culture based on recent data, G-Mamba was more accurate than any of its competitors. It also excelled at looking further ahead, predicting the trends of the next several hours with greater reliability. The most significant finding, however, appeared when the model was asked to predict both the cell density and the tyrosine concentration at the same time. In these multi-objective tasks, the system showed a unique ability to understand the trade-off between the two. It recognized that the growth of the bacterial cells and the production of the product are linked in a complex dance of resource competition. By predicting both together, the model improved its accuracy for the harder-to-predict product concentration, using the easier-to-predict cell growth as a guide. This suggests that the system successfully captured the underlying biological tension between the organism's survival and its production capabilities.

The study also highlighted the efficiency of the new approach. While other powerful models required significant computing resources and memory, especially as the time window for prediction grew longer, G-Mamba maintained a low computational cost. Its performance did not degrade as the input data became more extensive, a common failure point for older technologies. The researchers found that the system's ability to incorporate the known rules of the fermentation process was essential; when they removed this biological guidance, the model's accuracy dropped significantly. Similarly, removing the ability to track long-term time trends also led to a sharp decline in performance. The combination of these two elements proved to be the key to success.

Ultimately, this work demonstrates that the future of industrial fermentation control lies in hybrid systems that respect both the data and the science. By teaching artificial intelligence the rules of the biological world, rather than letting it guess blindly, researchers can build tools that are not only faster and more accurate but also more trustworthy. For the engineers managing these complex bioreactors, such a tool offers the promise of real-time insight, allowing them to adjust the process before a problem arises, ensuring that every batch of bacteria produces the maximum amount of valuable product with the highest possible consistency. The study confirms that when we guide powerful new algorithms with established scientific knowledge, we can solve problems that were previously too complex for either method to handle alone.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →