← Latest papers
🤖 AI

Generative Artificial Intelligence in Bioinformatics: A Systematic Review of Models, Applications, and Methodological Advances

This systematic review evaluates the transformative impact of generative artificial intelligence on bioinformatics, highlighting its superior performance in specialized applications like genomics and drug discovery while identifying key challenges such as data bias and scalability.

Original authors: Wasimul Karim, Riasad Alvi, Sayeem Been Zaman, Arefin Ittesafun Abian, Mohaimenul Azam Khan Raiaan, Saddam Mukta, Md Rafi Ur Rashid, Md Rafiqul Islam, Yakub Sebastian, Sami Azam

Published 2026-07-27
📖 8 min read🧠 Deep dive

Original authors: Wasimul Karim, Riasad Alvi, Sayeem Been Zaman, Arefin Ittesafun Abian, Mohaimenul Azam Khan Raiaan, Saddam Mukta, Md Rafi Ur Rashid, Md Rafiqul Islam, Yakub Sebastian, Sami Azam

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the human body as a massive, bustling library. Inside, there are billions of books written in a strange code. Some books are long scrolls of DNA (the instruction manual for building you), others are protein recipes (the workers that do the actual jobs), and some are RNA notes (the messengers passing instructions around). For decades, scientists tried to read these books using rigid, rule-based dictionaries. They could find a specific word here or a pattern there, but they struggled to understand the whole story, especially when the books were messy, incomplete, or written in a language no one had seen before.

Enter Generative Artificial Intelligence (GenAI). Think of GenAI not as a dictionary, but as a super-smart, creative storyteller who has read every book in the library and learned the rhythm, grammar, and style of the language. Instead of just looking up a word, this storyteller can predict what comes next, write new chapters that sound exactly like the original authors, or even imagine entire new stories that fit perfectly within the library's rules. This paper explores how this new "storyteller" is changing the game in bioinformatics—the science of decoding life's library. It asks: Can this AI actually write new proteins? Can it understand the complex stories of our cells better than old methods? And most importantly, is it ready to help us cure diseases, or is it just a very good mimic that sometimes makes things up?


The Great Bio-Storyteller: A Review of AI in Life Science

This paper is a massive "systematic review," which is a fancy way of saying the authors went on a treasure hunt through thousands of scientific studies to see exactly how Generative AI is being used to decode life. They didn't just look at one or two cool experiments; they gathered 82 high-quality studies from 2021 to 2026 to get the full picture. Their goal was to answer six big questions about how these AI models work, where they shine, and where they might trip up.

The New Superpowers: What Can GenAI Actually Do?

The authors found that GenAI is doing things that used to be impossible or took humans years to figure out.

1. Reading the DNA Script (Genomics)
Imagine trying to find a specific instruction in a book that has no spaces between the words. That's what DNA looks like to a computer. Old methods were like using a magnifying glass to find a single word. GenAI, however, treats the whole DNA sequence like a sentence. Models like DNABERT have learned to read this "language" so well that they can predict where genes start and stop, or how a gene might behave, with incredible accuracy (around 91.8% to 91.9% in some tests). It's like the AI learned the grammar of life and can now guess the next word in a sentence of DNA before it's even written.

2. Designing New Proteins (Proteomics)
Proteins are the tiny machines that run your body. Usually, scientists have to wait for nature to invent a new one, or they try to tweak an existing one. GenAI changes this by acting like a creative architect. Models like ProGen and ProtGPT2 can design brand-new protein sequences from scratch. In one famous experiment, an AI designed a new enzyme (a protein that speeds up chemical reactions) that actually worked just as well as a natural one. It's as if the AI didn't just copy a blueprint; it drew a new house that was just as sturdy and functional as the ones humans built.

3. Understanding the Cell's Neighborhood (Single-Cell Analysis)
Your body has trillions of cells, and they are all different. Old tools looked at a whole crowd and gave an average answer. GenAI models like scGPT can look at individual cells and understand their unique "personalities." They can predict how a cell will react to a drug or how it will change as you get sick. It's like having a translator that can listen to a whisper from a single cell in a noisy stadium and understand exactly what it's saying.

4. Finding New Medicines (Drug Discovery)
Finding a new drug is like trying to find a specific key that fits a lock, but you don't know what the lock looks like. GenAI helps by generating millions of potential "keys" (molecules) and testing them virtually. Tools like DrugAssist let scientists chat with the AI, saying, "Make me a molecule that dissolves well in water and can cross the brain barrier." The AI then designs it. While these designs are promising, the paper notes that most are still just digital sketches; they need real-world lab testing to prove they actually work in a human body.

The Secret Sauce: Specialized vs. General AI

One of the paper's biggest discoveries is that specialized AI is better than general AI for biology.

Think of a general AI (like the ones you might chat with online) as a brilliant student who has read every book in the library but has never worked in a lab. They know a lot of facts but might get confused by the specific jargon of biology.
In contrast, specialized models (like ESM-2 for proteins or DNABERT for DNA) are like students who have spent their entire lives in the lab. They were trained specifically on biological data. The paper shows that these specialized models consistently outperform the general ones. For example, when predicting how a mutation in a protein might change its function, the specialized models got it right much more often. The authors suggest that while general AI is great for summarizing text or writing code, you need a "biologist-trained" AI to truly understand the complex, messy language of life.

The Hiccups: Where the AI Gets Stuck

Despite the excitement, the paper is very careful not to call this a "solved problem." The authors point out several serious limitations that scientists still need to fix:

  • The "Hallucination" Problem: Just like a creative writer might invent a fake fact that sounds real, GenAI can sometimes generate a protein sequence that looks perfect on a computer but doesn't actually fold into a shape that works in real life. The paper warns that we can't trust these digital creations until they are tested in a wet lab (a real laboratory with test tubes and cells).
  • The Data Bias: The AI learned from the books that were already written. But most of those books are about common organisms like humans and mice. The AI is terrible at understanding rare animals or rare diseases because it hasn't read enough about them. It's like a student who only studied for the math test but forgot to read the biology chapter.
  • The Cost of Brains: Training these massive AI models requires supercomputers that use a huge amount of electricity. The paper notes that this is expensive and creates a lot of carbon emissions, making it hard for smaller labs to use these tools.
  • The "Black Box" Mystery: Sometimes the AI gives a correct answer, but no one knows why. In medicine, knowing why is just as important as the answer itself. If an AI says a drug will work, doctors need to understand the reasoning to trust it. Currently, many of these models are "black boxes" where the inner logic is hidden.

The Future: Building a Better Team

So, where do we go from here? The authors suggest that the future isn't about building one giant, perfect AI. Instead, they propose a modular approach. Imagine a team where a general AI acts as a manager, taking your question and delegating the hard parts to specialized tools. If you ask, "How does this gene affect the heart?", the manager AI could ask a protein model for the structure, a drug model for interactions, and a literature model for past studies, then put the answers together for you.

They also emphasize the need for efficiency. We need to build smarter, smaller models that don't require a supercomputer to run, so more scientists can use them. And finally, they stress the need for real-world validation. The paper concludes that while GenAI is a powerful tool for generating ideas and designs, it is not yet a replacement for the scientific method. The AI can write the story, but we still need to check if the story is true in the real world.

In short, Generative AI is a revolutionary new tool in the biologist's toolbox. It can write new proteins, decode complex genes, and design drugs faster than ever before. But like any powerful tool, it needs to be used carefully, checked constantly, and guided by human experts to ensure that the stories it tells are not just creative, but true.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →