← Latest papers
🧬 biology

Single-Cell Foundation Models in Biotechnology: A Scoping Review and Bibliometric Analysis of Multimodal AI for Cell-State, Perturbation, and Biomarker Prediction

This scoping review and bibliometric analysis of 1,042 studies reveals that while single-cell foundation models are rapidly expanding in biotechnology, the field suffers from inconsistent reproducibility, poor benchmark rigor, and a lack of open resources, with simple baselines often proving competitive against complex architectures.

Original authors: Behin Omidi, Mohammmad Mahdi Hemati Aalm, Arshia Farmahini Farahani, Amir Reza Hassan Poor Barkadehi

Published 2026-06-30
📖 5 min read🧠 Deep dive

Original authors: Behin Omidi, Mohammmad Mahdi Hemati Aalm, Arshia Farmahini Farahani, Amir Reza Hassan Poor Barkadehi

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Imagine a massive library of biological "stories." Each story is a single cell, and the text of the story is its genetic code (RNA). For years, scientists have been trying to read these stories one by one to understand how our bodies work, how diseases start, and how to fix them.

Recently, a new kind of "super-reader" has entered the library: Foundation Models. Think of these like the "GPT" or "ChatGPT" of biology. They are giant AI brains that have read millions of these cell-stories before. The hope is that because they've seen so much, they can understand new, unseen stories instantly, helping scientists predict how cells will react to drugs or find the root causes of disease.

This paper is a massive audit of the library. The authors didn't just read a few books; they used a computer program to scan over 1,000 research papers published between 2020 and 2026 to see what's actually happening in this field.

Here is what they found, broken down simply:

1. The Explosion of Hype

The field is growing like a wildfire. Between 2020 and 2025, the number of papers on this topic grew 24 times.

  • The Analogy: It's like when everyone suddenly started building a new type of robot. In 2020, there were a few hobbyists; by 2025, there were thousands of factories. Everyone is rushing to build these "super-readers."

2. What Are They Actually Building?

The authors looked at the blueprints of these AI models.

  • The Mix: Most of the new models are trying to be "multitaskers." They aren't just reading RNA; they are trying to read RNA, protein levels, and spatial location (where the cell is in the body) all at once.
  • The Goal: The most common jobs these models are hired for are:
    • Labeling: Figuring out what type of cell it is (e.g., "This is a liver cell").
    • Prediction: Guessing what happens if you poke the cell with a drug or a gene edit.
    • Mapping: Reconstructing the 3D map of tissues.

3. The "Big Claim" vs. The "Small Reality"

This is the most critical part of the paper. The authors asked: Do these giant, expensive AI models actually do a better job than the simple tools we already have?

  • The Claim: Many papers claim their "Foundation Model" is a miracle worker that solves problems simple tools can't.
  • The Reality: The authors found that simple tools often win.
    • The Analogy: Imagine a race between a Formula 1 car (the Foundation Model) and a very well-tuned bicycle (a simple statistical model). The paper found that on many standard tracks, the bicycle is just as fast, sometimes even faster, and it costs a fraction of the price.
    • The Problem: Many researchers only show the race where the F1 car wins. They rarely show the race where the bicycle wins, or they don't even race them against each other. The paper says only about 26% of the studies actually compare their fancy AI to a simple baseline.

4. The "Secret Recipe" Problem (Reproducibility)

In science, if you invent a new cake, you should share the recipe so others can bake it too.

  • The Finding: The authors checked if the researchers shared their "recipes" (the computer code, the trained AI brain, and the data).
    • Code: Only about 40% shared their code.
    • The AI Brain (Weights): Only 15% shared the actual trained model.
    • The Data: Only 19% shared the data used to train it.
  • The Consequence: If you can't see the recipe or taste the cake, you can't verify if it's actually good. The paper concludes that for most of these new models, nobody else can check if they actually work.

5. The "Benchmark" Issue

How do we know a model is good? We test it on a "benchmark" (a standard test).

  • The Finding: The quality of these tests is often low. The average score for how well a paper designed its test was 1 out of 4.
  • The Analogy: It's like a student taking a test where they are allowed to look at the answer key, or where the teacher only asks questions the student already knows the answer to. Many papers claim their model is a genius, but they haven't actually tested it on a hard, fair exam.

Summary: What Does This Mean?

The paper concludes that while the field is growing incredibly fast and the technology is exciting, we are in a "hype phase."

  • The Good: We have powerful new tools and a lot of data.
  • The Bad: We are over-promising on what these giant AI models can do. Simple tools are often just as good, but they get less attention.
  • The Ugly: Too many researchers are keeping their "recipes" secret and not testing their models fairly against simple alternatives.

The Bottom Line: The authors are calling for a "reality check." They want scientists to stop just building bigger, flashier models and start sharing their code, testing their models fairly against simple tools, and proving that these complex systems are actually necessary for the job. Until then, the "Foundation Models" might be more like a very expensive, very loud bicycle than the magic flying car everyone hopes they are.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →