← Latest papers
⚛️ biophysics

Information-theoretic Limits on Programmatic Specification of Biological Systems

This paper uses information theory to demonstrate that biological organisms lack sufficient internal information to fully pre-specify their microscopic organization, establishing a coarse-graining threshold that necessitates the genome act as a coarse specification compiled by shared physical dynamics and environmental randomness rather than a deterministic program.

Original authors: Kiiskinen, T., Kivinen, O., Rivas, M. A.

Published 2026-07-30
📖 5 min read🧠 Deep dive

Original authors: Kiiskinen, T., Kivinen, O., Rivas, M. A.

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Imagine trying to build a massive, self-repairing city using only a single, tiny instruction manual. You might think that if the manual is detailed enough, it could tell you exactly where every single brick, window, and streetlight should go, down to the millimeter, for every building in the city. This is the dream behind the idea that our DNA is a "master blueprint" that programs every single move our bodies make. But to understand why this might be impossible, we need to look at two big ideas. First, there's information: think of this as the total number of bits (zeros and ones) you can fit into a message. A short text message has very little information; a high-definition movie has a lot. Second, there's complexity: the sheer number of ways things can be arranged. A simple toy has few arrangements; a bustling city with millions of people moving around has an almost infinite number of possible arrangements. Scientists have long wondered: does the "instruction manual" inside a living cell (the genome) have enough information to dictate the exact position and movement of every single molecule in an organism, or is the manual just a rough sketch that relies on the laws of physics to fill in the details?

This question matters because it changes how we see life itself. If the DNA is a full program, then life is like a robot executing a pre-written script. If it's just a sketch, then life is more like a jazz band: the sheet music gives the notes and the rhythm, but the musicians (the physical laws) improvise the exact performance in real-time. A new paper by Tuomo Kiiskinen, Oscar Kivinen, and Manuel A Rivas uses the math of information theory to settle this debate. They treat the genome and the environment as a limited "budget" of information and ask if that budget is big enough to pay for the exact, microscopic details of a living body.

The authors' main finding is a resounding "no." They prove that an organism simply does not contain enough information to specify its own microscopic organization down to the exact position of every atom. They show that while the genome is brilliant at specifying the types of parts and the general rules of the game, it is far too small to act as a detailed map for every single move. Instead of a pre-written script, the genome acts more like a game plan or a set of rules for a "universal compiler" (the laws of physics) to run. The compiler takes the rough instructions and, using the laws of chemistry and physics, figures out the exact details in real-time.

To make this concrete, the authors use a fun analogy: think of the genome not as a figure-skating choreographer who dictates every twist and turn of a skater's body, but as a hockey coach's game plan. The coach specifies the roster, the tactics, and who plays on which line. But the coach doesn't tell the puck exactly where to go; physics does that at runtime. The game looks like hockey because of the team's strategy (the genome), not because every coordinate was pre-programmed.

The paper explicitly rules out the idea of "programmed microstate determinism." This is the fancy way of saying that the DNA cannot possibly contain a list of instructions that tells every molecule exactly where to be at every moment in time. The authors argue that if you try to write down the exact 3D coordinates of every atom in a protein or a cell, the amount of information required is vastly larger than the amount of information stored in the entire DNA of the organism. For example, they calculate that specifying the exact atomic structure of all the proteins in a human would require about 41% of the entire human genome's information capacity just for that one task, and that's without even counting the rest of the cell! If you try to specify the exact path of a developing embryo, the information needed is millions of times larger than the genome can hold.

The authors are very sure about this. They didn't just guess; they built a mathematical proof based on two solid rules: first, that the genome has a finite size (it can only hold so many bits), and second, that information can only travel locally (a molecule can't know what's happening on the other side of the cell instantly). They also ran simulations and checked real-world data. They looked at protein folding using AlphaFold-2 (a famous AI for predicting protein shapes) and found that the structural information in proteins is about 27 times larger than the information in the DNA that codes for them. They also simulated a whole bacterial cell (JCVI-syn3A) and found that as the cell grows, the information needed to track its exact spatial layout quickly exceeds the genome's capacity.

So, what does this mean? It means the genome is a generator specification, not a trajectory program. It tells the cell what to build and how to build it in a general sense, but it relies on the shared laws of physics to do the heavy lifting of figuring out the exact details. The genome is the blueprint for the ensemble (the group of possible outcomes), not the script for the individual trajectory (the exact path). This doesn't mean life is random; it just means that the "program" is much coarser than we thought, and the rest is handled by the physics of the universe itself. The paper confirms that no known natural environmental signal can fill this gap either; the environment helps, but it's not a secret second hard drive for the cell. The result reframes biology: we are not robots running a pre-written code, but rather the result of a compact set of rules interacting with the infinite possibilities of physics.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →