Thermodynamic cost of inference and learning in physical neural networks
This paper demonstrates that while quasi-static inference in physical neural networks incurs no thermodynamic cost, learning parameters carries an irreducible energy expense, establishing that the fundamental thermodynamic price of neural computation is determined by memory storage rather than arithmetic processing.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Every time a computer solves a problem or learns a new pattern, it burns energy. We see this in the massive power plants that feed our data centers and the heat radiating from our laptops. For decades, scientists have wondered how much of this energy cost is simply a flaw in our current technology, and how much is a fundamental law of nature that no machine can ever escape. The standard answer for digital computers comes from a rule called Landauer's principle. It states that if a machine erases a piece of information, it must release a tiny, unavoidable amount of heat. This rule applies to the binary switches inside our processors, where information is stored as distinct, stable states like "on" or "off." However, many researchers are now building a different kind of computer. These are physical neural networks, devices made of light, mechanical parts, or electronic circuits that compute by relaxing into a state of balance, much like a ball rolling down a hill to find the lowest point. Because these machines do not rely on flipping rigid switches, the old rules about erasing bits might not apply to them in the same way. The question is whether these new machines can compute and learn with almost no energy at all, or if they too are bound by a hard physical limit.
A researcher at Brookhaven National Laboratory has now mapped out exactly how much energy these physical networks require. They did not build a physical device in a lab; instead, they created a detailed mathematical model that treats a neural network as a system of springs and weights. In their model, the connections between the layers of the network are not rigid instructions but elastic constraints, like springs that pull the system toward a specific shape. When the network is given an input, it simply relaxes into the shape that satisfies all the springs at once. This approach allowed the scientist to calculate the energy cost of two distinct processes: inference, which is the act of solving a problem, and learning, which is the act of training the machine to get better at solving problems. Their findings reveal a surprising split in the physics of these machines. They discovered that the act of thinking, or inference, can theoretically be done with zero energy if the machine moves slowly enough. The cost only appears when the machine is forced to move quickly. In contrast, the act of learning carries a permanent energy price tag that cannot be avoided, no matter how slowly the machine operates.
The researcher found that for inference, the energy cost is not determined by the number of connections or "weights" inside the network, as it is in digital computers. Instead, the cost depends on the width of the network, or how many neurons are active at once. If the machine is allowed to relax slowly, moving from one input to the next at a leisurely pace, it requires no work at all. This is because the machine is always in a state of equilibrium, sliding along a smooth valley of energy without ever having to climb a hill or erase a memory. The only time energy is spent is when the machine is pushed to move faster than its natural speed of relaxation. In this fast regime, the energy required is proportional to the distance the machine must travel through its state space, divided by the time it takes. The simulations showed that even at the fastest useful speed, the energy cost is remarkably low, roughly equivalent to the thermal energy of the environment for every dimension of the widest layer in the network. This means the cost scales with the size of the network's width, not its total complexity. For a network with a million active neurons, the energy cost is about a million times the tiny thermal energy of a single particle, a figure that is billions of times smaller than what current digital hardware consumes.
Learning, however, tells a different story. When the machine is trained, both the input and the desired output are forced onto the system at the same time. This creates a conflict, or strain, within the springs of the network, because the current settings cannot satisfy both the input and the target simultaneously. To learn, the machine must adjust its internal weights to relieve this strain. The researcher found that this process of adjusting the weights carries an irreducible energy cost. Unlike inference, this cost does not disappear even if the training is done infinitely slowly. Every time a weight is changed, a small amount of energy is dissipated as heat, roughly a few times the thermal energy of the environment per parameter. This cost is incurred once for the entire set of parameters that define the machine's knowledge. It does not depend on how many examples the machine sees during training, but rather on the number of weights it has to store. This stands in sharp contrast to digital training, where the machine must erase and rewrite information for every single example it sees, leading to a massive energy bill that scales with the size of the dataset.
The researcher tested these ideas by simulating a small network trained to recognize handwritten digits, a standard task in machine learning. They watched how the machine behaved as they switched between different numbers and as it adjusted its weights to improve its accuracy. The results matched their theoretical predictions perfectly. When the machine was asked to infer a digit quickly, the energy it consumed rose sharply, and its accuracy dropped as the work it did became comparable to the random jiggling of heat in the system. But when the machine was allowed to move slowly, it solved the problem with almost no energy, and its accuracy remained high. During training, they observed that the energy spent on adjusting the weights followed a precise pattern, confirming that a small, fixed amount of heat is released for every bit of learning that occurs. The simulations showed that once the machine had learned the task, the energy required to maintain that knowledge was negligible, but the energy required to write that knowledge in the first place was a fundamental, unavoidable cost.
These findings suggest that the thermodynamic price of a neural network is set by its memory rather than its arithmetic. The act of computing a result is nearly free, limited only by how fast the machine is forced to move. The act of learning, however, requires paying a small but permanent fee to write the parameters into the machine's physical structure. This distinction offers a new perspective on the future of computing. While digital machines are bound by the cost of erasing bits for every operation, physical machines could potentially perform complex calculations with an energy efficiency that is orders of magnitude better. The remaining gap between current hardware and these theoretical limits is not due to a flaw in the laws of physics, but rather to the inefficiencies of our current materials and designs. The study provides a clear roadmap for what is possible, showing that if we can build machines that truly relax into their solutions, the energy cost of intelligence could be reduced to the bare minimum allowed by the universe.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.