Criticality in Neural Network Function Space through Wilsonian Fixed Points and Finite-Width Corrections
This paper proposes a Wilsonian framework that interprets finite-width neural networks as interacting quantum field theories emerging from a Gaussian fixed point, where criticality arises from finite-width corrections that introduce non-Gaussian interactions, thereby linking network architecture and training dynamics to the emergence of complexity, expressivity, and phase-like behaviors in function space.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Deep learning has transformed how machines see, speak, and reason, yet the mathematical machinery behind these systems remains a black box to most. At the heart of this mystery lies a fundamental question: what actually happens inside a neural network when it learns? Traditionally, scientists have viewed these systems as vast collections of adjustable numbers, or parameters, where the goal is to tune these numbers to minimize error. However, a more abstract perspective has emerged, suggesting that a neural network is better understood not by its internal settings, but by the probability distribution of the functions it can produce. In this view, the network is a generator of possibilities, and its behavior is defined by the shape of the functions it creates rather than the specific knobs it turns. This shift in perspective allows researchers to apply tools from statistical physics and quantum field theory to understand how these systems organize themselves, particularly when they are extremely wide or when they are pushed to the edge of stability.
A team of researchers at Macquarie University and Charles Sturt University has taken this abstract view and developed a new framework to explain why neural networks behave the way they do. They propose that the behavior of these networks can be understood through the lens of "criticality," a state found in physical systems where matter is poised between order and chaos. In their analysis, they treat the idealized, infinitely wide neural network as a simple, predictable baseline, similar to a calm, flat landscape. They then examine what happens when the network is finite in size, which introduces small, complex interactions that disrupt this simplicity. The researchers argue that these finite-size effects are not merely technical errors to be ignored, but are actually the source of the network's ability to learn complex patterns and express rich behaviors.
The core of their work involves reinterpreting the relationship between the size of a network and its complexity. For years, the prevailing intuition in the field has been that adding more parameters makes a system more complex and harder to control. The authors challenge this by showing that in the language of function space, making a network wider actually simplifies its behavior. As the network grows infinitely wide, its output becomes perfectly predictable and follows a simple statistical rule known as a Gaussian process, where all complex interactions vanish. In this infinite limit, the network is essentially "free" and lacks the intricate connections needed to model difficult real-world problems. The researchers suggest that the true power of a neural network comes from the finite size of its layers, which reintroduces these necessary interactions.
To make sense of this, the authors use a method borrowed from physics called the Wilsonian approach, which studies how systems change as you look at them from different scales. They identify the infinite-width limit as a fixed point, a stable state that the system tends toward. The finite width of real-world networks acts as a perturbation, a small push away from this perfect stability. The researchers classify these pushes into three categories: some fade away as the network gets wider, some grow to dominate the system, and others sit on a delicate boundary where they can be tuned to create critical behavior. They find that the most interesting and useful properties of neural networks—such as their ability to generalize from limited data and their capacity to learn complex structures—emerge when the network operates near this critical boundary.
The paper provides several concrete ways to measure this behavior. The researchers show that as a network becomes wider, the difference between its actual output and the idealized infinite-width prediction shrinks in a predictable way, following a specific mathematical decay. They demonstrate that the "connected" parts of the network, which represent complex, non-linear interactions between different inputs, become weaker as the width increases. This confirms that overparameterization, often thought of as adding complexity, is actually a process of suppressing these interactions and moving the system toward a simpler, more Gaussian state. Conversely, a network with finite width retains these interactions, allowing it to break free from simple statistical rules and capture the nuances of real data.
A key finding of the study is the redefinition of the "edge of chaos," a concept describing the boundary between a network that is too rigid to learn and one that is too chaotic to control. The authors map this boundary onto their framework, showing that it corresponds to a critical surface where correlations between inputs and outputs persist over many layers without collapsing or exploding. In this near-critical regime, the network is sensitive enough to learn new patterns but stable enough to retain what it has learned. They argue that training a neural network is effectively a process of deforming its underlying statistical structure, moving it from a random starting point toward a region where these critical interactions are balanced.
The researchers also address the mystery of generalization, or why large networks often perform better on unseen data despite having more parameters than training examples. Their framework suggests that generalization occurs when training suppresses the specific, noisy interactions that fit the training data too closely, while preserving the broader, stable structures that apply to the real world. Overfitting, in contrast, happens when the network amplifies these specific, unstable interactions. By viewing the network through the lens of effective field theory, the authors provide a way to distinguish between useful complexity and harmful noise, suggesting that the best-performing networks are those that sit in a controlled, critical regime where finite-width effects are strong enough to be expressive but not so strong as to cause instability.
Ultimately, this work offers a new vocabulary for understanding neural networks, shifting the focus from the number of parameters to the structure of the functions they produce. It suggests that the magic of deep learning does not come from having more knobs to turn, but from finding the right balance where the system is complex enough to learn but simple enough to generalize. The authors propose that future research should focus on measuring these critical properties directly, looking at how correlations and interactions evolve during training, rather than just tracking the loss function. By treating neural networks as interacting physical systems, this approach opens the door to a more systematic understanding of how artificial intelligence learns, offering a path to design better architectures and training methods based on the fundamental principles of criticality and scale.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.