Parameter-Level Attribution of Symmetry in Trained Networks Though Parameter-Wise Functional Sensitivity
This paper establishes that a smooth parameter-space action realizing a function's symmetry exists if and only if the symmetry's tangent orbit lies within the image of the network's functional sensitivities, providing a framework to identify local parameter directions that either follow the symmetry orbit or descend toward the equivariant subspace in trained networks.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the vast landscape of modern artificial intelligence, neural networks are often treated as black boxes: complex systems that take in data and produce answers, yet offer little insight into how they arrive at those conclusions. These systems are built from layers of mathematical connections, each holding a specific numerical value, or parameter, that the machine adjusts during training. A central question for scientists studying these systems is whether the internal structure of a trained network reflects the rules governing the world it learned. If a network learns a task that involves symmetry—such as recognizing that a shape looks the same after being rotated—does the network's internal wiring actually contain a mechanism that mirrors that rotation? Researchers have long suspected that the answer is yes, but proving exactly which parts of the network are responsible for this behavior has remained elusive. This uncertainty matters because understanding the internal geometry of these models could reveal how they generalize knowledge and whether they truly grasp the physical laws they are trained to mimic.
A team of researchers at the University of Oxford and the Heidelberg Institute for Theoretical Studies has developed a new way to look inside these networks, moving beyond simple statistics to examine the precise, local geometry of how parameters influence a network's output. They focused on a concept they call functional sensitivity, which measures how much a network's final answer changes when a single internal number is tweaked slightly. By mapping these tiny changes, the researchers created a detailed picture of the directions in which the network can move its output. They asked a specific question: if a network has learned a function with a known symmetry, can that symmetry be traced back to a specific motion within the network's parameters? In other words, is there a path through the network's internal settings that corresponds exactly to rotating or shifting the function it performs?
The researchers formulated this as a lifting problem, asking whether the smooth transformations seen in the final output can be "lifted" back to a smooth transformation in the network's internal settings. They found that for such a motion to exist, the directions in which the function can change due to symmetry must be reachable by the network's current parameters. If the network's internal settings are too rigid or poorly aligned, the symmetry in the output cannot be fully realized by moving the parameters. To test this, the team developed a method to calculate the best possible direction to move the parameters to either follow a symmetry or to correct a lack of symmetry. They identified two distinct paths: one that moves the network along the symmetry orbit, effectively rotating the learned function without changing its fundamental nature, and another that moves the network toward a state of perfect symmetry, reducing any errors where the function fails to respect the symmetry.
To see if these theoretical directions worked in practice, the researchers trained neural networks on two different tasks. First, they used a simple image classifier designed to distinguish between a full circle and a wedge-shaped slice. They then applied their calculated directions to the trained network. When they moved the network's parameters along the symmetry path, the network's decision boundary rotated exactly as predicted, preserving its accuracy while changing its orientation. When they moved the parameters along the path toward symmetry, the network's output became more consistent with the rules of rotation, reducing the error in its behavior. Crucially, they found that these directions only worked reliably when they were recalculated at every step of the movement. If they calculated the direction once at the beginning and kept it fixed, the network quickly drifted off course, failing to maintain the symmetry or the correction. This showed that the relationship between the parameters and the symmetry is highly local and changes as the network moves.
The team repeated this process with a more complex system designed to learn the physics of a rotating potential, known as a Hamiltonian system. Even though the network's architecture did not explicitly force it to respect rotational symmetry, the trained network still learned the underlying physics. When the researchers applied their method here, they observed the same pattern: recalculating the direction at every step allowed the network to rotate its learned potential or make it more symmetric, while keeping the direction fixed led to a gradual drift away from the correct behavior. The study demonstrates that while neural networks can learn to encode symmetries, the specific parameters responsible for this behavior are not static. The ability to trace these symmetries back to the network's internal structure depends on constantly re-evaluating the local geometry of the parameters. The findings suggest that the internal representation of symmetry is a dynamic feature, one that requires continuous adjustment to remain aligned with the mathematical rules the network has learned to follow.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.