← Latest papers
🤖 machine learning

Beyond Uniform Local Isometry and Topology: FactoMap for Disentangled Representations

This paper introduces FactoMap, a novel disentanglement method that learns interpretable prototypes on a factor-space lattice to capture complex geometric structures like wrapping, collapsing, and non-uniform scales, thereby enabling the separation of statistically independent factors that are not geometrically separable in standard Euclidean spaces.

Original authors: Sohini Gupta, Bahareh Tolooshams

Published 2026-08-26
📖 5 min read🧠 Deep dive

Original authors: Sohini Gupta, Bahareh Tolooshams

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the quest to make artificial intelligence truly understand the world, researchers have long pursued a specific kind of clarity. Imagine trying to describe a complex scene, like a busy street, to a computer. If the computer simply memorizes the pixels, it cannot easily tell the difference between a red car and a blue car, or a large truck and a small one. To solve this, scientists developed a method called disentangled representation learning. The goal is to teach a machine to break down an image into its independent building blocks, or "factors," such as color, size, and position. In an ideal system, changing just the color of an object in the data should change only one specific number in the computer's internal code, leaving the numbers for size and position completely untouched. This separation is crucial because it allows machines to generalize better, make fairer decisions, and be controlled more precisely. However, a persistent assumption has guided this field: that these building blocks behave like a flat, uniform grid, where every step in any direction feels the same distance.

This assumption, however, does not always hold true in the real world. A new study from the University of Alberta challenges the idea that these factors can always be mapped onto a simple, flat surface. The researchers, Sohini Gupta and Bahareh Tolooshams, discovered that even when factors are statistically independent—meaning they are chosen without regard to one another—they can still be geometrically tangled. They demonstrated that the way a factor affects an image can change depending on the current state of other factors. For instance, changing the size of an object might alter how much its color appears to shift in the final image, creating a relationship where the "distance" between colors grows or shrinks as the object gets larger. This means that a fixed, uniform map cannot accurately represent the data. To solve this, the team introduced a new method called FactoMap, which builds a flexible, structured map that can stretch, wrap, and collapse to match the true shape of the data.

The core of the problem lies in how different factors interact with the process of rendering an image. In many previous approaches, scientists assumed that if you moved a certain amount in the "color" direction, it would always produce the same visual change, regardless of the object's size. The researchers showed that this is often false. When they simulated objects with varying hues and scales, they found that the visual impact of changing the hue depended entirely on the scale. As an object grew larger, the effect of changing its color became more pronounced relative to the effect of changing its size. This created a situation where the geometry of the data was not a simple rectangle or a flat sheet, but something more complex, where the "distance" between two colors was not constant but varied with the size of the object. Furthermore, some factors, like color, are naturally circular; a hue of zero is the same as a hue of one, creating a loop. Traditional methods often cut this loop open to fit it onto a flat line, which breaks the natural continuity of the data.

To address these issues, the researchers proposed a new framework that combines three elements: the topology, or the overall shape of the space; the identifications, which determine how different points connect or wrap around; and the local scales, which dictate how much visual change occurs at any given point. They realized that to truly separate these factors, the computer's internal map must mirror this complexity. They developed FactoMap, a system that learns a set of prototypes, or representative examples, arranged on a lattice that matches the underlying structure of the data. Instead of forcing the data into a rigid grid, FactoMap allows the lattice to take on the shape of a cone or a cylinder. In their experiments, they created a synthetic dataset called FactoShapes, featuring objects with varying hues, sizes, and positions. They found that when they used a standard rectangular grid to represent the data, the system failed to separate the factors cleanly. The map could not handle the circular nature of the color or the way the scale of the color changed with the object's size.

However, when the researchers supplied FactoMap with a lattice shaped like a cone, the results changed dramatically. In this structure, the circular direction of the lattice wrapped around to represent the hue, preserving its continuous nature without a break. Simultaneously, the cone's shape allowed the circumference of the circle to expand or contract as the lattice moved along the scale axis, perfectly matching the way the visual impact of color changed with size. This alignment allowed the system to learn a representation where the color and size factors were cleanly separated. The prototypes learned by the system organized themselves so that moving along one axis changed only the color, while moving along the other changed only the size, even though the underlying geometry was non-uniform. The study measured this success using standard metrics for disentanglement, showing that the matched cone structure significantly outperformed the mismatched grid.

The findings suggest that the path to better artificial intelligence lies not in forcing data into simple, uniform shapes, but in understanding and respecting the unique geometry of the world the data comes from. The researchers showed that statistical independence does not guarantee geometric simplicity. Factors can be sampled independently yet remain coupled through the mechanisms that generate the data. By building a map that can adapt its shape to these couplings—allowing for periodic loops, collapsing points, and varying scales—FactoMap achieves a level of clarity that rigid, flat models cannot. This approach does not just separate the factors; it does so by acknowledging that the space in which they live is not a flat sheet, but a dynamic, structured landscape. The work provides a concrete demonstration that to truly understand the factors of variation, one must first understand the shape of the space they inhabit.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →