← Latest papers
⚛️ quantum physics

HomID : Benchmarking Intrinsic Dimension Estimators on Homogenous Manifolds with Anisotropic Embeddings

This paper introduces HomID, a benchmark of homogeneous manifolds with anisotropic embeddings, to demonstrate that current intrinsic dimension estimators systematically fail under anisotropic distortions due to induced shifts in their underlying distributional assumptions.

Original authors: Aritra Das, Joseph T. Iosue, Victor V. Albert

Published 2026-10-08
📖 4 min read🧠 Deep dive

Original authors: Aritra Das, Joseph T. Iosue, Victor V. Albert

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine trying to understand the shape of a hidden object by only touching its surface. In the world of machine learning, data is often thought of as a vast, high-dimensional cloud of points, but researchers believe these points actually lie on a much simpler, lower-dimensional shape hidden inside that cloud. This idea, known as the manifold hypothesis, suggests that even though a dataset might have thousands of features, the true complexity of the data is far smaller. To make sense of this hidden complexity, scientists use tools to estimate the "intrinsic dimension," which is essentially a count of the minimum number of coordinates needed to describe the data without losing any information. If a dataset is truly simple, these tools should easily find that low number. However, when researchers apply these tools to real-world data, like images of handwritten digits, the results are often a chaotic mess, with different tools giving wildly different answers. This lack of agreement has made it difficult to know if the tools are broken or if the data is just too tricky to measure.

A team of researchers at the University of Maryland and NIST set out to solve this mystery by building a new kind of test ground. They realized that many existing tests used shapes that were too perfect, like smooth spheres where the surface looks the same in every direction. Real data, they suspected, is rarely so uniform. Instead, real data often has "anisotropy," a fancy word for a shape that stretches or squashes differently depending on the direction you look at it. To test how well current measurement tools handle this unevenness, the team created a new benchmark called HomID. This collection consists of mathematical shapes known as homogeneous spaces, which are perfectly uniform in their internal structure but are embedded into a larger space in a way that creates these directional stretches. It is a controlled environment where the researchers know the exact answer, allowing them to see exactly where and why the measurement tools fail.

When the researchers ran their standard measurement tools on these new shapes, the results were stark. Tools that performed perfectly on simple, round shapes like spheres began to stumble and fail on the HomID shapes, even when given the same amount of data. The tools systematically produced wrong answers, often underestimating or overestimating the true complexity of the data. The researchers found that this failure wasn't random; it was a direct consequence of the directional stretching. They discovered that the tools rely on assumptions about how data is distributed, such as the idea that neighbors are equally spaced in all directions. When the data is stretched, these assumptions break down, causing the tools to misinterpret the local neighborhood and calculate the wrong dimension.

To prove that this directional stretching was the culprit, the team performed a series of stress tests. They took simple, round shapes and deliberately distorted them by stretching them in specific directions, mimicking the behavior of the HomID shapes. As soon as they introduced this distortion, the performance of the measurement tools dropped significantly, mirroring the failures seen on the more complex benchmarks. They also added noise to the data to see if random errors were to blame, but found that the specific type of directional distortion was the primary cause of the errors, not just general messiness. In a final check, they applied these same distortions to real-world image datasets, such as pictures of clothing and cars, and observed the same pattern: the tools struggled when the data was stretched, confirming that the issue is relevant to practical applications.

The study concludes that the current generation of tools for measuring data complexity is surprisingly fragile. They work well on idealized, perfectly round shapes but degrade quickly when faced with the directional biases found in more realistic geometries. The researchers suggest that the field has been overestimating the robustness of these methods because previous tests were too simple. By introducing HomID, they have provided a way to expose these hidden weaknesses. The work does not claim to have fixed the tools, but it has clearly identified a specific geometric property—anisotropy—that causes them to fail. This insight offers a clear path forward for developing better methods that can handle the uneven, stretched nature of real-world data, ensuring that when we try to measure the complexity of the world, our tools are up to the task.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →