Iterative Erasure Count Is Not an Affine-Invariant Concept Dimension
This paper demonstrates that iterative erasure counts used to estimate concept dimensions are not affine-invariant and thus depend on the specific measurement procedure and representation geometry rather than reflecting an intrinsic semantic property of the neural representation.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Hidden Trap in Counting Neural "Directions"
Imagine you are trying to understand how a super-smart robot sees the world. You might think, "If I can ask the robot a question and get an answer, the information must be stored in a specific direction inside its brain." Scientists have developed a clever trick called "iterative erasure" to count these directions. It works like this: you find the direction the robot uses to answer a question, then you mathematically "erase" that direction and ask the question again. If the robot still knows the answer, you erase another direction. You keep counting until the robot finally says, "I don't know." The number of erasures you performed was thought to be the true "dimension" of the concept—like counting how many distinct colors are mixed to make a specific shade of purple.
But here is the catch: what if the robot's brain isn't a fixed, rigid box, but a stretchy, shape-shifting balloon? If you squeeze or stretch the balloon (a mathematical move called "reparameterization"), the directions inside change their angles and lengths, even though the robot still knows exactly the same things. This paper dives into a corner of artificial intelligence called "representation learning," where researchers study how AI models organize information. The big question is: when we count how many times we have to erase a direction to stop the AI from knowing something, are we counting a real, unchangeable feature of the concept, or are we just counting how the balloon happens to be stretched at that exact moment?
The Great Stretchy-Balloon Experiment
This paper pulls back the curtain on a popular method used to measure AI concepts, revealing a surprising and somewhat mischievous truth: the count you get depends entirely on how you stretch the data, not on the concept itself.
The authors, Tingan Jin, Shuhang Dong, Haosong Li, and Chung-Hsien Chou, argue that the "iterative erasure count"—the number of directions you have to delete to make an AI forget a concept—is not a fixed, intrinsic property of the AI's knowledge. Instead, it is a "procedure-relative" number, meaning it changes based on the specific mathematical tools and rules you use to do the counting.
To prove this, they built a mathematical "lab" where they could control every variable. They created a scenario where a concept (like a simple "yes/no" signal) was generated by just one hidden source. In the real world, this is like a single light switch turning a lamp on. However, they then applied a mathematical "shear" (a slanting stretch) to the data, mixing that single switch with some random noise. Even though the concept was still generated by only one source, and the AI could still be "guarded" (made to forget) by removing just one direction, the erasure counting game gave a different answer.
When they used a standard, straight-line (Euclidean) counting method:
- If the data was perfectly straight (no stretch), the counter stopped at 1.
- If they applied even a tiny stretch (a shear factor of just 0.5 or 2), the counter suddenly jumped to 2.
It's as if you have a single light switch, but if you look at the room from a tilted angle, you suddenly think you need to unplug two different wires to turn off the light. The paper proves mathematically that for a specific type of AI probe, the count can jump from 1 to 2 (or from 2 to 4 in more complex setups) simply by changing the coordinate system, even though the underlying "concept dimension" never changed.
The "Stress Test" with Real Robots
The authors didn't just stop at math; they tested this on real, frozen AI models that analyze videos and images (specifically V-JEPA2 and DINOv2). They looked at a concept called "hand-object contact"—basically, is a hand touching an object?
They took the AI's internal features (its "brain" state) and applied different mathematical stretches (called "dense maps") to them. They then ran the erasure counting game again.
- The Result: When they used a standard, unstretched view, the AI seemed to need a certain number of erasures to forget the contact. But when they stretched the data (using factors like 2, 3, 5, or 10), the number of erasures required to make the AI "forget" changed significantly.
- The Proof: In one experiment with 4,000 samples, a standard (identity) setup stopped after 1 update in all 20 runs. But when they applied a stretch (a shear of 1), every single one of the 20 runs accepted at least 2 updates, and some went as high as 8.
This shows that the "count" isn't a stable measure of how many "physical variables" are involved in the concept. It's a measure of how the math interacts with the specific shape of the data.
The Silver Lining: A New Way to Count
The paper doesn't just say "this method is broken"; it offers a way to fix the math so it does work, but only if you follow strict rules. They prove that if you carry your "ruler" (the metric) along with the stretch, the count stays the same. This is called "affine equivariance."
Think of it like this: If you stretch a rubber sheet with a drawing on it, the drawing gets distorted. If you use a ruler that also stretches with the rubber, your measurements stay accurate. If you use a rigid metal ruler, your measurements will be wrong. The authors show that if you use the "right" ruler (a transported metric) and follow the right steps, the count becomes stable. However, they emphasize that the popular methods currently used in the field often use the "rigid metal ruler," leading to misleading counts.
What This Means for You
The main takeaway is a warning for anyone trying to count the "dimensions" of an AI's thoughts. If you see a paper claiming that a concept uses "dozens of directions" because an erasure algorithm counted that many, you should be skeptical. That number might just be an artifact of how the data was stretched or how the algorithm was set up.
The authors conclude that the "iterative erasure count" is not a magical, coordinate-free truth about the AI's mind. It is a result that is jointly determined by the representation's geometry and the full measurement procedure. To get a real answer, you have to declare your rules, your ruler, and your stopping point in advance. Without those commitments, the number you get is just a reflection of your own math, not the AI's secret.
In short: Don't trust the count unless you know exactly how the balloon was stretched.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.