Claim-Specific Admissibility of PCA Biplot Interpretations: Target Alignment, Spectral Identifiability, and Projection Adequacy
This paper argues that computational correctness in PCA biplots is insufficient for scientific interpretability and proposes a formal framework to assess the admissibility of specific claims based on target alignment, spectral identifiability, and projection adequacy.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to solve a mystery using a map. In the world of data science, there is a famous tool called Principal Component Analysis (PCA) that acts like a magical mapmaker. Its job is to take a huge, messy pile of information—like thousands of measurements about plants, people, or planets—and squash it down into a simple, two-dimensional picture. This picture, called a biplot, lets you see patterns, groups, and relationships at a glance. It's like turning a 3D sculpture into a flat shadow on the wall so you can quickly understand its shape.
But here is the catch: before the mapmaker draws the map, you have to decide how to measure the objects. Do you measure them in their raw, natural units (like centimeters, dollars, or kilograms)? Or do you first shrink and stretch everything so that every object has the exact same "size" or variance? This second step is called standardization. It's like taking a giant elephant and a tiny mouse, and resizing them both to be exactly the same height so you can compare their shapes without the elephant's size dominating the view. The problem is that this resizing changes the geometry of the map. A straight line in the original world might become a curve in the resized world, or a 90-degree angle might twist into something else entirely. Scientists have long known that standardization changes the math, but they often assume the resulting picture still tells the truth about the original, un-resized world.
This paper asks a critical question: What happens when we use a resized map to tell a story about the original, un-resized world? The authors, Luiz Roberto Martins Pinto and Carlos Tadeu dos Santos Dias, argue that just because a computer calculates the map perfectly doesn't mean the story you tell about it is true. They propose a new set of rules to check if a scientific claim is "admissible"—meaning, is it actually allowed to be made based on the picture you are looking at?
The Three Rules of the Map
The authors suggest that before you write a scientific story based on a PCA biplot, you must pass three strict tests. If you fail even one, your story is "ill-posed," meaning it's built on shaky ground, even if the math was done correctly.
1. The Target Alignment Test (Are you looking at the right map?)
Imagine you are trying to study the speed of cars. If you resize the data so that a bicycle and a Ferrari both have the same "speed variance," your map will show you how they compare to each other, not how fast they actually are. The paper argues that if your scientific question is about the original, raw speeds (the covariance), but you used a resized map (the correlation), you are looking at the wrong target. The map might be mathematically perfect for the resized data, but it is lying to you about the original cars. The authors show that standardization replaces the "covariance geometry" (the real-world distances) with "correlation geometry" (the relative shapes). If your question is about the real world, but your map is about the resized world, the story doesn't fit.
2. The Spectral Identifiability Test (Is the map pointing to a real place?)
Sometimes, the data is so perfectly balanced that the mapmaker gets confused. Imagine a spinning top that is perfectly symmetrical. If you try to point to its "front," "back," "left," or "right," you can't, because it looks the same from every angle. In the math world, this happens when two or more "eigenvalues" (which measure how much information a direction holds) are exactly the same. When this happens, the specific lines (axes) on the map are arbitrary. The map could rotate 45 degrees, and it would still be mathematically correct, but the specific labels you put on the lines would be nonsense. The paper warns that if you try to tell a story about a specific line (like "PC2 points to the north") when the math says that line is just a random choice, your story is invalid. You can only talk about the group of directions (the subspace), not the specific line.
3. The Projection Adequacy Test (Is the shadow hiding the truth?)
Finally, remember that a biplot is a flat shadow of a 3D object. If you look at a shadow of a cube, you might see a square. But if the cube is actually a complex shape with a hidden hole, the shadow won't show it. The paper introduces a way to measure how much "truth" is lost when squashing the data into two dimensions. They use a mathematical tool called a "residual-Gram bound" to check if the angles and distances you see on the 2D map are close enough to the real 3D relationships. If the map shows two variables as being perfectly aligned (0 degrees apart), but the real data says they are actually 60 degrees apart, the map is misleading. The authors show that a "good" map (one that keeps most of the data's energy) can still completely distort the specific relationship you care about.
The Evidence: Six Controlled Scenarios
To prove their point, the authors didn't just guess; they built six specific, controlled "worlds" (scenarios) where they knew the exact truth before they even started.
- The "No Connection" World: They created a world where four variables had no connection to each other. In the raw data, the map showed them as distinct, separate lines. But when they applied standardization, the map became a perfect circle of randomness. Any specific angle drawn on this circle was just a random accident of the computer's calculation, not a real pattern.
- The "Hidden Angle" World: They created a world where two variables had a clear 60-degree relationship in the real world. When they made a standard map, the two variables looked like they were pointing in the exact same direction (0 degrees). The map had collapsed the truth. The 60-degree relationship was only visible if you looked at a different part of the map that the standard analysis ignored.
- The "Symmetry" World: They built a world where the data was perfectly symmetrical. Here, the map showed that no specific direction was special. Any story claiming "Line A is the most important" was proven false because the math said any line in that group was equally valid.
The Verdict
The paper concludes that computational correctness is not enough. Just because a computer spits out a pretty picture with perfect numbers doesn't mean the story you tell about it is scientifically valid.
The authors provide a formal checklist for scientists:
- Declare your target: Are you asking about the raw world or the resized world?
- Check the map: Does the map actually represent that target?
- Check the lines: Are the specific lines you are talking about real, or are they just random choices made by the math?
- Check the shadow: Does the 2D picture preserve the specific angles and distances you need for your story?
If you can't answer "yes" to all of these, the paper argues you must either change your story, change your map, or admit that the picture cannot support your conclusion. It's a call for scientists to stop treating these colorful charts as magic truth-tellers and start treating them as tools that require careful, specific calibration. The goal isn't to stop using PCA, but to stop using it blindly.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.