← Latest papers
🧬 biology

A robust, scalable K-statistic for quantifying immune cell clustering in spatial proteomics data

This paper introduces KAMP, a robust and computationally efficient method that corrects for spatial inhomogeneity in Ripley's K-statistic to accurately quantify immune cell clustering and colocalization in spatial proteomics data, thereby preventing biased downstream survival analyses.

Original authors: Julia Wrobel, Hoseung Song

Published 2026-07-15
📖 5 min read🧠 Deep dive

Original authors: Julia Wrobel, Hoseung Song

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Imagine you are a detective trying to solve a mystery inside a giant, crowded city. The city is a slice of human tissue, and the citizens are cells. Some citizens are the "good guys" (immune cells like B cells and macrophages), and the rest are the "background crowd." Your job is to figure out if the good guys are hanging out together in tight-knit groups (clustering) or if they are just wandering around randomly.

For a long time, detectives used a tool called Ripley's K. Think of this tool as a ruler that measures how close neighbors are. But here's the catch: the old ruler had a huge blind spot. It assumed the city was perfectly flat and uniform, like a giant, empty parking lot. If the city had a pothole, a broken sidewalk, or a torn-up section (which happens in real tissue samples due to tearing or degradation), the old ruler would get confused. It would look at a torn-up area where cells are naturally bunched up and scream, "Clustering! Clustering!" even if the cells were just reacting to the damage, not each other. This led to false alarms and wrong conclusions about how the immune system works.

The Big Discovery: The "KAMP" Detective
Enter KAMP (K adjustment by Analytical Moments of the Permutation distribution). The authors, Julia Wrobel and Hoseung Song, invented a super-smart, super-fast new way to measure cell groups that doesn't get fooled by the potholes.

Instead of assuming the city is a perfect parking lot, KAMP looks at the entire crowd of citizens (both the good guys and the background crowd) to see what a "normal" day looks like in that specific, messy city. It asks: "If we shuffled the names of the citizens randomly, keeping their locations exactly where they are, how often would we see these groups?"

By doing this mathematically (using a clever shortcut called "analytical moments" so it doesn't have to shuffle the cards millions of times), KAMP builds a custom "null" ruler for every single sample. It subtracts the background noise from the signal.

  • The Result: If the immune cells are still bunched up more than the shuffled crowd would be, KAMP says, "Okay, that's real biological clustering!" If they are just bunched up because of a tear in the tissue, KAMP says, "Nope, that's just the background noise."

What They Ruled Out
The paper explicitly argues against using the old, standard Ripley's K without correction for these messy tissue samples. They show that if you use the old method on tissue with tears or degradation, you will overestimate how much the cells are clustering. It's like thinking a crowd is cheering for a band when they are actually just huddled together because it's raining under a broken awning. The authors also argue against using a method that tries to guess the density of the crowd locally (called Kinhom), showing in their simulations that it can actually underestimate the clustering or get the numbers wrong.

How Sure Are They?
The authors are very confident in their math and their simulations, but they are careful with their real-world conclusions.

  • The Math & Simulations: They proved mathematically that their new method works and ran 48 different simulation scenarios (including 1,000 datasets for testing) to show it. In these simulations, KAMP correctly identified when cells were not clustering (keeping the error rate low) and when they were clustering, even when the tissue was torn up. It was also over 100 times faster than the old "shuffle everything" method (which took 562 minutes for their dataset, while KAMP took just 3.6 minutes).
  • The Real Data (Ovarian Cancer): They applied this to a real study of 128 women with high-grade serous ovarian cancer. After cleaning the data, they focused on 71 patients who had enough B cells and macrophages to analyze.
    • They found that using the old method (K) showed no statistically significant link between B cell and macrophage grouping and patient survival.
    • However, using KAMP, they found evidence suggesting that when B cells and macrophages are close together (colocalized) at certain distances (specifically at radii of 5, 10, 20, and from 35 to 120 units), it is associated with better survival.
    • The authors are careful to say this is exploratory. They found that the KAMP-based models predicted survival better (higher c-index) than the old models, but because the sample size was modest (71 patients) and they looked at many different distances, they don't claim this is a final, solved medical breakthrough. They suggest it's a strong hint that needs more study.

The Bottom Line
The paper suggests that the old way of measuring cell groups was often lying to us because it couldn't handle messy, real-world tissue. KAMP fixes this by using the background cells as a reference point, allowing researchers to see the true biological signal. In their ovarian cancer study, this new lens revealed a potential link between specific immune cell teamwork and patient survival that the old lens completely missed. While the math is rock-solid and the simulations are perfect, the real-world medical application is a promising lead that invites further investigation, not a final diagnosis.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →