← Latest papers
📊 statistics

Geometric mean-based pairwise comparison method with the reference values -- statistical approach

This paper presents a statistical approach to the pairwise comparison method using reference values and geometric means to simultaneously capture inconsistency and preference distances, while defining interpretable indicators to measure the quality of the resulting weight vector.

Original authors: Konrad Kułakowski, Kamil Pustelnik, Jacek Szybowski

Published 2026-07-31
📖 6 min read🧠 Deep dive

Original authors: Konrad Kułakowski, Kamil Pustelnik, Jacek Szybowski

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are the captain of a school club, and you need to pick the best three activities for the upcoming year. You ask your friends to vote, but instead of just picking a favorite, you ask them to compare every single activity against every other one. "Is 'Hiking' twice as fun as 'Baking'?" "Is 'Coding' three times better than 'Art'?" This method of comparing things in pairs is a classic tool used by experts to make tough decisions, from choosing a new city mayor to picking the best medical treatment. It's like trying to figure out the exact height of a mountain by only measuring the slope between two trees at a time.

The problem is that humans aren't perfect calculators. One friend might say Hiking is better than Baking, but then say Baking is better than Coding, and Coding is better than Hiking. That's a logical loop, or "inconsistency." For decades, mathematicians have had ways to smooth out these loops and find a final ranking, but they mostly treated the errors as messy noise to be ignored. They could tell you what the ranking was, but they couldn't easily tell you how sure they were about it. It was like getting a weather forecast that says "It will rain" without saying if there's a 51% chance or a 99% chance. This paper steps into that gap, treating those messy human comparisons not just as errors, but as data points in a statistical experiment, allowing us to measure the confidence in our decisions just like a scientist measures the confidence in a lab result.

The Paper's Big Idea: Turning Guesswork into a Math Lab

The authors of this paper, Konrad Kułakowski, Kamil Pustelnik, and Jacek Szybowski, are looking at a specific decision-making method called "Heuristic Rating Estimation" (HRE). Imagine you have a list of unknown items (like new video games) and a few known items (like your favorite classic game, which you know is a perfect 10/10). You ask an expert to compare the new games against the classics and against each other. The goal is to figure out the "weight" or score of the new games.

Traditionally, this is done using a geometric mean (a specific type of average that handles ratios well). The authors realized something brilliant: this whole process is actually a linear regression problem. That's a fancy way of saying it's the same math used to draw a "line of best fit" through a scatter plot of dots.

Here is the analogy: Imagine you are trying to guess the weight of several mystery boxes. You have a few known weights (reference values) on a scale. You don't know the mystery weights, but you have a friend who gives you clues like, "Box A is about half as heavy as Box B," or "Box C is twice as heavy as the 5kg reference weight." Your friend isn't perfect; sometimes they guess "twice as heavy" when it's actually "1.9 times."

The paper suggests we shouldn't just average these guesses. Instead, we should treat every single comparison as a "measurement" in a science experiment. The "unknown weights" are the variables we are trying to solve for, and the "friend's guesses" are our data points. By using the tools of statistics (specifically, the least squares method), we can find the most likely weights for the mystery boxes.

What They Found: Confidence Intervals and "Tie" Clusters

The real magic of this paper isn't just finding the weights; it's about knowing how much to trust them. Because they treated the problem as a statistical experiment, they could calculate things that were previously impossible:

  1. Confidence Intervals: Just like a poll might say "50% support, plus or minus 3%," this method can tell you that a game's score is likely between 0.15 and 0.25. If the range is huge, you know the data is shaky. If it's tiny, you can trust the ranking.
  2. Probability of Rank Reversal: This is the most exciting part. The authors calculated the exact probability that the order of two items is correct. For example, they found that for two specific items in their example, there was a 99.46% chance that one was better than the other. But for another pair, the chance was only 58.18%. That's barely better than flipping a coin!
  3. The "Tie" Solution: This is where the paper gets really practical. If the probability that Item A is better than Item B is low (say, below 75%), the authors suggest we shouldn't force a ranking. Instead, we should group them into a "tie cluster." Imagine a race where two runners are so close that the finish line camera can't tell who won. Instead of declaring a fake winner, you say, "They tied." The paper proposes an algorithm to automatically group these uncertain items together and give them an average score. This prevents the decision-maker from making a bold choice based on weak evidence.

What They Don't Claim (and What They Rule Out)

It is important to note what this paper does not do. It doesn't claim to fix the human brain or make experts stop making mistakes. The "noise" or inconsistency in the data is still there; the paper just measures it better.

The authors explicitly argue against the idea that we should just ignore the uncertainty or rely solely on traditional "inconsistency indices" (old math tools that give a single number to say how messy the data is). They show that a single number isn't enough because two different pairs of items can have the same "messiness" score, but one pair might be clearly separated (easy to rank) while the other is almost identical (hard to rank). Their new method captures both the "messiness" and the "distance" between the items.

They also don't claim this is a magic bullet for every situation. Their results are based on mathematical proofs and simulations (like the example with 7 alternatives they worked through). They suggest that in group settings, where multiple experts provide data, this method can be even more powerful because it can combine everyone's "measurements" to get a clearer picture, even if each expert missed some comparisons.

The Takeaway

In simple terms, this paper takes the messy, subjective world of "I think this is better than that" and turns it into a rigorous, data-driven science. It gives us a way to say, "We are 99% sure A is better than B, but we are only 58% sure about C and D, so let's call C and D a tie." It turns a rigid list of rankings into a flexible, honest representation of what we actually know, helping decision-makers avoid the trap of pretending to be more certain than they really are.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →