← Latest papers
📊 statistics

Conflict inherits the frame: bin geometry and BPA construction in evidential fusion of quantised continuous evidence

This paper demonstrates that in Dempster-Shafer evidential fusion of quantized continuous data, the geometry of the binning frame artificially confounds conflict measures with evidence position, leading to a proposed excess-conflict statistic to correct this bias while revealing that standard combination rules often reduce to Bayesian products with negligible efficacy differences in predicting software vulnerability exploitation.

Original authors: Elena Udrescu, Alexandru Udrescu, Ana-Maria Suduc, Mihai Bîzoi

Published 2026-09-07
📖 6 min read🧠 Deep dive

Original authors: Elena Udrescu, Alexandru Udrescu, Ana-Maria Suduc, Mihai Bîzoi

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine trying to decide how dangerous a software flaw is by asking three different experts. Each expert gives a number, but they rarely agree. Sometimes one says a flaw is a minor nuisance, while another calls it a critical threat. To make sense of this disagreement, researchers often use a mathematical system designed to combine uncertain opinions into a single, clearer picture. This system works by taking a continuous range of possibilities—like a scale from zero to ten—and chopping it into distinct buckets, such as "low," "medium," "high," and "critical." The idea is that once the numbers are sorted into these buckets, the system can calculate how much the experts disagree and use that disagreement to judge the risk. For decades, this method has been a standard tool in fields ranging from medical diagnosis to security analysis, trusted to turn messy, conflicting data into a reliable verdict.

However, a new study suggests that the way these buckets are sized might be secretly distorting the results. Researchers Elena Udrescu, Alexandru Udrescu, Ana-Maria Suduc, and Mihai Bîzoi from Valahia University of Targoviste investigated this process using real-world data on software vulnerabilities. They looked at thousands of records where different organizations had assigned severity scores to the same security flaws. Their goal was to see if combining these scores using the standard mathematical rules actually helped identify which flaws were being actively exploited by hackers, or if the method was simply creating an illusion of insight. What they found was that the shape of the buckets themselves was doing most of the work, often leading the system to report high disagreement in places where the experts actually agreed, and vice versa.

The researchers began by examining the data from the Common Vulnerability Scoring System, which is the standard language used to describe software flaws. This system divides scores into four bands: low, medium, high, and critical. The problem is that these bands are not the same size. The "low" band covers a wide range of numbers, while the "critical" band is very narrow. When the researchers fed the data into the standard combination system, they discovered a hidden flaw in the logic. The system was treating the narrowness of a bucket as a sign of conflict. If two experts gave scores that fell near the edge of a narrow bucket, the system calculated a high level of disagreement, even if the experts were actually very close in their assessment. Conversely, if the scores fell in the middle of a wide bucket, the system reported low disagreement, even if the scores were quite far apart. The geometry of the buckets was overriding the actual opinions of the experts.

To prove this, the team created a new way to measure disagreement that stripped away the influence of the bucket sizes. They compared the standard method against this corrected version. The results were striking. The standard method suggested that experts disagreed more about the most severe flaws than the less severe ones, a pattern that seemed to make sense on the surface. But the corrected method showed the exact opposite: experts actually disagreed more about the moderate flaws and agreed more often on the critical ones. The original method had been misled by the fact that the critical category was so narrow that any small difference in scores pushed the calculation into a state of high conflict. By removing this geometric bias, the researchers showed that the standard tool was reporting the shape of the categories rather than the true state of the evidence.

The study also looked at how the experts' scores were converted into the system's input. A common practice is to take a specific score and assign it entirely to the single bucket it falls into, ignoring the fact that the score is a point on a continuous line. The researchers found that this "crisp" approach reduced the entire concept of disagreement to a simple yes-or-no flag. If the experts landed in the same bucket, the system saw zero conflict. If they landed in different buckets, it saw maximum conflict. This binary view missed all the nuance in between. When the researchers used a "graded" approach, which spreads the score across neighboring buckets to reflect uncertainty, the system regained its ability to see subtle differences. However, even with this improvement, the underlying geometry of the buckets still skewed the results unless the new correction was applied.

When the team tested whether combining the experts' views actually helped predict which vulnerabilities were being exploited in the real world, the results were sobering. They analyzed over 13,000 pairs of assessments from different organizations, checking if the combined score was better at identifying the 70 known exploited flaws than simply trusting the single best expert. The answer was no. No matter how they combined the scores, whether using the standard rules or the new corrections, the fused result did not outperform the single best source. The only method that did slightly better was a simple trick of taking the lower of the two scores, but the researchers cautioned that this might be an artifact of how the data was collected rather than a true insight. In fact, the study found that the standard combination rules were mathematically equivalent to a simple average in many cases, meaning the complex machinery was not adding any real value.

A significant portion of the study also uncovered a data issue that had been hiding in plain sight. Many of the "second opinions" in the database were not actually independent assessments from the organizations that originally reported the flaws. Instead, they were automated updates from a government agency that often just copied the original score. When the researchers filtered these out, the disagreement between the true independent sources jumped significantly. This revealed that previous studies might have underestimated how much experts actually disagree because they were inadvertently comparing a source to a copy of itself.

The researchers concluded that the tools used to fuse uncertain evidence are far more sensitive to the design of the categories than previously thought. They showed that the standard measure of conflict is often a mirror of the bucket sizes rather than a reflection of human disagreement. While the complex mathematical rules for combining evidence did not improve the ability to predict real-world danger in this specific case, the study provided a vital correction for how to read the data. By adjusting for the size of the buckets, analysts can now see the true level of disagreement between sources. The lesson is that in the quest to combine uncertain information, the way we define the categories matters just as much as the information itself. Without paying attention to the shape of the buckets, we risk seeing patterns that are merely the shadow of our own measurement tools.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →