How small a difference can conductometry establish? An uncertainty budget, resolution curves and the arithmetic cost of four design faults
This paper constructs a GUM-compliant uncertainty budget and resolution curves to demonstrate that common experimental design faults, particularly uncontrolled temperature and statistical missteps, drastically inflate false-positive rates in conductometric studies, ultimately proving that the author's own previously reported data fails to meet the rigorous detection limits required to claim altered physicochemical properties.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Water is often treated as a blank slate in the laboratory, a neutral medium that simply holds other substances. Yet, for decades, researchers have claimed that water can be fundamentally altered by physical or chemical treatments, acquiring new properties that persist long after the treatment ends. The evidence usually cited for these changes is a tiny shift in how well the water conducts electricity. When a treated sample conducts electricity slightly better or worse than an untreated one, scientists have often taken this as proof that the water's structure has changed. The differences reported are small, typically ranging from one to five percent. The scientific community has spent years debating whether known physics could possibly create such changes, but a more basic question has been largely ignored: is the measuring tool itself capable of seeing such a tiny difference, or is the signal just noise?
This question matters because if a measurement cannot reliably distinguish a real change from random fluctuations, the result tells us nothing about the water, regardless of how carefully it was recorded. The paper by Vadim Kovalev tackles this problem not by arguing about the nature of water, but by doing the arithmetic of measurement itself. The author builds a detailed accounting of every possible source of error that could creep into a conductivity experiment. This includes the precision of the balance used to weigh the salt, the accuracy of the glassware used to measure the water, the stability of the instrument over time, and most critically, the temperature. It is a well-established fact that the ability of water to conduct electricity changes by about two percent for every single degree Celsius the temperature rises. This means that a difference of just half a degree between two samples can mimic a one percent change in conductivity, a shift that looks exactly like the effects researchers claim to find.
To understand the limits of these experiments, the author constructed a comprehensive error budget. This is a methodical tally of all the uncertainties involved in a single measurement. The analysis shows that under ideal, controlled conditions, the uncertainty of a single water sample is about 1.3 percent. This means that even with perfect equipment, a single reading could be off by that much simply due to the natural limits of the process. However, the real challenge arises when comparing two groups of water samples. The variability between different batches of water—differences in how they were weighed, mixed, and prepared—is often much larger than the uncertainty of the instrument itself. In many studies, this natural variation between batches is around two percent. When the author calculated how many independent samples would be needed to reliably spot a one percent difference against this background noise, the number was startlingly high: sixty-four separate preparations for each group. Most published studies in this field use only three to five samples, a number far too small to distinguish a real effect from the natural jitter of the experiment.
The paper then examines four common mistakes that researchers make, often without realizing it, which can make a non-existent effect look real. The first is treating multiple readings from the same sample as if they were separate samples. If a researcher measures one cup of water three times and counts those three numbers as three independent data points, they artificially shrink the apparent noise, making a difference look significant when it is not. The second fault is stopping the experiment as soon as a result looks promising. If a researcher keeps adding samples and testing them one by one, stopping the moment they see a "significant" result, they are essentially gambling until they win, which guarantees a false alarm. The third error is measuring all the control samples first and all the treated samples last. If the instrument drifts slightly over time, this fixed order ensures that the drift looks like a difference between the groups. The final and most dangerous fault is failing to control the temperature. Even a tiny, unnoticed difference in room temperature between the two groups can create a false signal that is indistinguishable from a real chemical change.
When the author simulated a study that contained all four of these common faults, the results were dramatic. In a scenario where no real effect existed, the flawed design found a "significant difference" in 78 percent of the runs. This means that a study with these hidden errors is more likely to report a discovery than to correctly report that nothing happened. The author applied these strict criteria to their own previously published work, which had reported conductivity differences of the kind discussed in the literature. The analysis showed that those earlier results did not survive the test; they could not be distinguished from the noise of the experimental design. Consequently, those specific values were withdrawn from the author's research program. The paper does not claim that water cannot be changed, nor does it say that the effects are impossible. It simply establishes that the current methods used to detect them are too blunt to see them.
The study concludes with a set of six requirements for any future experiment to be considered valid. These include recording the exact temperature for every single reading, counting only independent preparations rather than repeated readings, randomizing the order of measurements to cancel out instrument drift, and stating in advance how many samples will be used. Most importantly, researchers must calculate and report the smallest difference their specific design could possibly detect. If a study claims to find a two percent change but its design can only resolve differences larger than six percent, the claim is meaningless. The paper argues that without these records, a result is not weak evidence; it is no evidence at all. The remedy is not to believe or dismiss the findings, but to demand the missing data that would allow the numbers to be judged fairly. By applying these rigorous standards, the scientific community can finally determine which studies are capable of answering the question and which are simply measuring the limits of their own tools.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.