← Latest papers
📊 statistics

Do not log(x+1) transform: the bias and alternatives

This paper argues that the widely used log(x + 1) transformation introduces substantial bias and misleading inferences, particularly for exponentially distributed data, and recommends abandoning it in favor of alternative numerical methods like Generalized Linear Models (GLM).

Original authors: Vasco Vieira, Cátia Bartilotti, Jorge Lobo-Arteaga, Rafael Santos, David Leitão-Silva, Arthur Veronez, Joana Neves, Ana Brito, Rui Cereja, Diogo Paulo, Francisco Leitão

Published 2026-08-20
📖 6 min read🧠 Deep dive

Original authors: Vasco Vieira, Cátia Bartilotti, Jorge Lobo-Arteaga, Rafael Santos, David Leitão-Silva, Arthur Veronez, Joana Neves, Ana Brito, Rui Cereja, Diogo Paulo, Francisco Leitão

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the natural world, measurements often behave in ways that defy simple straight lines. When scientists count the number of fish in a pond, measure the height of a forest canopy, or track the temperature of the ocean, the data rarely spreads out evenly. Instead, small values are common, while large values are rare, creating a lopsided shape that makes standard mathematical tools struggle to find patterns. To fix this, researchers have long relied on a mathematical trick called a logarithmic transformation. This process reshapes the data, compressing the huge numbers and stretching out the tiny ones, turning a jagged, uneven landscape into a smooth, flat plain where relationships become clear and easy to study. However, this trick hits a hard wall when the data contains zeros. Since the mathematical operation cannot handle a zero, scientists have traditionally added a small, arbitrary number to every single measurement before applying the trick. For decades, the most common choice has been to simply add one to every value, a method known as the log(x+1) transformation. It seemed like a harmless, practical workaround, allowing researchers to include zero counts in their studies without discarding valuable information.

A team of researchers from Portugal and the University of the Algarve has now shown that this long-standing habit is fundamentally flawed. By testing the method with both computer-generated numbers and real-world ecological data, they demonstrated that adding a constant number to the data does not just fix the zero problem; it actively distorts the truth. The study reveals that this simple addition introduces a massive bias, warping the relationships between variables and leading scientists to draw incorrect conclusions about how nature works. The researchers found that the distortion is not a minor glitch but a severe error that changes the slope of relationships, hides the true spread of data, and can even reverse the direction of a trend. Their work suggests that the scientific community needs to stop using this specific fix immediately and look for more robust methods that do not rely on adding arbitrary numbers to the data.

To prove their point, the researchers first created a perfect, artificial dataset where the relationship between two variables was known exactly. They designed a scenario where one variable grew exponentially based on the other, a pattern common in nature, but they included a few zero values to mimic real-world conditions. When they applied the standard log(x+1) transformation to this perfect data, the resulting model was wildly inaccurate. The curve that should have been smooth and predictable became bent and broken. The researchers found that the choice of the number one was the culprit. Because one is a large number compared to the tiny values often found in nature, adding it completely overwhelmed the actual data. For very small measurements, the added one became the dominant factor, making the tiny differences between them disappear. For larger measurements, the addition had a different effect, creating a mismatch that the mathematical model could not correct. Even when the researchers tried using much smaller numbers instead of one, the problem persisted. If they chose a number too small, the zeros in the data were transformed into values so far away from the rest of the data that they created a new kind of distortion, pulling the analysis off course.

The team then moved beyond artificial numbers to examine real ecological studies, including one about seagrass flowering and another about water quality in Portuguese estuaries. In the seagrass study, researchers were trying to understand how rising ocean temperatures triggered a massive bloom of flowers. The relationship between the heat and the flowering was clearly non-linear, growing faster as the heat increased. When the original researchers used the log(x+1) method to analyze this, the resulting curve failed to capture the true intensity of the response. It flattened out the peak and missed the rapid acceleration of the bloom. In contrast, when the new team re-analyzed the same data using different methods that did not involve adding a constant, the true, sharp relationship between heat and flowering emerged clearly. Similarly, in the estuary study, scientists were looking at the ratio of above-ground to below-ground plant biomass to determine the health of the seagrass. Because some plants were growing while others were shrinking, the data contained values both above and below one. The log(x+1) method treated these values unevenly, compressing the healthy, growing plants and exaggerating the struggling ones, which led to a distorted view of the ecosystem's health. The researchers showed that this distortion was not just a statistical curiosity; it changed the interpretation of whether the seagrass was thriving or dying.

The study also tested whether other common mathematical tricks could solve the problem. They tried using square roots, cube roots, and other complex adjustments that are often used to smooth out uneven data. While some of these alternatives worked better than the log(x+1) method in specific situations, none of them could fully fix the problem when the data contained zeros and followed an exponential pattern. The researchers found that when data behaves in an exponential way, no amount of reshaping the numbers with a simple formula can make it fit the standard linear models without introducing error. However, they clarified that this does not mean every transformation is doomed to fail; for datasets that do not strictly follow an exponential model, other methods like the Box-Cox transformation can be successful. In these specific exponential cases, the paper concludes that the best solution is not to force the data into a straight line at all. Instead, scientists should use statistical models designed specifically for non-linear data, such as generalized linear models, which can handle the zeros and the exponential growth directly without needing to add arbitrary numbers.

The implications of this finding are significant for environmental science, where data often includes zeros because a species might be absent from a site or a pollutant might be undetectable. For years, the log(x+1) transformation has been the default tool for handling these zeros, used in countless studies on biodiversity, climate change, and pollution. The researchers argue that this widespread use has likely led to a vast number of studies producing biased results, where the strength of a relationship is underestimated or overestimated, or where the direction of a trend is misunderstood. The bias is not subtle; it is large enough to change the outcome of a study. The authors emphasize that while the log(x+1) method is easy to use, it is not a valid scientific solution. They urge the scientific community to abandon this practice and to adopt more rigorous methods that respect the true nature of the data. By doing so, researchers can ensure that their conclusions about the natural world are based on reality, not on the mathematical artifacts of a convenient but flawed shortcut.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →