← Latest papers
⚛️ high-energy experiments

Green BOA: Determining the environmental break-even point for ML-based data compression

This paper evaluates the environmental sustainability of ML-based data compression by calculating the carbon-equivalent break-even point where energy savings from reduced storage offset the carbon footprint of training and inference infrastructure.

Original authors: Caterina Doglioni, Akshat Gupta, Thomas Elliott, Hanzila Hussain, Sanjiban Sengupta

Published 2026-08-21
📖 4 min read🧠 Deep dive

Original authors: Caterina Doglioni, Akshat Gupta, Thomas Elliott, Hanzila Hussain, Sanjiban Sengupta

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the vast, humming halls of modern science, experiments like those at the Large Hadron Collider are generating data at a pace that threatens to overwhelm our storage capabilities. Scientists are recording several exabytes of information, a volume so immense that it demands not just massive budgets but also significant environmental resources to keep it safe. As the world seeks to reduce its carbon footprint, the energy required to store this digital mountain has become a pressing concern. Traditionally, researchers have relied on standard methods to shrink these files, but a new wave of technology uses machine learning to compress data even further. However, these smart algorithms are computationally hungry; they require powerful computers to learn how to squeeze the data and then to run the compression. This creates a complex trade-off: does the energy spent teaching a computer to compress data actually save more energy than it costs to store the larger, uncompressed files?

A team of researchers at the University of Manchester set out to answer this question by calculating the environmental break-even point for a specific machine learning compression tool called BOA. They wanted to know exactly how much data would need to be processed before the carbon emissions from training and running the machine learning model were offset by the carbon savings from needing fewer hard drives and tape cartridges. The study focused on a lossless compression algorithm, which means it shrinks files without losing any information, a critical requirement for scientific data. The researchers compared the carbon cost of the electricity used to train the model and run it on a graphics processing unit against the carbon cost of manufacturing and operating the physical storage devices that would be saved if the data were compressed.

The investigation revealed that the answer depends heavily on where the computing takes place. The carbon footprint of electricity varies significantly from country to country, depending on how much of the power grid comes from clean sources versus fossil fuels. The researchers found that the location of the data center is the most sensitive factor in determining whether the machine learning approach is environmentally beneficial. If the compression is performed in a region with a carbon-intensive energy mix, the cost of training the model can easily outweigh the savings from reduced storage. However, in regions with cleaner energy, the balance shifts, and the machine learning method can become the greener option. The study also highlighted that while machine learning compression often achieves better shrinking ratios than standard algorithms, it does so at a slower speed and with higher energy use per unit of data processed.

To reach these conclusions, the team simulated the entire process using a specific graphics card, the Nvidia T4, and tracked the energy consumption of the central processor, the graphics card, and the memory. They modeled the carbon impact of storing data for five years on both hard disk drives and magnetic tape, which is commonly used for long-term archives. They calculated the emissions from manufacturing these devices and the electricity needed to keep them running. The results showed that magnetic tape, while slower to access, has a lower carbon footprint than hard disk drives. This means that for the machine learning compression to be worth the environmental cost, the savings must be substantial enough to displace a significant amount of physical storage. The researchers noted that their calculation represents a worst-case scenario for energy use, as they assumed the model was trained on the entire dataset from scratch, whereas in practice, the model could likely be trained on a smaller sample and then applied to much larger amounts of data.

The findings suggest that there is no single, universal answer to whether machine learning compression is environmentally friendly. Instead, the decision must be made based on the specific context of the data center's location and the type of storage being replaced. For the machine learning approach to justify its environmental cost, the volume of data being compressed must be large enough to offset the initial energy investment in training the model. The team emphasized that while the current throughput of the algorithm is lower than standard methods, improving this speed could make the technology more viable. Ultimately, the study serves as a proof of concept, demonstrating that the environmental benefits of advanced compression are not automatic but depend on a delicate balance between computational energy and storage savings. As science continues to generate more data, understanding this break-even point will be essential for ensuring that the pursuit of knowledge does not come at an unsustainable cost to the planet.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →