← Latest papers
💻 computer science

The Hidden Cost of Thinking: Energy Use and Environmental Impact of LMs Beyond Pretraining

This paper reveals that the environmental impact of modern language model development is dominated by post-training experimentation and complex reasoning pipelines—accounting for 82.2% of total compute and making reasoning models 17 times more expensive to post-train than instruction-tuned ones—highlighting a critical, unreported gap in current environmental reporting standards.

Original authors: Jacob Morrison, Noah A. Smith, Emma Strubell

Published 2026-05-05
📖 4 min read☕ Coffee break read

Original authors: Jacob Morrison, Noah A. Smith, Emma Strubell

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Picture: The Iceberg of AI Costs

Imagine you are building a massive, high-tech factory to bake the perfect cake (a smart AI model). For years, everyone has only been counting the electricity used to bake the final cake and putting that number on the label.

This paper argues that this is like looking at an iceberg and only counting the tip above the water. The authors went underwater to measure the entire iceberg. They found that the "baking" of the final model is actually a tiny part of the story. The real energy, carbon emissions, and water usage happen during the messy, experimental process of figuring out how to bake the cake in the first place.

1. The "Thinking" Models Are Energy Hungry

The researchers studied a new family of AI models called Olmo 3. They built two types:

  • The "Instruct" Model: Follows orders well (like a helpful assistant).
  • The "Think" Model: Solves complex problems by "thinking" through steps (like a mathematician working out a proof).

The Analogy:
Imagine the "Instruct" model is a sprinter. It runs a short race and finishes. The "Think" model is a marathon runner who stops to solve puzzles along the way.

  • The Finding: Training the "Think" model required 17 times more energy than the "Instruct" model.
  • Why? The "Think" model has to generate massive amounts of "practice runs" (called rollouts) to learn how to reason. It's like a student who has to write out 100 practice essays just to learn how to write one good one. The paper calls this the "cost of thinking."

2. The "Hidden" Cost: Trial and Error

The biggest surprise in the paper is about development costs.

  • Old Belief: We thought most energy was used for the final training run (the "final exam").
  • New Reality: 82% of the total energy was used for experimentation. This includes failed attempts, testing different settings, and trying out different data mixes before the final model was even built.

The Analogy:
Think of it like a chef trying to create a new signature dish.

  • The Final Dish: The one served to the customer (the final model).
  • The Development: The chef burning 100 pots of soup, throwing away 50 batches of cookies, and tasting 20 different spice blends to get it right.
  • The Paper's Point: If you only count the electricity used to cook the one final dish served to the customer, you are ignoring the 100 pots of soup that were burned in the kitchen. The paper found that for modern AI, the "burned soup" (failed experiments) costs 5 times more than the final dish.

3. The Water Problem: It's Not the Data Center's Fault

The paper also looked at water usage. They found the process used a huge amount of water (enough for one person to drink for 140 years).

The Analogy:
Many people think AI uses water because the computers get hot and need to be cooled down with water (like a car radiator).

  • The Reality: The data center used a "closed-loop" system, meaning it didn't lose water to cooling.
  • The Real Culprit: The water was used to make the electricity that powered the computers. Power plants (coal, gas, nuclear) use massive amounts of water to generate the power needed to run the AI.
  • The Lesson: If you want to save water, you don't just need to build better computers; you need to change how we generate the electricity that powers them.

4. The Total Bill

When the researchers added up everything—the final training, the failed experiments, the data generation, and the hardware manufacturing—the total cost for building the Olmo 3 models was:

  • Energy: About 12.3 million kilowatt-hours (enough to power 847 US homes for a whole year).
  • Carbon: About 4,251 tons of CO2.
  • Water: About 15,887,000 liters.

Why This Matters

The paper concludes that the way we currently report AI's environmental impact is broken. We are only reporting the "final run" cost, which makes the problem look much smaller than it actually is.

As AI models get smarter and require more "thinking" (reasoning), and as the development process gets more complex, these hidden costs will skyrocket. The authors are asking developers to stop hiding the "burned soup" and start reporting the full cost of the entire kitchen, not just the final meal.

In short: Building smart AI is expensive, but the process of figuring out how to build it is even more expensive, and the "thinking" models are the most energy-hungry of all.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →