← Latest papers
⚡ electrical engineering

Provisioning to Runtime Optimization of a +100 MW AI Cluster

This paper presents the first end-to-end account of power management for a hyper-scale AI datacenter, detailing the process from early power planning for next-generation accelerators to runtime optimization of a 150 MW facility hosting 83,000 GB200 GPUs.

Original authors: Ehsan K. Ardestani, Leonardo Piga, Jovan Stojkovic, Pavan Balaji, Mustafa Ozdal, Mikel Jimenez Fernandez, Mihaela Dimovska, Luka Tadic, Hao Shen, Devika Vishwanath, Richa Mishra, Melaku Mihret, Valent
Published 2026-05-26
📖 6 min read🧠 Deep dive

Original authors: Ehsan K. Ardestani, Leonardo Piga, Jovan Stojkovic, Pavan Balaji, Mustafa Ozdal, Mikel Jimenez Fernandez, Mihaela Dimovska, Luka Tadic, Hao Shen, Devika Vishwanath, Richa Mishra, Melaku Mihret, Valentin Andrei, Mauricio Cespedes, Julien Prigent, James Monahan, Tyler Graf, Bin Li, Charles Marquez, Shobhit Kanaujia, Kaushik Veeraraghavan, Chunqiang Tang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to build the world's largest, most powerful supercomputer to solve the hardest problems in Artificial Intelligence. You have bought thousands of the fastest computer chips (GPUs) money can buy. But there's a catch: you don't have enough electricity to run them all at full speed.

In fact, the paper argues that running out of power is now a bigger problem than running out of computer chips. It's like having a fleet of 100 race cars but only enough gas to fill half of them.

This paper describes how Meta (the company behind Facebook) solved this problem for a massive 150-megawatt data center. They didn't just plug everything in and hope for the best; they treated electricity like a precious, limited resource that needed to be managed in three distinct stages.

Here is the story of how they did it, using simple analogies.

The Three Stages of Power Management

The authors break the process down into three phases: Planning, Checking, and Running.

Phase 1: Planning (The "Budgeting" Phase)

Before the computers even arrived, the team had to figure out how many they could fit into their fixed electricity budget.

  • The Old Way: Usually, engineers look at a chip's "maximum speed" (its Thermal Design Power, or TDP) and assume that's what it will always use. If a chip is rated for 1,200 Watts, they plan for 1,200 Watts.
  • The Meta Way: They realized that if they forced every chip to run at its absolute maximum, they could only fit a few chips in the building. Instead, they asked: "What if we run all the chips slightly slower?"
  • The Analogy: Imagine a buffet with a fixed amount of food.
    • Strategy A: Give everyone a massive, 10-course meal. You can only feed 50 people.
    • Strategy B: Give everyone a slightly smaller, 8-course meal. Now you can feed 60 people.
    • The Result: Even though each person got slightly less food, the total amount of food eaten by the group is higher, and more people are fed.
  • The Discovery: They found that by lowering the power limit of each GPU from 1,200 Watts to about 1,000 Watts, they could fit many more GPUs into the data center. The slight drop in speed for each individual chip was more than made up for by having more chips working together. This gave them a 10% boost in total computing power.

Phase 2: Checking (The "Reality Check" Phase)

Once the computers were installed, the team realized their plans didn't match reality perfectly.

  • The Problem: The computers' internal power meters (PSUs) were lying. They were being overly cautious, reporting that the machines were using more power than they actually were. It's like a car's fuel gauge that always says "Empty" when you still have a quarter tank left.
  • The Fix: They compared the internal meters to the building's main power sensors (which are very accurate). They found that if they ignored the "worst-case" readings and looked at the "average" usage (specifically the 70th percentile), they could trust the numbers more.
  • The Result: Because they stopped trusting the overly cautious internal meters, they realized they had a little bit of "hidden" power they weren't using. They turned the power back up slightly (from 1,000W to 1,020W), squeezing out another 2% performance boost.

Phase 3: Running (The "Traffic Cop" Phase)

Now the data center is running live AI training jobs. The problem is that AI workloads are "synchronous," meaning all the chips have to wait for each other. If one chip is slow, the whole job slows down.

  • The Problem: Sometimes, the power grid gets a sudden spike. If the system detects a spike, the old way was to immediately shut down or slow down one specific chip to save power. But in a synchronized race, if you slow down one runner, the whole team loses.
  • The Solution (Dimmer): They built a smart system called "Dimmer." Instead of punishing one chip, it gently lowers the power of every chip in the job by a tiny, equal amount.
  • The Analogy: Imagine a group of runners in a relay race. If the track gets slippery (power limit), instead of tripping one runner (which stops the whole team), the coach tells everyone to run just a tiny bit slower. The team stays together, and the race continues without stopping.
  • The Result: This prevents the system from crashing and allows them to use the last few percent of "stranded" power that was previously wasted. This added another ~2% boost.

Other Key Insights

  • The "Bottleneck" Effect: The data center is a mix of different hardware (computers, cooling, network switches). Sometimes, the power limit isn't set by the computers, but by the cooling system or the network cables. If one part of the chain is weak, it limits the whole chain.
  • Power Swings: When AI chips work, they sometimes pause for a split second to talk to each other, causing the power usage to dip and then spike. The team created a "software smoother" that fills these tiny gaps with dummy work, keeping the power usage steady like a calm river instead of a choppy sea. This protects the building's electrical grid from shocks.
  • Newer is Better: They compared their new chips (GB200) to older ones (H100). The new chips are so efficient that even though they use less power per chip, the new cluster is nearly twice as fast as the old one would have been in the same building.

The Bottom Line

By treating electricity as a flexible resource rather than a fixed limit, and by managing it carefully from the planning stage to the daily operations, Meta managed to get roughly 14% more computing power out of their 150-megawatt data center without building a single new wall or buying a single new wire.

They didn't find a magic battery; they just learned how to use the electricity they already had much more wisely.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →