Data Center Electricity Consumption: The Forecast Was Wrong Last Time

The number everybody quotes is 945 terawatt-hours by 2030, which the International Energy Agency described as slightly more than Japan’s total electricity consumption today. Data centers used about 415 terawatt-hours in 2024, roughly 1.5 percent of global electricity. The projection more than doubles that in six years.

It is a good number. It comes from a credible institution, it is built on a documented methodology, and it has been updated since, with the agency’s April 2026 revision moving the central case to 950. It is also a forecast, produced by people who had to make assumptions about adoption rates nobody can observe yet, and there is a specific reason to hold it loosely.

In 2007, the Environmental Protection Agency reported that American data centers had roughly doubled their electricity use between 2000 and 2006 and warned the trajectory was unsustainable. The extrapolation was reasonable, the underlying data was sound, and the forecast was substantially wrong. Consumption did not continue doubling. It flattened for the better part of a decade while the amount of computing being done grew by a factor of several hundred percent.

Understanding why that happened, and whether the same mechanism is available now, is the single most useful thing anybody can do before forming a view about data center electricity consumption. It is also the closest thing this subject has to a controlled experiment, since the forecast, the mechanism that falsified it, and the eventual measurement are all documented.

Where the buildings came from

The lineage runs through four distinct architectures, and each transition changed the energy profile in a way the previous generation would not have predicted.

The mainframe room came first. A single large machine in a purpose-built space, with raised floors to route cabling and enormous air conditioning because the machine rejected all its energy as heat into a room that had to stay cool for the equipment rather than for the people. Every design element of a modern facility traces to that room, including the raised floor, which persisted for decades after the cabling reason for it disappeared and was repurposed as a plenum for cold air. Institutional inertia in building design is a real force, and a good deal of data center convention is archaeology rather than engineering.

Client-server distribution came next and was an efficiency disaster nobody recognized at the time. Organizations bought many small servers instead of one large one, deployed them in closets and converted offices, and ran them at utilization rates in the single digits because each application got its own machine. The industry spent the 1990s accumulating an installed base of idle hardware drawing full power, since a server at five percent utilization draws well over half its peak power doing essentially nothing. That gap between idle and peak draw is the single largest source of waste in the history of the industry, and closing it is what virtualization actually did.

Colocation and the internet buildout produced the first purpose-built facilities at scale, and the first serious attention to the cost of powering them. Then virtualization arrived and allowed many logical servers to share one physical machine, which raised utilization dramatically and is the single largest efficiency gain in the history of the industry. It also established the commercial logic that made cloud possible, since a provider that can pack many customers onto shared hardware has a cost structure no single enterprise can match.

Hyperscale is the current architecture and it changed the economics completely. A handful of operators running enormous facilities with custom hardware, custom power distribution, and the capital to optimize at a scale no enterprise data center could justify. That consolidation is the reason the 2007 forecast failed.

The AI facility is arguably a fifth architecture rather than a variation on the fourth, and the distinction is physical. A hyperscale cloud hall runs racks at five to fifteen kilowatts, air cooled, with workloads that can be scheduled around each other and moved between sites. An AI training hall runs racks above a hundred kilowatts, liquid cooled by necessity, with a workload that wants every accelerator running simultaneously on the same job. The thermal and electrical requirements are different enough that the buildings are not interchangeable, which is why so much of the current construction is greenfield rather than retrofit.

Why the last forecast was wrong

The mechanism deserves precision because the current argument turns on whether it still operates.

Between 2010 and 2018, global data center compute output grew by roughly 550 percent while electricity consumption grew by around six percent. Storage capacity grew by a factor of twenty-five. Energy use stayed nearly flat. That finding, from a recalibration of global data center energy-use estimates published in Science, is the most important fact in the history of this subject, and it was produced by a group of researchers who had spent a decade building bottom-up inventories of what was actually installed rather than extrapolating from service demand.

Four things produced it. Server efficiency improved continuously, with each hardware generation delivering more computation per watt. Virtualization raised utilization from single digits toward a large fraction of capacity, which meant fewer physical machines doing more work. Workloads migrated from inefficient enterprise closets into hyperscale facilities running at far better efficiency. And facility overhead collapsed, with power usage effectiveness improving from typical values approaching 2.0 in the mid-2000s toward figures near 1.1 at the best hyperscale sites.

That last one is worth dwelling on. A PUE of 2.0 means every watt reaching a processor requires another watt for cooling, power conversion, and lighting. A PUE of 1.1 means ten percent overhead. Going from 2.0 to 1.1 nearly halves total facility consumption for identical computing, and the industry did that across roughly fifteen years.

Worth noting what drove it, since the mechanism was economic rather than regulatory. Electricity is a substantial operating cost for a facility running continuously, and at hyperscale it is large enough that a fractional efficiency improvement justifies dedicated engineering teams. The operators that consolidated the workloads were exactly the operators with the scale to fund that engineering, which is why consolidation and efficiency arrived together rather than independently.

So the forecast failed not because the demand projection was wrong but because the efficiency assumption was, and the people making the projection had no way to know that a set of unrelated improvements would compound into an offsetting force of exactly the right magnitude.

Why this time might be different

The argument that the historical pattern will not repeat is stronger than the reflexive skepticism allows, and it rests on the observation that all four of those levers are largely spent.

Facility overhead cannot repeat its trick. Going from 2.0 to 1.1 saved forty-five percent. Going from 1.1 to 1.0 saves nine percent, and 1.0 is thermodynamically unreachable. The largest single efficiency gain in the industry’s history is exhausted by arithmetic.

Virtualization cannot repeat either. Utilization improved from single digits to high levels, and there is no second order of magnitude available.

The consolidation migration is largely complete. Most workloads that were going to move from enterprise closets to hyperscale facilities have moved, which means the population of grossly inefficient installations available to be improved has shrunk.

And per-transistor efficiency gains slowed with the end of Dennard scaling. Chips still improve, and the improvement now comes substantially from specialization and packaging rather than from the automatic power reduction that accompanied each process node for decades, which is why advanced packaging rather than transistor scaling has become the industry’s binding constraint.

Against those, the counter-argument is that the efficiency work has simply moved up the stack. Algorithmic improvements, model architecture changes, quantization to lower-precision formats, and inference optimization are all reducing energy per unit of useful output, and several of them are producing gains larger than any hardware generation delivered. The IEA’s own lower case assumes something along these lines, noting that similar to trends in the early 2010s, efficiency improvements could offset most of the impact of increased utilization.

Which is the honest state of the argument. The hardware and facility levers that flattened the last curve are exhausted. Software and algorithmic levers of unknown magnitude have replaced them. Nobody knows the size of the second set, and both camps are extrapolating from a base of one historical episode.

There is one more asymmetry worth registering. The hardware and facility gains were largely automatic, in the sense that buying newer equipment delivered them without anybody changing what they were computing. Algorithmic gains are not automatic; they require the workload itself to change, which means they accrue to operators who adopt them and not to the installed base. An efficiency improvement that requires re-engineering a model reaches the fleet far more slowly than one that arrives with the next server refresh, and the fleet is what determines aggregate consumption.

What the data center electricity consumption numbers say

Stating the figures with their units, stages, and dates attached is most of the analytical work here.

Global data center electricity consumption reached approximately 415 terawatt-hours in 2024, around 1.5 percent of world electricity. The IEA base case projects roughly 945 terawatt-hours by 2030, just under three percent of global consumption, with the April 2026 update at 950. Growth runs around fifteen percent annually across that period, more than four times faster than electricity demand growth from all other sectors combined.

The concentration matters as much as the total. The United States, Europe, and China account for about eighty-five percent of current consumption. The United States and China together account for nearly eighty percent of projected growth, with American consumption rising by around 240 terawatt-hours and Chinese by around 175. In the United States, data centers account for nearly half of all electricity demand growth to 2030, and by the end of the decade the country is projected to consume more electricity for data centers than for aluminium, steel, cement, chemicals, and all other energy-intensive goods combined. That comparison is the most striking single line in the IEA report and it is a statement about American deindustrialization as much as about computing.

AI is the driver and it is not the whole. AI servers accounted for about twenty-four percent of server electricity demand and fifteen percent of total data center energy in 2024, with AI-focused consumption projected to triple or quadruple by 2030 depending on the case. Data center electricity consumption grew seventeen percent in 2025 while AI-focused facility consumption rose about fifty percent.

And a single facility is now a meaningful unit. The IEA’s analysis puts a hyperscale AI-focused data center at roughly the electricity consumption of a hundred thousand households. One building. That figure is the one worth carrying into any local argument, because a national percentage is abstract and a comparison to a hundred thousand households is what a county commission is actually being asked to approve.

The three percent problem

Here is the figure that reframes the entire public argument, and it cuts against alarm in one direction and complacency in the other.

Even in the doubling scenario, data centers reach just under three percent of global electricity consumption in 2030. That is not an existential share. Industry, transport, buildings, and agriculture dwarf it, and a sector at three percent is not going to determine the trajectory of global emissions on its own, particularly against industrial processes with far larger footprints.

The reason it matters anyway is entirely about concentration and rate. Three percent globally spread evenly would be invisible. Three percent arriving in a handful of grid regions over five years, after two decades of flat demand growth that caused every actor in the electrical supply chain to size capacity accordingly, is what produces interconnection queues, transformer shortages, and capacity prices that show up on bills. The supply chains that would relieve any of that were themselves sized against the flat decade, which is the compounding problem.

That distinction is the single most common failure in coverage of this subject. A national or global percentage answers a climate question. A local percentage answers a grid question. They are different questions and the same figure cannot serve both, and anybody quoting a share without specifying which one they are addressing is not making an argument.

Where the estimates disagree, and why

The spread between credible projections is wide enough to matter, and the reasons are methodological rather than political.

A critical review of models put plausible AI data center consumption in 2030 at 200 to 400 terawatt-hours, roughly thirty-five to fifty percent of overall projected data center energy, against a range across the literature of 200 to 900 for AI specifically. That is a factor of four and a half between the low and high ends of published work.

Two methodologies compete. Bottom-up approaches inventory installed equipment, apply power draw assumptions, and aggregate, which is more accurate and requires data that operators disclose sparingly. Top-down approaches interpolate from service demand, which is easier and has become unreliable precisely because service demand decoupled from electricity use during the flat decade.

Historical error is instructive here. A 2016 report extrapolated high-end server power draw forward thirteen years at seven percent annual growth, reaching twenty kilowatts by 2020. Actual high-end servers were closer to ten. Long extrapolations of power draw assumptions have a track record of overshooting. The current equivalent assumption is rack density, and the roadmap figures being planned against run from a hundred and forty kilowatts today toward several hundred and eventually a megawatt, which is exactly the kind of extrapolation that went wrong last time.

Disclosure is the underlying problem and it has not improved. Operators report energy use in limited and inconsistent ways, PUE is reported without standardized boundaries, and a substantial share of the global installed base belongs to operators who report nothing, including much of the capacity being built outside the reporting jurisdictions entirely. Every figure in this subject is a model calibrated against partial data, which is not a reason to dismiss the figures and is a reason to state which model produced any given one. The same disclosure problem governs water and cost figures at these facilities, and it has the same cause: no operator is obliged to publish, and the ones that do choose the boundary.

What the buildings are actually for

A brief detour into workload composition, because the AI framing obscures how much of this is not AI.

The IEA projections cover data centers generally, which include streaming, cloud storage, enterprise applications, financial transaction processing, and the ordinary infrastructure of the internet. AI accounted for around fifteen percent of total data center energy in 2024. The growth is concentrated in AI and the base is not, which means a scenario where AI demand disappoints still leaves a large and growing conventional load underneath it, and the infrastructure being built to serve the peak does not become worthless.

Within AI, training and inference have different profiles. Training is periodic, enormously intensive, and concentrated in a small number of facilities. Inference is continuous, distributed, and scales with usage rather than with model development, and it is memory-bandwidth-bound rather than compute-bound, which means its energy profile per unit of output is set by different hardware characteristics than training’s. Forecasts have progressively shifted weight toward inference as the dominant long-run driver, which changes what infrastructure is needed and where, since inference is more latency-sensitive and more geographically distributed than training.

That composition shift matters for everything downstream. A training-dominant future concentrates enormous loads in a few locations near cheap power. An inference-dominant future distributes moderate loads near users, which is a different siting and grid problem with different politics attached. It also changes the materials and equipment demand profile, since many moderate facilities near population centres consume different infrastructure than a few enormous ones near cheap generation.

The rebound question

Efficiency improvements do not reliably reduce total consumption, and the reason has a name and a long history.

Jevons observed in the nineteenth century that improvements in the efficiency of coal use increased rather than decreased total coal consumption, because cheaper coal expanded the range of applications where using coal made sense. The mechanism is general: when the cost per unit of a service falls, demand for the service frequently rises by more than the efficiency gain.

Applied here, an efficiency improvement that halves the energy per query does not halve energy consumption if it more than doubles the number of queries. Cheaper inference makes applications viable that were not viable before, and there is no economic law setting the elasticity below one. The history of resource efficiency in industrial systems is largely a history of this, and the cases where efficiency did reduce total consumption generally involved a saturating end use rather than an expanding one.

The historical case cuts both ways and should be read carefully. During the flat decade, efficiency gains did offset demand growth, which is evidence that offsetting is possible. But compute demand during that period was growing at a rate set by conventional workloads, and the current period has a demand driver that did not exist then.

Nobody knows the elasticity. Anyone who tells you efficiency will solve this, and anyone who tells you efficiency is irrelevant, is asserting a value for a parameter that has not been measured.

The unit problem

Before any of the arguments can be evaluated, the units have to be separated, and they almost never are in general coverage.

Terawatt-hours measure energy over a period and answer questions about fuel, emissions, and annual cost. Megawatts measure power at an instant and answer questions about grid capacity, transformers, and interconnection. A facility described as one gigawatt is a capacity figure; the same facility’s annual consumption depends on utilization and is a different number entirely. Confusing the two produces most of the incoherent comparisons in circulation.

Capacity and consumption diverge further because announced capacity is frequently not built, energized capacity lags contracted capacity by years, and facilities rarely run at nameplate. A gigawatt of announced projects, a gigawatt of contracted power, a gigawatt of installed capacity, and a gigawatt-year of consumption are four claims of decreasing size and increasing verifiability.

Boundary matters as much as unit. Site consumption excludes the water and energy consumed generating the electricity. It excludes the energy embodied in manufacturing the hardware, which for a facility replacing its accelerators every few years is not negligible. It excludes construction. Every reported figure draws a boundary and the boundary is chosen by whoever is reporting, which is the same structural issue that makes water and emissions accounting at these facilities so difficult to compare across operators.

And the denominator is a choice. A share of global electricity, a share of national electricity, a share of a grid region’s peak, and a share of a utility’s load growth are four fractions with the same numerator and wildly different magnitudes, and each supports a different rhetorical conclusion.

The claims that do not hold up

An audit, because the figures in this subject circulate with unstated units more than in almost any other technical domain.

Data center electricity consumption will equal Japan’s is a 2030 projection under a base case, not a current fact, and the comparison is to Japan’s consumption today rather than to Japan’s consumption in 2030.

Data centers use a huge share of global electricity overstates a figure that was 1.5 percent in 2024 and is projected just under three percent in 2030.

Data centers are only three percent so this is a non-issue understates a concentration and rate problem that is entirely real in specific grid regions.

AI is consuming all the electricity conflates AI with data centers generally, when AI was roughly fifteen percent of data center energy in 2024.

Efficiency will solve it as it did before assumes levers that are substantially exhausted at the hardware and facility layers.

Efficiency is irrelevant now ignores that algorithmic and architectural gains have been large and are continuing.

The forecasts are reliable is contradicted by a spread of more than four times across credible published work and by a documented history of overshooting extrapolations.

The forecasts are worthless is the mirror error, since the bottom-up methodology is genuinely rigorous and the direction of travel is not in dispute.

What data center electricity consumption is telling us

The useful discipline this subject teaches is how to read an infrastructure forecast, and the rules generalize well beyond data centers.

Ask what is being projected and over what boundary, because global data center consumption, American consumption, AI-specific consumption, and single-facility consumption are four different quantities that appear in the same sentences.

Ask what the efficiency assumption is, because the difference between the high and low cases in every published projection is substantially an efficiency assumption rather than a demand assumption, and the efficiency assumption is the one nobody can validate.

Ask who produced the projection and what they needed it for, since an agency, an operator, a utility, and a short seller all have positions and all publish numbers. Ask what happened the last time this forecast was made, because the answer is that it was wrong in the direction of overshoot and the mechanism that made it wrong is documented.

And ask whether the question being answered is a climate question or a grid question, because a three percent global share and a hundred-thousand-household building are both true and they support completely different conclusions.

The ten-lecture briefing on how AI data centers work runs the physics, the money, and the politics in the order the constraints bind, and the sequence starts here because the scale figure is what everything else is arguing about. The thermal density determines the load, the load meets a supply chain sized for flat demand, and the resulting shortage becomes a rate case, a ballot measure, and eventually a proposal to put the whole thing in orbit. Every one of those is downstream of the scale figure, which is why getting the figure’s units right is not pedantry.

A forecast made in 2007 said this was unsustainable. It was reasonable, it was well-sourced, and consumption went flat for eight years while computing grew five and a half times over. The current forecast may well be right, and the thing worth remembering is that the last one was made by serious people using good data and it was not.


Comments

Leave a Reply

Your email address will not be published. Required fields are marked *