News / Updates / Blog:

  • The AI Infrastructure Bubble: What Survives Is Not What Fails

    Eighty million miles of fiber optic cable were laid across the United States in the late 1990s. Four years after the crash, eighty-five percent of it was still dark. Bandwidth prices fell ninety percent, in what remains the closest available structural parallel to the current buildout. Global Crossing, WorldCom, and most of their peers went bankrupt, the industry ended up owing something near a trillion dollars, and equity holders were wiped out entirely.

    The fiber is still there. It carries streaming video, cloud computing, and every AI training run currently in progress, and the companies using it bought it from distressed sellers at a fraction of replacement cost.

    That is the shape of an infrastructure bubble, and it is the reason the question people keep asking about the AI infrastructure bubble is the wrong one. Whether there is an AI infrastructure bubble is close to unanswerable and mostly semantic, since the word describes a valuation judgment that can only be confirmed afterward. What can be analyzed is more specific and more useful: which assets outlive their financing, who holds the loss when the financing fails, and what the residue is worth to whoever buys it afterward.

    Those three questions have different answers for the buildings, the power contracts, and the silicon, and separating them is the entire analysis of what this buildout actually is.

    The numbers, stated with their dates

    Start with the magnitudes, because the argument is frequently conducted without them.

    Morgan Stanley estimates the five largest hyperscalers will spend roughly $805 billion on capital expenditure in 2026, up from $261 billion in 2024, with projections near $1.1 trillion for 2027. Aggregate hyperscaler capex is on track to roughly triple across two years, which is a rate of industrial capital formation with few peacetime precedents, though the growth rate is decelerating, from around seventy-three percent in 2025 to something near thirty-six percent in 2026.

    The funding source is what changed. Through the early part of the cycle, hyperscalers funded capex from operating cash flow, which imposed a natural rate limit. That broke in 2025. PIMCO estimates combined hyperscaler capex will consume roughly ninety-four percent of operating cash flow across 2026 and 2027. Free cash flow at the largest technology companies has fallen to its lowest share of sales since early 2024. Alphabet and Meta halted share buybacks; Apple and Microsoft scaled theirs back. That is the most legible behavioural evidence available, because a buyback halt is a public statement that capital has a better use, made by companies that had run buybacks continuously for years.

    So the sector moved from asset-light to capital-intensive, and then to debt-financed. The largest capex providers reported around $385 billion of total debt at the end of 2025 and had issued more debt by mid-March 2026 than in all of 2025. Morgan Stanley and JP Morgan projections suggest the technology sector may need $1.5 trillion of new debt over the next few years.

    Two things follow immediately. Exposure is no longer confined to equity, since credit markets now hold a substantial position. And the buildout has moved past the point where it self-limits, because a company spending from cash flow stops when cash flow stops and a company spending from debt markets stops when the debt markets stop. That is a different stopping condition with a different trigger, since credit availability responds to interest rates, spreads, and sentiment on a schedule unrelated to whether the underlying business is working. The supply chains being expanded against this demand are committing multi-year capital on the assumption that the funding holds.

    What an AI infrastructure bubble unwind looks like

    The mechanism matters more than the probability, and it is worth tracing because it is mechanical rather than speculative.

    The loop currently running is recursive. Rising valuations justify heavier capex. Rising capex is read as a signal of explosive future demand. That signal reinforces the valuations. Every step is individually rational and the whole is self-referential, which means it holds until revenue growth fails to steepen on schedule.

    When it breaks, the sequence is legible. Reported margins compress first, because depreciation on the assets already purchased compounds at thirty to forty percent annually as the fleet ages, and that compression arrives whether or not demand disappoints. This is the mechanical part of the bear case and the part least dependent on any view about AI: the depreciation on hardware already purchased shows up on the income statement regardless of what happens next, and the useful-life assumption determining its size is the contested input. Chief financial officers facing margin questions historically respond by moderating capex growth, and Amazon’s 2022 and 2023 behaviour is the template.

    Capex moderation at the hyperscalers transmits immediately to everybody downstream: accelerator vendors, memory suppliers, equipment manufacturers, construction firms, and the neoclouds whose entire revenue base is hyperscaler contracts. The most leveraged participants in the AI infrastructure bubble fail first, because that is what leverage does. The order is predictable: equity in the most levered operators, then their credit, then the suppliers whose order books were built against those operators, then the equipment manufacturers who expanded capacity on the strength of the same backlog.

    Then the assets change hands. Buildings, substations, interconnection rights, and power contracts get sold by administrators to buyers with cash, at prices reflecting distress rather than replacement cost. Distressed debt funds bought telecom bonds at pennies on the dollar in the early 2000s and new operators acquired network assets at fractions of replacement, which is how the fiber ended up in the hands of the companies that eventually used it. And the acquirers operate infrastructure that cost somebody else three times what they paid for it.

    That is not a prediction. It is the standard shape of a capital cycle in a long-lived asset class, documented across every extractive and infrastructure industry, and it has happened in railways, in electrical generation, in telecoms, and in fiber. The interesting variable is not whether it happens but which parts of the asset stack retain value.

    The three clocks, and which assets survive

    This is the part that determines everything, and it follows directly from a fact established earlier in this subject: the components of a data center have wildly different useful lives.

    The building is a thirty-year asset. Shell, foundations, floor loading, and structure do not become obsolete because a chip generation changed.

    The electrical infrastructure is a twenty-to-forty-year asset. Substations, transformers, switchgear, and interconnection agreements retain value regardless of what computes inside, and given the three-to-five-year lead times on that equipment, an energized site with executed interconnection is worth substantially more than the same site without one. In a distressed sale, this is the crown jewel, and it is the reason distressed data center assets will trade well above what the buildings alone would fetch.

    The power contract is a fifteen-to-twenty-year asset, transferable in principle and subject to counterparty and regulatory conditions.

    The cooling infrastructure is somewhere between, since coolant distribution units and heat rejection plant serve multiple hardware generations if the thermal envelope was specified generously.

    The silicon is four to six years by the generous estimate and two to three by the skeptical one, and it is the asset that does not survive. A four-year-old accelerator in a distressed sale competes against current-generation hardware on cost per token, and it loses.

    So a correction here would look different from the fiber case in a specific way. Dark fiber sat unused for a decade and then became enormously valuable because glass does not age. The buildings and substations in this buildout will behave like fiber. The accelerators will not, and they are the largest single line item in the capital stack. Roughly $180 billion of accelerator spend in a single year, against memory and packaging capacity that was expanded to serve it, is the portion of this buildout with no second owner.

    Where the loss lands

    Tracing the exposure is more useful than characterizing the sentiment, and the positions are unusually stratified.

    Hyperscaler equity holders sit at the top and are best protected. Alphabet’s net income means substantial depreciation increases compress margins without threatening solvency, and analysts forecast depreciation rising by something like $57 billion over four years. These companies do not fail. Their multiples reprice.

    Credit markets are the newly exposed party. Bond investors who bought technology paper on the strength of historically fortress balance sheets now hold claims against companies whose capex exceeds cash generation, and that exposure is distributed through index funds to holders who did not choose it.

    Neoclouds are the most levered and least protected, borrowing at nine to ten percent against collateral with no residual value curve, serving a small number of counterparties, on contracts shorter than the depreciation schedule. This is where the equity gets destroyed first.

    Vendors carry a subtler exposure through circular financing. Interlocking commitments among chip suppliers, model developers, and cloud operators, involving equity stakes, take-or-pay compute agreements, and debt-funded hardware purchases, can make end demand look larger and more independent than it is. The comparison drawn most frequently is to the vendor financing dynamics that ran through Lucent and Nortel in 1999 through 2001, and it is a fair comparison in structure while remaining a different situation in scale and creditworthiness.

    Ratepayers hold a position nobody assigned them. Utilities building generation and transmission on thirty-to-forty-year cost recovery for customers on shorter contracts have created a stranded asset exposure that lands on the remaining rate base if the load leaves.

    Construction and equipment suppliers hold an order-book position, having expanded capacity against backlogs that include speculative reservations, which is the double-ordering problem resolving in the unfavourable direction.

    And local governments hold the tax abatement position, having forgone revenue for a decade or more against a facility whose assessed value declines with the depreciation of the equipment inside it.

    Notice the pattern. The parties with the strongest balance sheets bear the least risk, and the risk migrated toward neoclouds, credit markets, ratepayers, and counties, none of which chose the exposure in the way an equity investor does. That migration is not a conspiracy and it is a predictable consequence of every party negotiating rationally: the strongest counterparty extracts the best terms, and the best terms are the ones that put the risk somewhere else. The same pattern governs cost allocation in the rate cases currently being litigated.

    Why the analogy might not hold

    The historical parallels are instructive and they are not identical, and the differences run in both directions.

    The strongest argument that this is different concerns the demand feedback. WorldCom’s claim that internet traffic doubled every hundred days was simply false, with actual traffic doubling roughly annually, and the entire fiber overbuild rested on a misrepresentation. Bandwidth also could not create its own demand, since a faster pipe does not generate reasons to use it.

    AI capability plausibly does. Improvements in capability create new applications, which create demand for more capability, which is a feedback loop with few complete analogues. Whether that loop is strong enough to absorb the capacity being built is the entire question, and it is genuinely open rather than rhetorically open.

    The second difference is creditworthiness. The 1990s telecom buildout was financed by companies with no earnings against assets with no alternative use. The current buildout is substantially financed by the most profitable companies in the world, and the assets sit on balance sheets that can absorb impairment.

    The differences running the other way deserve equal weight. Asset life is much shorter here, since fiber does not obsolete and accelerators do. Duplication is substantial, with multiple organizations training similar models on overlapping infrastructure and bidding up the same constrained inputs, principally high-bandwidth memory, power capacity, and data center space. And the megacap governance structures insulate the largest participants from the market discipline that eventually reasserted itself in previous cycles, which delays correction rather than preventing it. A dual-class share structure does not repeal capital cycles; it changes who can force a decision and how long the decision takes.

    The utilization problem

    Underneath every position in this argument sits a data gap that makes the debate close to unresolvable with public information.

    Nobody publishes utilization. A hundred gigawatts of capacity is scheduled to come online between 2026 and 2030, and no source provides systematic data on how heavily the existing installed base is being used. Overbuild concerns and demand confidence are both being asserted without the one measurement that would settle them.

    Depreciation estimates compound the opacity. Published figures range from twenty percent annually to thirty or forty percent in the first year alone, with accelerators variously described as retaining half their value after three years and a fifth after five. Those are not necessarily contradictory, since first-year depreciation is always steepest, and they create genuine ambiguity about which rate is appropriate for structuring debt against the collateral.

    Off-balance-sheet commitments add another layer. Moody’s has flagged roughly $662 billion of signed but not commenced data center leases, an obligation larger than the annual capex figure everybody quotes, which is an obligation that exists and does not appear where a casual reader would look for it.

    Which means the honest position on the central question is that it cannot be resolved from outside. Both the bull and bear cases are constructed from the same public filings, and the disagreement is about parameters that only the operators can observe. That is worth stating as a limit on the analysis rather than as a complaint, since the equivalent opacity in materials markets makes those forecasts equally contested and for the same reason.

    The observable gates

    Since the argument cannot be settled by assertion, the useful thing is to identify what would actually move it, and the indicators are specific.

    Useful life disclosures are the highest-signal item and the least watched. Hyperscalers restate depreciation assumptions in estimate-change paragraphs buried in annual filings each January and February. A company shortening its server life estimate is telling you, in the most legally constrained language available, that it believes the assets earn for less time than it previously claimed. Amazon already did this once, moving from six years to five and taking a $700 million operating income hit.

    Capex guidance revisions are the second, and they are visible quarterly rather than annually, which makes them noisier and faster. The current growth rate is already decelerating, and the question is whether that reflects maturation or the beginning of moderation.

    The secondary market for accelerators is the third and the most direct. Nothing about residual value is established until a large fleet of four-year-old hardware is sold or re-leased at an observable price, and that transaction has not occurred at scale. When it does, it will establish a residual value curve for an asset class that currently has none, and every debt structure written against accelerator collateral will be repriced against it in the following quarter.

    Contract renewals are the fourth, and they are where the take-or-pay structures stop protecting anybody. Take-or-pay agreements signed in 2023 and 2024 begin reaching renewal decisions in the late 2020s, and the terms at renewal are the market’s verdict on the depreciation schedule.

    And utilization disclosure, if it ever arrives, would settle more than everything else combined.

    What the reckoning would leave behind

    Assume for a moment that the pessimistic case runs. What exists afterward is worth specifying, because it is the part that determines whether this was waste.

    Gigawatts of interconnected, energized electrical capacity exist that would otherwise have taken a decade of ordinary demand growth to justify. Transmission upgrades built for data centers serve whatever comes next. Generation brought online, including restarted nuclear plants and new firm capacity, continues generating, and the firm capacity procurement that funded it has pulled forward projects that would otherwise not have been financed this decade.

    Buildings exist, with floor loading and cooling infrastructure specified for densities that took decades to become normal in any other industry.

    A memory industry exists at a scale that would not otherwise have been financed, with three suppliers having expanded stacking and packaging capacity against demand that may not persist, along with advanced packaging capacity that took years to build and that every other semiconductor application now benefits from.

    A supply chain for high-voltage electrical equipment exists, expanded against demand that may or may not persist, which is the capacity every other electrification project has been waiting for, and which resolves a shortage that predates AI entirely.

    A domestic supply chain for specialized minor metals and process inputs exists at greater scale than the previous decade could justify.

    And a generation of engineers exists who know how to build and operate high-density liquid-cooled facilities, which is human capital that does not depreciate on any schedule and which the wider electrification buildout has been short of for a decade.

    What does not survive is the equity, a meaningful share of the credit, and the accelerators. The 1990s left dark fiber that became the internet backbone. This would leave energized substations, high-density buildings, and a hardware fleet that ages out.

    The revenue question underneath everything

    Every position in this argument reduces to one unobservable, and it is worth stating plainly rather than leaving implicit.

    The capital being deployed is justified by future revenue from AI services. That revenue currently exists at a scale far smaller than the capital, which is normal for an infrastructure buildout and is also the thing that has to change. The bull case is that adoption compounds and the revenue curve steepens to meet the capex. The bear case is that adoption is constrained by cost, trust, regulation, and organizational capacity, and the curve steepens too slowly.

    The observable signals are mixed in a way that supports neither camp cleanly. Cloud revenue growth at the major providers has been strong. Enterprise adoption surveys report high experimentation and considerably lower production deployment. Consumer usage is large and monetization per user is small relative to the infrastructure cost of serving it.

    What makes this harder than the equivalent question in previous cycles is that the unit economics are moving underneath the measurement. Inference cost per token has fallen substantially through hardware improvement, numerical precision changes, and software optimization, which means revenue per unit of compute and cost per unit of compute are both changing faster than the reporting period. A business that looks unprofitable at one cost structure can look fine at the next one, and the reverse.

    Which is why the efficiency vector cuts in both directions here and gets used selectively by both sides. Cheaper inference expands the addressable market, which is bullish for revenue and bearish for compute demand per unit of revenue. Nobody has established which effect dominates, and the same person will frequently cite whichever supports the position they already hold.

    The claims that do not hold up

    An audit, since this subject generates more confident assertion than any other in the buildout.

    The AI infrastructure bubble will collapse like the dot-com bubble overstates the parallel, since the financing is substantially different and the largest participants are profitable at scale.

    There is no bubble because the companies are profitable understates that profitability at the top does not protect neoclouds, credit holders, ratepayers, or counties.

    Hyperscalers will go bankrupt is not supported by any plausible scenario. Margins compress and multiples reprice; the companies persist.

    The infrastructure will be worthless is contradicted by the physical asset lives, and the infrastructure will all be valuable is contradicted by the silicon.

    Circular financing proves fraud overstates arrangements that are ordinary in capital equipment industries and that do make aggregate demand harder to read.

    Depreciation understatement of $176 billion is a fact treats an estimate by an investor with a position as a measurement.

    Utilization is high, or utilization is low, are both assertions unsupported by public data.

    The buildings will be stranded assets ignores that a permitted, energized, interconnected site is the scarcest thing in the sector and will be bought by somebody.

    Dark fiber proves overbuilding works ignores that the equity holders who funded it were wiped out, that the value accrued to second owners a decade later, and that fiber does not obsolete the way silicon does.

    What the AI infrastructure bubble question is telling us

    The framing that survives is that infrastructure buildouts routinely destroy the capital that funds them while creating assets that outlast everybody involved, and the two facts are not in tension.

    The 1840s railway investors lost their money and the rails are still carrying freight. The 1920s electrical overbuild looked reckless in the crash and the grids were running near capacity by the 1950s. The 1990s fiber investors were wiped out and the fiber carries this sentence. In each case the technology thesis was correct, the timing was wrong, and the mistake was building the future faster than demand arrived rather than building the wrong thing. That distinction is the whole of it. A buildout that constructs useful assets too early destroys capital and leaves infrastructure. A buildout that constructs the wrong assets destroys capital and leaves nothing, which is what happened to the towns built around a single resource that ran out.

    Which reframes what anybody should be watching. Not whether the AI infrastructure bubble bursts, which is a question about a word. But the ratio between how long the assets last and how long the financing does, because that ratio determines whether a correction transfers ownership or destroys value.

    For the buildings and the substations, that ratio is favourable and a correction transfers them cheaply to somebody who will use them. For the accelerators, it is not, and a correction destroys them. The buildout is therefore both things at once: a durable expansion of industrial capacity and a probable destruction of the capital that funded it, in proportions determined by a depreciation schedule nobody can currently verify.

    The ten-lecture briefing on how AI data centers work runs the physics, the money, and the politics in the order the constraints bind, and the money lens exists to insist that every figure state its unit, its stage, and its date. An $805 billion capex number, a $662 billion off-balance-sheet lease figure, and a $176 billion depreciation estimate are three different kinds of claim, and only one of them is a measurement. Which one it is depends on where in the sequence you look, and the discipline of asking is worth more than any particular answer.

    Eighty-five percent of the fiber was dark in 2005 and none of it is dark now. The people who paid for it never saw a return, and everyone on the internet today is using it. Whether that counts as a bubble depends entirely on where you were standing.

  • Data Center Electricity Consumption: The Forecast Was Wrong Last Time

    The number everybody quotes is 945 terawatt-hours by 2030, which the International Energy Agency described as slightly more than Japan’s total electricity consumption today. Data centers used about 415 terawatt-hours in 2024, roughly 1.5 percent of global electricity. The projection more than doubles that in six years.

    It is a good number. It comes from a credible institution, it is built on a documented methodology, and it has been updated since, with the agency’s April 2026 revision moving the central case to 950. It is also a forecast, produced by people who had to make assumptions about adoption rates nobody can observe yet, and there is a specific reason to hold it loosely.

    In 2007, the Environmental Protection Agency reported that American data centers had roughly doubled their electricity use between 2000 and 2006 and warned the trajectory was unsustainable. The extrapolation was reasonable, the underlying data was sound, and the forecast was substantially wrong. Consumption did not continue doubling. It flattened for the better part of a decade while the amount of computing being done grew by a factor of several hundred percent.

    Understanding why that happened, and whether the same mechanism is available now, is the single most useful thing anybody can do before forming a view about data center electricity consumption. It is also the closest thing this subject has to a controlled experiment, since the forecast, the mechanism that falsified it, and the eventual measurement are all documented.

    Where the buildings came from

    The lineage runs through four distinct architectures, and each transition changed the energy profile in a way the previous generation would not have predicted.

    The mainframe room came first. A single large machine in a purpose-built space, with raised floors to route cabling and enormous air conditioning because the machine rejected all its energy as heat into a room that had to stay cool for the equipment rather than for the people. Every design element of a modern facility traces to that room, including the raised floor, which persisted for decades after the cabling reason for it disappeared and was repurposed as a plenum for cold air. Institutional inertia in building design is a real force, and a good deal of data center convention is archaeology rather than engineering.

    Client-server distribution came next and was an efficiency disaster nobody recognized at the time. Organizations bought many small servers instead of one large one, deployed them in closets and converted offices, and ran them at utilization rates in the single digits because each application got its own machine. The industry spent the 1990s accumulating an installed base of idle hardware drawing full power, since a server at five percent utilization draws well over half its peak power doing essentially nothing. That gap between idle and peak draw is the single largest source of waste in the history of the industry, and closing it is what virtualization actually did.

    Colocation and the internet buildout produced the first purpose-built facilities at scale, and the first serious attention to the cost of powering them. Then virtualization arrived and allowed many logical servers to share one physical machine, which raised utilization dramatically and is the single largest efficiency gain in the history of the industry. It also established the commercial logic that made cloud possible, since a provider that can pack many customers onto shared hardware has a cost structure no single enterprise can match.

    Hyperscale is the current architecture and it changed the economics completely. A handful of operators running enormous facilities with custom hardware, custom power distribution, and the capital to optimize at a scale no enterprise data center could justify. That consolidation is the reason the 2007 forecast failed.

    The AI facility is arguably a fifth architecture rather than a variation on the fourth, and the distinction is physical. A hyperscale cloud hall runs racks at five to fifteen kilowatts, air cooled, with workloads that can be scheduled around each other and moved between sites. An AI training hall runs racks above a hundred kilowatts, liquid cooled by necessity, with a workload that wants every accelerator running simultaneously on the same job. The thermal and electrical requirements are different enough that the buildings are not interchangeable, which is why so much of the current construction is greenfield rather than retrofit.

    Why the last forecast was wrong

    The mechanism deserves precision because the current argument turns on whether it still operates.

    Between 2010 and 2018, global data center compute output grew by roughly 550 percent while electricity consumption grew by around six percent. Storage capacity grew by a factor of twenty-five. Energy use stayed nearly flat. That finding, from a recalibration of global data center energy-use estimates published in Science, is the most important fact in the history of this subject, and it was produced by a group of researchers who had spent a decade building bottom-up inventories of what was actually installed rather than extrapolating from service demand.

    Four things produced it. Server efficiency improved continuously, with each hardware generation delivering more computation per watt. Virtualization raised utilization from single digits toward a large fraction of capacity, which meant fewer physical machines doing more work. Workloads migrated from inefficient enterprise closets into hyperscale facilities running at far better efficiency. And facility overhead collapsed, with power usage effectiveness improving from typical values approaching 2.0 in the mid-2000s toward figures near 1.1 at the best hyperscale sites.

    That last one is worth dwelling on. A PUE of 2.0 means every watt reaching a processor requires another watt for cooling, power conversion, and lighting. A PUE of 1.1 means ten percent overhead. Going from 2.0 to 1.1 nearly halves total facility consumption for identical computing, and the industry did that across roughly fifteen years.

    Worth noting what drove it, since the mechanism was economic rather than regulatory. Electricity is a substantial operating cost for a facility running continuously, and at hyperscale it is large enough that a fractional efficiency improvement justifies dedicated engineering teams. The operators that consolidated the workloads were exactly the operators with the scale to fund that engineering, which is why consolidation and efficiency arrived together rather than independently.

    So the forecast failed not because the demand projection was wrong but because the efficiency assumption was, and the people making the projection had no way to know that a set of unrelated improvements would compound into an offsetting force of exactly the right magnitude.

    Why this time might be different

    The argument that the historical pattern will not repeat is stronger than the reflexive skepticism allows, and it rests on the observation that all four of those levers are largely spent.

    Facility overhead cannot repeat its trick. Going from 2.0 to 1.1 saved forty-five percent. Going from 1.1 to 1.0 saves nine percent, and 1.0 is thermodynamically unreachable. The largest single efficiency gain in the industry’s history is exhausted by arithmetic.

    Virtualization cannot repeat either. Utilization improved from single digits to high levels, and there is no second order of magnitude available.

    The consolidation migration is largely complete. Most workloads that were going to move from enterprise closets to hyperscale facilities have moved, which means the population of grossly inefficient installations available to be improved has shrunk.

    And per-transistor efficiency gains slowed with the end of Dennard scaling. Chips still improve, and the improvement now comes substantially from specialization and packaging rather than from the automatic power reduction that accompanied each process node for decades, which is why advanced packaging rather than transistor scaling has become the industry’s binding constraint.

    Against those, the counter-argument is that the efficiency work has simply moved up the stack. Algorithmic improvements, model architecture changes, quantization to lower-precision formats, and inference optimization are all reducing energy per unit of useful output, and several of them are producing gains larger than any hardware generation delivered. The IEA’s own lower case assumes something along these lines, noting that similar to trends in the early 2010s, efficiency improvements could offset most of the impact of increased utilization.

    Which is the honest state of the argument. The hardware and facility levers that flattened the last curve are exhausted. Software and algorithmic levers of unknown magnitude have replaced them. Nobody knows the size of the second set, and both camps are extrapolating from a base of one historical episode.

    There is one more asymmetry worth registering. The hardware and facility gains were largely automatic, in the sense that buying newer equipment delivered them without anybody changing what they were computing. Algorithmic gains are not automatic; they require the workload itself to change, which means they accrue to operators who adopt them and not to the installed base. An efficiency improvement that requires re-engineering a model reaches the fleet far more slowly than one that arrives with the next server refresh, and the fleet is what determines aggregate consumption.

    What the data center electricity consumption numbers say

    Stating the figures with their units, stages, and dates attached is most of the analytical work here.

    Global data center electricity consumption reached approximately 415 terawatt-hours in 2024, around 1.5 percent of world electricity. The IEA base case projects roughly 945 terawatt-hours by 2030, just under three percent of global consumption, with the April 2026 update at 950. Growth runs around fifteen percent annually across that period, more than four times faster than electricity demand growth from all other sectors combined.

    The concentration matters as much as the total. The United States, Europe, and China account for about eighty-five percent of current consumption. The United States and China together account for nearly eighty percent of projected growth, with American consumption rising by around 240 terawatt-hours and Chinese by around 175. In the United States, data centers account for nearly half of all electricity demand growth to 2030, and by the end of the decade the country is projected to consume more electricity for data centers than for aluminium, steel, cement, chemicals, and all other energy-intensive goods combined. That comparison is the most striking single line in the IEA report and it is a statement about American deindustrialization as much as about computing.

    AI is the driver and it is not the whole. AI servers accounted for about twenty-four percent of server electricity demand and fifteen percent of total data center energy in 2024, with AI-focused consumption projected to triple or quadruple by 2030 depending on the case. Data center electricity consumption grew seventeen percent in 2025 while AI-focused facility consumption rose about fifty percent.

    And a single facility is now a meaningful unit. The IEA’s analysis puts a hyperscale AI-focused data center at roughly the electricity consumption of a hundred thousand households. One building. That figure is the one worth carrying into any local argument, because a national percentage is abstract and a comparison to a hundred thousand households is what a county commission is actually being asked to approve.

    The three percent problem

    Here is the figure that reframes the entire public argument, and it cuts against alarm in one direction and complacency in the other.

    Even in the doubling scenario, data centers reach just under three percent of global electricity consumption in 2030. That is not an existential share. Industry, transport, buildings, and agriculture dwarf it, and a sector at three percent is not going to determine the trajectory of global emissions on its own, particularly against industrial processes with far larger footprints.

    The reason it matters anyway is entirely about concentration and rate. Three percent globally spread evenly would be invisible. Three percent arriving in a handful of grid regions over five years, after two decades of flat demand growth that caused every actor in the electrical supply chain to size capacity accordingly, is what produces interconnection queues, transformer shortages, and capacity prices that show up on bills. The supply chains that would relieve any of that were themselves sized against the flat decade, which is the compounding problem.

    That distinction is the single most common failure in coverage of this subject. A national or global percentage answers a climate question. A local percentage answers a grid question. They are different questions and the same figure cannot serve both, and anybody quoting a share without specifying which one they are addressing is not making an argument.

    Where the estimates disagree, and why

    The spread between credible projections is wide enough to matter, and the reasons are methodological rather than political.

    A critical review of models put plausible AI data center consumption in 2030 at 200 to 400 terawatt-hours, roughly thirty-five to fifty percent of overall projected data center energy, against a range across the literature of 200 to 900 for AI specifically. That is a factor of four and a half between the low and high ends of published work.

    Two methodologies compete. Bottom-up approaches inventory installed equipment, apply power draw assumptions, and aggregate, which is more accurate and requires data that operators disclose sparingly. Top-down approaches interpolate from service demand, which is easier and has become unreliable precisely because service demand decoupled from electricity use during the flat decade.

    Historical error is instructive here. A 2016 report extrapolated high-end server power draw forward thirteen years at seven percent annual growth, reaching twenty kilowatts by 2020. Actual high-end servers were closer to ten. Long extrapolations of power draw assumptions have a track record of overshooting. The current equivalent assumption is rack density, and the roadmap figures being planned against run from a hundred and forty kilowatts today toward several hundred and eventually a megawatt, which is exactly the kind of extrapolation that went wrong last time.

    Disclosure is the underlying problem and it has not improved. Operators report energy use in limited and inconsistent ways, PUE is reported without standardized boundaries, and a substantial share of the global installed base belongs to operators who report nothing, including much of the capacity being built outside the reporting jurisdictions entirely. Every figure in this subject is a model calibrated against partial data, which is not a reason to dismiss the figures and is a reason to state which model produced any given one. The same disclosure problem governs water and cost figures at these facilities, and it has the same cause: no operator is obliged to publish, and the ones that do choose the boundary.

    What the buildings are actually for

    A brief detour into workload composition, because the AI framing obscures how much of this is not AI.

    The IEA projections cover data centers generally, which include streaming, cloud storage, enterprise applications, financial transaction processing, and the ordinary infrastructure of the internet. AI accounted for around fifteen percent of total data center energy in 2024. The growth is concentrated in AI and the base is not, which means a scenario where AI demand disappoints still leaves a large and growing conventional load underneath it, and the infrastructure being built to serve the peak does not become worthless.

    Within AI, training and inference have different profiles. Training is periodic, enormously intensive, and concentrated in a small number of facilities. Inference is continuous, distributed, and scales with usage rather than with model development, and it is memory-bandwidth-bound rather than compute-bound, which means its energy profile per unit of output is set by different hardware characteristics than training’s. Forecasts have progressively shifted weight toward inference as the dominant long-run driver, which changes what infrastructure is needed and where, since inference is more latency-sensitive and more geographically distributed than training.

    That composition shift matters for everything downstream. A training-dominant future concentrates enormous loads in a few locations near cheap power. An inference-dominant future distributes moderate loads near users, which is a different siting and grid problem with different politics attached. It also changes the materials and equipment demand profile, since many moderate facilities near population centres consume different infrastructure than a few enormous ones near cheap generation.

    The rebound question

    Efficiency improvements do not reliably reduce total consumption, and the reason has a name and a long history.

    Jevons observed in the nineteenth century that improvements in the efficiency of coal use increased rather than decreased total coal consumption, because cheaper coal expanded the range of applications where using coal made sense. The mechanism is general: when the cost per unit of a service falls, demand for the service frequently rises by more than the efficiency gain.

    Applied here, an efficiency improvement that halves the energy per query does not halve energy consumption if it more than doubles the number of queries. Cheaper inference makes applications viable that were not viable before, and there is no economic law setting the elasticity below one. The history of resource efficiency in industrial systems is largely a history of this, and the cases where efficiency did reduce total consumption generally involved a saturating end use rather than an expanding one.

    The historical case cuts both ways and should be read carefully. During the flat decade, efficiency gains did offset demand growth, which is evidence that offsetting is possible. But compute demand during that period was growing at a rate set by conventional workloads, and the current period has a demand driver that did not exist then.

    Nobody knows the elasticity. Anyone who tells you efficiency will solve this, and anyone who tells you efficiency is irrelevant, is asserting a value for a parameter that has not been measured.

    The unit problem

    Before any of the arguments can be evaluated, the units have to be separated, and they almost never are in general coverage.

    Terawatt-hours measure energy over a period and answer questions about fuel, emissions, and annual cost. Megawatts measure power at an instant and answer questions about grid capacity, transformers, and interconnection. A facility described as one gigawatt is a capacity figure; the same facility’s annual consumption depends on utilization and is a different number entirely. Confusing the two produces most of the incoherent comparisons in circulation.

    Capacity and consumption diverge further because announced capacity is frequently not built, energized capacity lags contracted capacity by years, and facilities rarely run at nameplate. A gigawatt of announced projects, a gigawatt of contracted power, a gigawatt of installed capacity, and a gigawatt-year of consumption are four claims of decreasing size and increasing verifiability.

    Boundary matters as much as unit. Site consumption excludes the water and energy consumed generating the electricity. It excludes the energy embodied in manufacturing the hardware, which for a facility replacing its accelerators every few years is not negligible. It excludes construction. Every reported figure draws a boundary and the boundary is chosen by whoever is reporting, which is the same structural issue that makes water and emissions accounting at these facilities so difficult to compare across operators.

    And the denominator is a choice. A share of global electricity, a share of national electricity, a share of a grid region’s peak, and a share of a utility’s load growth are four fractions with the same numerator and wildly different magnitudes, and each supports a different rhetorical conclusion.

    The claims that do not hold up

    An audit, because the figures in this subject circulate with unstated units more than in almost any other technical domain.

    Data center electricity consumption will equal Japan’s is a 2030 projection under a base case, not a current fact, and the comparison is to Japan’s consumption today rather than to Japan’s consumption in 2030.

    Data centers use a huge share of global electricity overstates a figure that was 1.5 percent in 2024 and is projected just under three percent in 2030.

    Data centers are only three percent so this is a non-issue understates a concentration and rate problem that is entirely real in specific grid regions.

    AI is consuming all the electricity conflates AI with data centers generally, when AI was roughly fifteen percent of data center energy in 2024.

    Efficiency will solve it as it did before assumes levers that are substantially exhausted at the hardware and facility layers.

    Efficiency is irrelevant now ignores that algorithmic and architectural gains have been large and are continuing.

    The forecasts are reliable is contradicted by a spread of more than four times across credible published work and by a documented history of overshooting extrapolations.

    The forecasts are worthless is the mirror error, since the bottom-up methodology is genuinely rigorous and the direction of travel is not in dispute.

    What data center electricity consumption is telling us

    The useful discipline this subject teaches is how to read an infrastructure forecast, and the rules generalize well beyond data centers.

    Ask what is being projected and over what boundary, because global data center consumption, American consumption, AI-specific consumption, and single-facility consumption are four different quantities that appear in the same sentences.

    Ask what the efficiency assumption is, because the difference between the high and low cases in every published projection is substantially an efficiency assumption rather than a demand assumption, and the efficiency assumption is the one nobody can validate.

    Ask who produced the projection and what they needed it for, since an agency, an operator, a utility, and a short seller all have positions and all publish numbers. Ask what happened the last time this forecast was made, because the answer is that it was wrong in the direction of overshoot and the mechanism that made it wrong is documented.

    And ask whether the question being answered is a climate question or a grid question, because a three percent global share and a hundred-thousand-household building are both true and they support completely different conclusions.

    The ten-lecture briefing on how AI data centers work runs the physics, the money, and the politics in the order the constraints bind, and the sequence starts here because the scale figure is what everything else is arguing about. The thermal density determines the load, the load meets a supply chain sized for flat demand, and the resulting shortage becomes a rate case, a ballot measure, and eventually a proposal to put the whole thing in orbit. Every one of those is downstream of the scale figure, which is why getting the figure’s units right is not pedantry.

    A forecast made in 2007 said this was unsustainable. It was reasonable, it was well-sourced, and consumption went flat for eight years while computing grew five and a half times over. The current forecast may well be right, and the thing worth remembering is that the last one was made by serious people using good data and it was not.

  • Space Data Centers: The Escape Hatch That Makes Cooling Harder

    The single most persistent misunderstanding about orbital compute is that space is cold, and therefore cooling is easy.

    Space is cold and it does not cool anything, because cooling requires a medium to carry heat away and vacuum has none. There is no convection. There is no conduction to anything beyond the spacecraft itself. The only mechanism available is thermal radiation, which is governed by the Stefan-Boltzmann law and which requires surface area in quantities that terrestrial engineers never have to think about.

    NVIDIA’s chief executive put it plainly: it is cold in space, and there is no airflow. A space station operator put it more bluntly, calling it counterintuitive that cooling in space is hard precisely because there is no medium to transmit hot to cold.

    Which produces the finding that should govern how anyone reads this subject. Space data centers relieve exactly one of the three constraints that have organized everything else, and they make the most fundamental one substantially worse. Every serious space data centers proposal is therefore a trade rather than an escape, and the question is only whether the trade is favourable.

    What space data centers actually solve

    The case is real and deserves stating at full strength before the objections, because the advantages are genuine and specific.

    Power is the big one. In a dawn-dusk sun-synchronous orbit a satellite sits near the terminator and receives nearly continuous solar illumination, with no night, no weather, and no atmospheric attenuation. Solar panels in that configuration produce substantially more energy per unit area than the same panels on the ground, with estimates running as high as eight times terrestrial output depending on the comparison, and the power is carbon-free without a fuel supply chain, a combustion permit, or a grid interconnection queue. It also requires no fuel cycle, no enrichment capacity, and no reactor licensing, which removes an entire category of multi-year dependency.

    Water use goes to zero, because there is no evaporation and nothing to evaporate.

    Land use goes to zero, and with it the entire apparatus of county boards, zoning hearings, ballot measures, and tax abatement negotiations that has become the binding political constraint on terrestrial siting.

    Transformer lead times, turbine backlogs, and electrician shortages become irrelevant, because none of that equipment exists in orbit. A satellite carries its own generation and distribution, which sidesteps the three-to-five-year transformer queues and multi-year turbine backlogs that currently determine which terrestrial announcements become buildings.

    That is a serious list. Three of the four are the exact constraints that have made terrestrial buildout difficult, and orbit removes them entirely rather than mitigating them.

    Worth naming the political point precisely, because it is the one operators discuss least publicly and value most. A terrestrial project can be stopped by a county commission, a ballot measure, a water permit, a rate case, or an air permit, and increasingly is. A satellite constellation is licensed by a federal regulator and an international spectrum body, and no locality has standing. That is worth weighing against the cost-allocation fights now consuming utility commissions, since a satellite has no rate case and no ratepayers to shift costs onto. The entire apparatus of local objection that has become the binding constraint on siting simply does not apply, which is a structural advantage independent of any physics.

    The heat problem, quantified

    Then the physics arrives. Radiative heat rejection scales with the fourth power of absolute temperature and linearly with area, and the numbers that produces are not intuitive.

    At around one hundred and twenty-seven degrees Celsius, roughly the practical upper limit for electronics, a radiator surface rejects approximately 1,450 watts per square meter. That is the theoretical best case, before accounting for view factors, radiator efficiency, the temperature drop between the chip and the radiator surface, and the fact that a radiator facing the sun or the illuminated Earth absorbs heat rather than rejecting it.

    Work the arithmetic. A one-megawatt cluster requires something in the range of three thousand to ten thousand square meters of radiator area, depending on operating temperature and configuration. A gigawatt-scale facility, which is the unit of ambition in this industry, needs radiator area measured in square kilometers.

    Every square meter of that has to be manufactured, folded into a fairing, launched, deployed reliably in orbit, and then survive micrometeoroid impacts and thermal cycling for years without a maintenance visit. Radiator mass and area do not scale gracefully, which is the central engineering fact about space data centers, and at megawatt scale they plausibly dominate the entire spacecraft.

    The scale distinction matters enormously and gets collapsed. At ten to five hundred watts per node, the thermal problem is largely solved with flight-proven technology carrying decades of heritage in low Earth orbit, and thousands of satellites already reject heat at that scale routinely. At a megawatt, the radiator becomes the spacecraft. At a gigawatt, the structure being described is a megastructure rather than a satellite, and no comparable object has ever been assembled in orbit. Collapsing a fifty-watt edge node and a gigawatt facility into one conversation about space data centers is how the difficulty gets hidden.

    There is a genuine mitigating trend and it deserves credit. Data center accelerators have become dramatically more tolerant of warm coolant. A 2014-generation GPU required coolant below fifteen degrees Celsius to avoid throttling. Current-generation platforms run efficiently at forty-five degrees. Because radiative rejection scales with the fourth power of temperature, a higher operating temperature is worth far more than the linear improvement it looks like, and that trend does real work for orbital feasibility.

    It does not eliminate the problem. It moves the radiator area required for a given load down by a meaningful factor while leaving the scaling relationship exactly where it was.

    Where the demonstrations actually are

    The distinction between what has flown and what has been announced is the whole analytical task, and it is unusually clean in this case because both are documented.

    What has flown: in November 2025 Starcloud placed a roughly sixty-kilogram satellite carrying an unmodified NVIDIA H100 into low Earth orbit at around three hundred and fifty kilometers, and trained a small language model on it. That is the first state-of-the-art data center GPU to operate in space and it is a genuine milestone. Axiom Space deployed orbital data center nodes in January 2026 with optical links in the low gigabits per second, which is real hardware in a real orbit and is bandwidth roughly four orders of magnitude below what a rack backplane moves internally.

    What is scheduled: Starcloud-2 in October 2026, carrying several H100s alongside Blackwell hardware, with plans to deploy AWS Outposts hardware in orbit. Google’s Project Suncatcher prototype, two satellites built with Planet Labs, targeting early 2027 to test Trillium TPUs, optical inter-satellite links, and thermal management, with Google in launch services discussions with SpaceX as of May 2026.

    What has been filed: SpaceX applications for up to one million data center satellites. Starcloud’s February 2026 filing for an 88,000-satellite constellation totaling roughly twenty gigawatts, with a stated vision of a five-gigawatt orbital hypercluster powered by a solar array spanning four square kilometers. A five-month-old company filing in June 2026 for up to 100,000 satellites at around ten gigawatts.

    The gap between those three categories is four to five orders of magnitude. One satellite with one GPU has flown. Filings describe constellations of a million. An FCC filing is a document with an author who wanted something, it costs comparatively little to submit, and it establishes a regulatory position rather than a capability.

    Google’s ground-based radiation testing produced a genuinely useful result, confirming that its TPU v6e can withstand the radiation environment of a five-year low Earth orbit mission, and its laboratory optical link demonstrations reached 1.6 terabits per second on a single transceiver pair. Those are real technical de-riskings of specific subsystems, and they are not the same as an operating cluster.

    Launch economics, and the number everything depends on

    The entire orbital case rests on launch cost falling, and the current numbers are further from the projections than the coverage suggests.

    As of 2026, a reused Falcon 9 delivers payload to low Earth orbit at roughly $2,700 to $3,100 per kilogram, which is a ninety to ninety-five percent reduction from the Space Shuttle era and a genuine achievement. Small payloads on rideshare missions run considerably higher, around six to seven thousand dollars per kilogram. The industrial supply chain underneath launch is also not infinitely elastic, and launch cadence has its own regulatory ceiling that analysts expect to bind through 2028 regardless of vehicle capability.

    Starship is the vehicle every orbital data center business case assumes. As of mid-2026 it had completed a series of test flights, with a record of roughly seven successes across twelve to thirteen attempts, and had deployed functional satellites on a suborbital trajectory rather than into a stable orbit. It has not reached stable orbit on any flight and has not begun selling launches to outside customers.

    Analyst estimates put current Starship flight costs in the range of eighty to one hundred million dollars, which against a hundred-tonne payload implies roughly eight hundred to a thousand dollars per kilogram, an order of magnitude above the target figure.

    The famous sub-hundred-dollar number traces to a 2019 projection of eventual operating cost, made years before the current vehicle existed. SpaceX’s own 2026 prospectus is more restrained, stating an aim to reduce the cost of reaching orbit by ninety-nine percent or more against a historical benchmark of $18,500 per kilogram, which computes to $185 per kilogram with room below.

    Every optimistic figure rests on the same conditions: both stages returning and reflying with minimal refurbishment, and a flight cadence high enough to spread pad, factory, workforce, and development costs across many launches. Neither has been demonstrated.

    Which means the honest framing is that orbital data centers are a bet on a launch vehicle achieving an operating profile it has not yet achieved, priced against a cost per kilogram that is currently a projection.

    Radiation, and the hardware nobody designed for this

    Commercial accelerators are not built for orbit and the environment does two distinct things to them.

    Single-event effects occur when an energetic particle strikes the silicon and flips a bit or induces a transient fault. These are survivable with error correction, redundancy, and checkpointing, all of which cost performance and memory overhead, and which interact badly with memory-bandwidth-bound inference since error-correcting overhead consumes exactly the resource that is already scarce.

    Total ionizing dose is the cumulative one and it is what limits mission life, and it is the constraint that makes commercial silicon awkward, since chips optimized for terrestrial density carry no radiation margin by design. Radiation gradually degrades semiconductor performance, and the degradation is not repairable in place. Google’s testing establishing TPU v6e tolerance across a five-year mission is the most useful public data point available, and it is a five-year number.

    That interacts badly with the refresh cadence. NVIDIA moved to an annual product cycle. A five-year radiation-limited service life against a one-year hardware generation means an orbital facility is running increasingly obsolete silicon for most of its operating life, with no possibility of replacing individual components. Terrestrial operators treat component failure as routine and budget for it; the replacement supply chain for accelerators and cooling hardware is a standing operational cost rather than an emergency. In orbit there is no such supply chain, and proposals for in-situ manufacturing and orbital fabrication are considerably earlier in development than the compute platforms they would service. Terrestrial facilities swap failed accelerators, drives, and power supplies continuously. In orbit, a failed unit is dead capacity until the entire satellite is deorbited and replaced.

    The distributed architecture is the answer to this, and it is why Google’s approach uses clusters of many small satellites replaceable incrementally rather than large monolithic platforms. Distributed failure characteristics are better and the upgrade path exists. It also multiplies the number of objects requiring launch, tracking, and eventual disposal, and it caps the size of any single coherent compute domain, since a model that would occupy a seventy-two-GPU NVLink domain on the ground has to be split across satellites connected by optical links with vastly lower bandwidth than a copper backplane.

    Latency, bandwidth, and what the workload has to look like

    Low Earth orbit adds roughly twenty to forty milliseconds round trip, plus Doppler compensation and handover between satellites as they move relative to a ground station.

    That rules out interactive inference, which is the highest-value and fastest-growing workload in the industry. It is acceptable for batch training and for processing data that originates in orbit, which is the genuinely defensible use case: Earth observation constellations already generate more imagery than they can downlink, and processing in place rather than transmitting raw data is a real argument with real economics behind it, and it is the version of orbital compute that would exist whether or not anybody had ever proposed a gigawatt constellation.

    The bandwidth constraint runs the same direction. Optical inter-satellite links have improved dramatically and space-to-ground optical links remain weather-dependent, since clouds block them. Radio frequency downlink has spectrum limitations, and spectrum is allocated internationally through a coordination process with its own multi-year timeline, which puts it in the same category as every other permitting constraint the terrestrial buildout runs into. Getting a training corpus up and a trained model down is a substantial data movement problem, and the scale-across communication penalty that already constrains multi-site terrestrial training is considerably worse across a link that is intermittent and weather-limited.

    So the workload profile that fits orbit is narrow: batch, latency-tolerant, and ideally operating on data already up there. That is a real market and it is not the market the gigawatt constellation filings describe.

    Debris, slots, and the governance problem

    A million-satellite constellation is a regulatory and orbital-mechanics proposition before it is an engineering one.

    Orbital debris accumulation is the obvious concern, and the mechanism that worries people is cascading collision, where fragments from one collision raise the probability of the next. Deployment on the scale being filed for would change the population of tracked objects in low Earth orbit by orders of magnitude.

    The regulatory apparatus was not designed for this. SpaceX has requested waiver of FCC milestone requirements that normally mandate half a constellation deployed within six years and full deployment within nine, which is an acknowledgment that the schedule implied by the filing is not achievable under existing rules.

    Spectrum coordination, international frequency allocation through the ITU, and end-of-life disposal obligations all scale with constellation size, and none of them has been tested at the numbers being proposed.

    There is also a jurisdictional question that mirrors the terrestrial one exactly. A facility in orbit is subject to the licensing state’s law, which makes orbital siting a jurisdiction and export control decision in the same way terrestrial siting has become one. China’s Three-Body Computing Constellation began launching in May 2025 as part of a planned multi-thousand-satellite program, which means the geopolitical vector arrived in orbit before the commercial one did. Orbital slots and spectrum are allocated on a first-come basis in practice, which turns a filing into a claim on a finite resource and explains why the applications describe constellation sizes nobody expects to build. The same behaviour appears wherever a scarce permitted position has option value.

    What the serious people are actually claiming

    Separating the operators’ claims from the analysts’ assessments clarifies where the disagreement sits.

    Elon Musk has projected cost parity between orbital and terrestrial compute within two to three years. Jeff Bezos has suggested gigawatt-scale orbital data centers within ten to twenty years. Google frames Suncatcher explicitly as early research toward eventual in-space scaling rather than as a near-term product, which is the most careful public framing any operator has offered and is worth noting as a contrast to the confident timelines elsewhere in the sector.

    Deutsche Bank puts cost parity well into the 2030s, which is a bank taking a position and should be read as one. Analysts covering the sector characterize orbital data centers as speculative near-term revenue, citing unproven economics, hardware aging, latency limits, and narrow use cases.

    Notice that the spread is not about physics. Everybody agrees on the Stefan-Boltzmann law, the radiation environment, and the latency. The disagreement is entirely about the launch cost curve and the timeline, which means the argument is a financial one wearing a technical costume. That is worth registering because technical-sounding disputes with financial content resolve on financial evidence, and the evidence here is a flight-rate curve rather than a thermal calculation.

    The capital has arrived regardless. Starcloud raised a $170 million Series A at a $1.1 billion valuation in March 2026 against roughly $200 million total raised. Another entrant reached a reported $2 billion valuation in late March 2026. Venture money is now underwriting the demonstrations, which means investors can get exposure to the outcome without funding the experiment themselves. That is the same structure as any speculative infrastructure buildout financed against a projected cost curve, and it resolves the same way: the demonstrations either hit their gates or the valuations reprice.

    The mass budget, and why it decides everything

    Everything above resolves into one number, and working it explicitly is the most useful thing anybody can do with this subject.

    A satellite carrying compute has to launch its processors, its solar array, its radiators, its structure, its attitude control, its communications hardware, and its propulsion for station-keeping and disposal. Of those, the radiators and the solar array scale with power while the rest scale more slowly, which means at high power the thermal and generation hardware dominate the mass.

    Take a rough case. A megawatt of orbital compute needs radiator area in the thousands of square meters. Deployable radiator panels for spacecraft run in the range of several kilograms per square meter depending on technology and durability requirements. That alone puts radiator mass in the tens of tonnes per megawatt before anything else is counted, and the solar array to generate the megawatt adds its own.

    Now apply launch cost. At the current demonstrated Falcon 9 figure near three thousand dollars per kilogram, tens of tonnes per megawatt implies launch costs in the range of a hundred million dollars per megawatt of compute, against a terrestrial data hall that runs roughly ten million dollars per megawatt all-in including the building. At the aspirational hundred dollars per kilogram, the same mass costs a few million per megawatt, and the comparison inverts.

    That single sensitivity is the entire orbital thesis. Everything else in the engineering is a detail relative to whether launch cost falls by a factor of thirty from a demonstrated figure to a projected one. Anyone modelling this should build the case as a function of dollars per kilogram and observe how little else matters, which is the same structure as the depreciation-schedule sensitivity that governs terrestrial GPU economics: one unobservable input determines whether the business exists.

    The claims that do not hold up

    An audit, because this subject generates more confident assertion per unit of demonstrated capability than anything else in the buildout.

    Space is cold so cooling is free is the foundational error and it inverts the actual difficulty.

    Orbital data centers are coming in two to three years describes filings and prototypes rather than operating capacity, and the operators making the claim have an interest in it being believed.

    Starship will cost a hundred dollars per kilogram is a projection resting on operating conditions not yet demonstrated, and current analyst estimates put the figure roughly an order of magnitude higher.

    A GPU has been operated in space so the technology is proven conflates a sixty-kilogram demonstration with a gigawatt facility, and the scaling problems are exactly the ones the demonstration did not test.

    Orbital compute eliminates environmental impact ignores launch emissions, atmospheric effects of large constellation reentry, orbital debris, and the manufacturing footprint of the satellites themselves, including the critical minerals in solar cells and spacecraft structures that carry their own extraction and processing burden.

    Latency does not matter because it is all batch workloads is true for the defensible use case and inconsistent with the revenue projections attached to the large constellation filings, which require serving general demand.

    The FCC filings show the industry is committed shows that filings are cheap and regulatory position is valuable.

    Space data centers will replace terrestrial ones is the strongest version of the claim and is not supported by anything demonstrated, since the workload profile that suits orbit is narrow and the workload profile driving the buildout is not.

    Nothing in orbit will ever be economic is the mirror error, since orbital edge compute for space-originated data has a genuine near-term case and the demonstrations are real.

    What space data centers are actually telling us

    The reason this belongs at the end of the sequence is that it functions as a test of everything established earlier.

    The terrestrial constraints are power, thermodynamics, industrial supply, and politics. Orbit removes politics entirely, removes the industrial supply constraint for electrical equipment, and improves power dramatically. It makes thermodynamics categorically harder, adds radiation, removes maintenance, adds latency, constrains bandwidth, and substitutes a launch cost curve that has not been demonstrated for a set of equipment lead times that have.

    That is not a solution to the problem. It is a different allocation of the same problem, which is the pattern this entire subject keeps producing. Move cooling water off the site and it reappears at the power plant. Move generation behind the meter and the cost allocation reappears in a rate case. Move the load offshore and the domestic political fight resolves at the cost of the strategic argument. Move the whole facility to orbit and the heat rejection problem, which was the original reason the buildings got expensive, becomes the dominant engineering constraint of the entire enterprise.

    That pattern has a name in every other capital-intensive industry, which is that constraints are conserved rather than eliminated. The materials sector spent a century discovering it: substituting one input for another moves the bottleneck rather than removing it, and the substitution is worth making only when the new bottleneck is genuinely cheaper to relieve than the old one. Whether orbit clears that bar depends entirely on the launch cost curve, which is why the mass budget is the whole argument.

    The genuinely useful version of orbital compute is the narrow one nobody is filing hundred-thousand-satellite applications for: processing data that originates in space, at kilowatt to hundreds-of-kilowatt scale, alongside defence and Earth-observation applications where the strategic value of processing in place exceeds the cost premium, where the thermal problem is solved with flight-proven hardware that has decades of heritage and the latency does not matter because nothing is waiting.

    That business is real, it is being built, and it is roughly six orders of magnitude smaller than the announcements. That is not an argument against it. It is an argument for reading the unit attached to any figure in this space, since a hundred-kilowatt orbital node and a five-gigawatt orbital hypercluster differ by a factor of fifty thousand and appear in the same articles.

    Which is the note the ten-lecture briefing on how AI data centers work ends on, having run the physics, the money, and the politics in the order they bind. Every figure carries a unit, a stage, and a date, and space data centers are where that discipline gets its hardest test, because the gap between what has flown and what has been filed is larger here than anywhere else in the subject. A gigawatt constellation with a 2028 date attached and a prototype flying in 2027 is three of those and only one of them is checkable.

    The heat has to go somewhere. That was true of the first mainframe in an air-conditioned room, it is true of a rack drawing a hundred and forty kilowatts in Arizona, and it is true of a satellite in a dawn-dusk orbit with nowhere to put the heat except empty sky and a fourth-power law that does not negotiate.

  • AI Rack Architecture: The Unit of Computing Changed

    For about sixty years the unit of computing was the chip. Then for about twenty it was the server. As of the current generation it is the rack, and that is not a packaging convention. It is an architectural claim.

    NVIDIA describes the GB200 NVL72 as seventy-two GPUs that behave as one accelerator, and the description is literally accurate rather than promotional. Seventy-two Blackwell GPUs and thirty-six Grace CPUs sit inside a single NVLink domain that the vendor describes as acting as one massive GPU, presenting 13.5 terabytes of unified HBM3e memory addressable as one pool, connected by a switch fabric delivering 130 terabytes per second of GPU-to-GPU bandwidth. A model does not run across seventy-two chips in that rack. It runs on one very large chip that happens to be assembled from seventy-two pieces and weighs about three thousand kilograms.

    https://open.spotify.com/show/0RX8glAwL5YhMbKzrYO9up?si=50216ed3eee14185

    AI rack architecture is the study of why that assembly became necessary, and the answer runs through a constraint most coverage of this industry gets backwards. The binding limit is not how fast the silicon can calculate. It is how fast data can be moved to it, and every design decision in a modern rack is a response to that.

    The claim is worth testing against the alternative reading, which is that this is a marketing construct and a rack is a rack. It is not, and the test is behavioural: a model too large for a single accelerator’s memory can be split across an NVL72 with a modest performance cost and across a networked cluster of equivalent chips with a severe one. Same silicon, same total memory, different result, and the difference is entirely in how the pieces are connected.

    The memory wall, and why FLOPS is the wrong number

    Every vendor datasheet leads with compute throughput, and for the dominant workload in 2026 it is close to irrelevant.

    Language model inference during the decode phase is memory-bound rather than compute-bound. Generating each token requires reading the model weights out of memory, and the arithmetic performed on those weights is trivial relative to the cost of fetching them. Tokens per second therefore tracks memory bandwidth far more closely than it tracks floating-point capability, which means a chip with more bandwidth can outperform a chip with more compute on the workload that actually pays the bills.

    The gap that produces this has been widening for decades. Processor throughput improved far faster than memory bandwidth across the whole history of computing, and the divergence is what engineers call the memory wall. Accelerators arrived at a point where they can calculate faster than anything can feed them.

    The framework engineers use to determine which regime a workload sits in is the roofline model, which plots achievable performance against arithmetic intensity, meaning operations performed per byte fetched. Below a threshold the workload is bandwidth-limited and additional compute is idle. Above it the workload is compute-limited and additional bandwidth is idle. Training sits closer to the compute-limited side because batch sizes amortize weight fetches across many examples. Decode-phase inference sits firmly on the bandwidth-limited side because each token requires a full pass over the weights for a single sequence.

    That distinction has a commercial consequence that the vendor comparisons obscure. As the industry shifts from training toward inference, which forecasts have becoming the primary driver of AI server demand toward the end of the decade, the metric that determines competitive position shifts with it, and a chip selected on compute benchmarks may be the wrong chip for the workload it ends up running.

    High bandwidth memory is the response, and the mechanism is geometric rather than clever. HBM stacks DRAM dies vertically and connects them through the silicon with through-silicon vias, placing the memory immediately adjacent to the processor die on the same package. That produces a 1,024-bit interface per stack against sixty-four bits for a conventional DDR5 channel, sixteen times wider, with a much shorter path and correspondingly lower latency.

    A B200 GPU carries 192 gigabytes of HBM3e at eight terabytes per second. The B300 in the GB300 generation carries 288, raising rack-level memory from 13.5 to 20.7 terabytes. HBM4, entering mass production in 2026, doubles the interface to 2,048 bits and targets around two terabytes per second per stack while maintaining transfer rates above eight gigabits per second.

    That is the actual specification race, and it is being run on memory rather than on transistors.

    HBM, and the three companies that gate the industry

    The consequence of memory being the constraint is that the memory suppliers became the chokepoint, and the market structure is uncomfortable.

    Three companies produce effectively all HBM: SK Hynix, Samsung, and Micron. SK Hynix holds roughly sixty-two percent share with NVIDIA accounting for something like ninety percent of its HBM output. HBM4 allocation is running roughly sixty to seventy percent SK Hynix, twenty-five to thirty percent Samsung, with Micron as the supplementary third source.

    All three reported full capacity allocation through 2026. Micron confirmed its entire year’s HBM production sold out under binding volume and price agreements struck in December 2025, with orders locked more than twelve months ahead of delivery. That is not tightness. That is an industry operating on allocation.

    The structural reason capacity cannot simply expand is that HBM shares fabrication lines with conventional DRAM. Diverting capacity to HBM tightens standard DRAM, which is why memory prices across the board have moved and why NVIDIA reportedly cut gaming GPU production substantially in the first half of 2026 on GDDR7 constraints. The AI buildout is consuming the memory industry’s output and the consumer market is absorbing the shortfall.

    Demand growth compounds it. HBM consumption grew more than a hundred and thirty percent year over year based on 2025 shipments, with 2026 growth still projected above seventy percent as B300, GB300, and Rubin platforms ramp alongside Google TPU and AWS Trainium transitioning to HBM3e.

    Which places a three-firm oligopoly, concentrated in a specific geography, at the base of the entire buildout. The same concentration pattern that governs critical materials applies with the same consequences, and the advanced packaging capacity that assembles HBM onto the processor die was itself the binding constraint on GPU supply until recently.

    The yield problem underneath HBM explains why capacity does not respond quickly. Stacking twelve or sixteen DRAM dies vertically and connecting them with through-silicon vias means a defect anywhere in the stack can fail the whole assembly, so yields on the newest generations start low and improve slowly with process maturity. Reports of base-die issues on early HBM4 production, subsequently resolved before qualification, are the normal shape of that curve. A supplier cannot simply run more wafers to fix a yield problem, and the specialty materials and process gases involved have their own constrained supply.

    Scale-up, scale-out, and why the distinction matters

    The rack exists because two different kinds of communication have wildly different costs, and modern models require the expensive kind.

    Scale-up means connecting accelerators tightly enough that they behave as one device, with shared memory addressing and very high bandwidth. Scale-out means connecting nodes over a network, which is far cheaper per unit and far slower.

    The NVL72 is a scale-up domain. Each GPU has eighteen NVLink connections distributed across nine dedicated switch boards, delivering 1,800 gigabytes per second of bidirectional bandwidth per GPU into a non-blocking topology. Grace CPUs connect to their paired GPUs through NVLink chip-to-chip at 900 gigabytes per second, which allows unified memory addressing so a GPU can reach CPU memory as if it were local rather than traversing PCIe.

    Scale-out uses conventional networking, with InfiniBand favoured for training clusters and Ethernet variants for multi-tenant environments, and both operate at a small fraction of NVLink bandwidth. The Ethernet camp has been closing the gap with AI-specific variants adding congestion control and lossless behaviour, which matters commercially because Ethernet has a vastly larger supplier base and the concentration risk in any single-vendor interconnect is exactly what large buyers try to avoid.

    The reason the distinction is architectural rather than incremental shows up in specific model behaviours. Tensor parallelism splits a single layer’s computation across devices and requires all-to-all communication at every step, which is catastrophic across a network and tolerable across NVLink. Mixture-of-experts models route tokens to different experts, producing exactly the all-to-all pattern that creates communication hotspots on a networked cluster and resolves cleanly on a full-mesh fabric.

    So the seventy-two-GPU domain is not a convenience. It is the boundary inside which a model can be split without paying a network penalty, and a model that exceeds a single accelerator’s memory has to be split somewhere.

    There is a third tier now being discussed as scale-across, meaning communication between data centers rather than between racks, which arises when a training run exceeds what one facility can host. At that distance the bandwidth and latency penalties are severe enough to constrain what parallelism strategies work at all, and it is the reason multi-site training is a networking problem before it is a compute problem, with the fibre routes and latency budgets between facilities becoming a siting criterion in their own right.

    What AI rack architecture actually contains

    Enumerating the contents makes the engineering legible and explains where the money goes.

    Compute trays hold two GB200 Grace Blackwell Superchips each, with each superchip containing two Blackwell GPUs and one Grace CPU. Each Blackwell GPU is itself two dies joined by a ten-terabyte-per-second chip-to-chip link, because a single die at that transistor count exceeds the reticle limit of the lithography process, which is a hard physical boundary set by the optics of the scanner rather than a design choice. Chiplet construction is therefore not an optimization; it is what happens once the ambition exceeds what one exposure can print, and it is why advanced packaging became the industry chokepoint rather than transistor fabrication. The interposer that carries the signals between chiplets and HBM stacks is itself a manufactured silicon component with its own capacity limits, which is one more layer of the supply chain that has to expand before anything else can. The GPU carries 208 billion transistors on a custom TSMC four-nanometer-class process. Each Grace CPU runs seventy-two Arm cores with up to 480 gigabytes of LPDDR5X, which functions less as a host processor than as a high-speed memory extension addressable by the GPUs, solving capacity rather than compute for models whose parameters exceed even the pooled HBM.

    Switch trays hold the NVLink fabric, with industry estimates suggesting eight NVSwitch chips per board across nine boards, each operating at 14.4 terabytes per second aggregate.

    Then the unglamorous majority: power distribution, busbars, the coolant distribution unit, manifolds, cold plates on every processor, networking interfaces for scale-out, and local storage. The magnets, specialty alloys, and minor metals distributed through the power and cooling equipment are a supply story of their own, and the copper alone across a large deployment is a commodity exposure most buyers never price.

    The physical numbers constrain everything downstream. Roughly three thousand kilograms, one hundred twenty to one hundred forty kilowatts, mandatory liquid cooling, and a system price around three million dollars before networking and storage. That mass and that power density are why a data hall built to previous-generation assumptions cannot host one regardless of floor space.

    The printed circuit board problem

    One constraint deserves isolating because it is invisible in every specification sheet and is genuinely at the edge of manufacturability.

    Moving 1,800 gigabytes per second per GPU across a rack requires signal integrity that pushes printed circuit board engineering to its commercial limits: extremely high layer counts, exotic low-loss dielectric materials, impedance control at tolerances that reject most of a production run, and connector systems that maintain signal quality across mechanical mating cycles.

    That matters commercially because it narrows the supplier base. A component that only a few manufacturers can produce at yield becomes a scarcity, and the substrate and packaging materials involved have their own concentrated supply chains. A rack architecture that pushes bandwidth this hard is therefore not merely expensive because the silicon is expensive. It is expensive because several of the boring components are near the limit of what anybody can make. Cabling and connectors are the same story, since a rack carrying this much interconnect contains kilometres of copper in configurations that have to be assembled by hand and tested individually. That labour is skilled, the pool is small, and it is the same specialist workforce shortage that gates the buildings themselves.

    Power delivery, which is becoming its own architecture

    Getting one hundred forty kilowatts into a rack and distributing it to processors drawing over a kilowatt each is a problem conventional data center power design does not solve.

    Traditional racks distribute alternating current to power supplies in each server. At current densities the conversion losses and the copper required for the currents involved become prohibitive, which is why the industry moved to higher-voltage direct current distribution within the rack and is now developing eight-hundred-volt DC architectures aimed at megawatt-class racks toward 2027.

    Higher voltage means lower current for the same power, which means less copper, less resistive loss, and less heat generated by the distribution itself. It also means new safety practices, new component qualification, and a supply base that has to be built. Solid-state transformers, DC protection devices, and the busbar systems to carry those currents are components with small existing markets and long qualification cycles, which puts them in the same lead-time category as everything else electrical.

    Power quality is the constraint nobody outside the industry discusses. An AI training run produces synchronized load steps as thousands of accelerators start and stop computation together, which creates transients the local grid has to absorb and which look nothing like the smooth baseload profile a data center historically presented. Operators now deploy energy storage inside the facility partly to smooth those steps, which means battery systems are becoming standard equipment for reasons having nothing to do with backup power.

    Per-processor power tells the same story from the other end. An H100 runs around seven hundred watts. A B200 runs one thousand to twelve hundred. The B300 generation reaches roughly 1.4 kilowatts per GPU. Each increment tightens the thermal and electrical design simultaneously, and the electrical equipment required to deliver it is on multi-year lead times.

    Numerical precision, which is the quiet efficiency story

    The largest performance gains in the current generation came from arithmetic rather than from transistors, and this is underappreciated.

    Blackwell added hardware support for four-bit floating point. Previous generations accelerated eight-bit. Halving the bits per value halves the memory required to store weights, halves the bandwidth required to move them, and roughly doubles the throughput of the arithmetic units.

    Given that inference is memory-bound, a format that halves memory traffic is worth more than a proportional increase in compute would be. Much of the headline generational improvement is a precision result, and it only materializes for workloads quantized to the format the hardware accelerates. A team not quantizing captures a fraction of the advertised gain, which is a detail that rarely survives into procurement conversations. Quantization also costs accuracy on some workloads, which means the headline efficiency figure carries a quality assumption that has to be validated per model rather than accepted per datasheet.

    The complementary software techniques operate on the same constraint. Paged attention manages the key-value cache more efficiently, prefix caching reuses computed cache entries across identical prompt prefixes, and continuous batching keeps the arithmetic units fed. Every one of them is a memory optimization rather than a compute optimization, which tells you where the bottleneck sits. Prefix caching in particular is close to free performance for workloads with repeated prompt structure, and it is the sort of efficiency gain that complicates any forecast built on compute demand scaling with usage.

    Alternatives, and where the architecture is being challenged

    The rack-scale NVLink approach is dominant and it is not the only design being funded, which matters for anyone assuming the current architecture is permanent.

    Wafer-scale integration takes the opposite approach to the memory wall, putting an enormous number of cores on a single wafer with on-chip SRAM and eliminating external memory access entirely for models that fit. Reported results include throughput several times a Blackwell system on certain model sizes, achieved by removing the bandwidth constraint rather than by adding compute.

    Processing-in-memory places compute elements inside the memory stack, attacking the same problem from the memory side. Major suppliers are developing variants.

    Inference-specific accelerators trade generality for efficiency on the decode workload, which is a defensible bet given that the economics of serving tokens differ entirely from the economics of training, and NVIDIA’s own modular platform now accommodates third-party inference hardware in reference configurations, which is a notable concession from a company whose position rests on an integrated stack.

    Custom silicon from the hyperscalers is the largest structural threat, and it is being pursued for the same reason any large buyer eventually integrates backward into a concentrated supplier. Google TPU and AWS Trainium are transitioning to HBM3e and represent internal demand that does not flow to NVIDIA, and every hyperscaler with a credible internal accelerator program has an incentive to reduce dependence on a single vendor whose gross margins are public. Those programs also compete for the same HBM allocation and the same packaging capacity, which means internal silicon relieves vendor concentration without relieving the physical constraint.

    None of that displaces the current architecture in the near term. All of it is aimed at the same constraint, which is the strongest evidence that the constraint is correctly identified. When wafer-scale integration, processing-in-memory, inference-specific silicon, and custom hyperscaler accelerators are all attacking memory bandwidth from different directions, the diagnosis is not in dispute even where the treatment is.

    The competitive dynamic worth watching is that AMD’s accelerators have carried a memory capacity advantage over comparable NVIDIA parts in several generations, which matters more operationally than the compute comparison implies given where the bottleneck sits. Whether that translates into share depends on software ecosystem maturity rather than on hardware, which is the moat that has held longest and which is the least physical constraint in this entire subject and therefore the most likely to erode.

    The cadence, and what AI rack architecture does to a building

    NVIDIA moved from a roughly two-year product cadence to an annual one, and the buildings have not.

    The GB200 NVL72 was announced in March 2024, ramped through late 2024 and 2025, and is the primary frontier platform in 2026. GB300 entered production in the third quarter of 2025 with fifty percent more memory. Rubin follows in the second half of 2026, and the roadmap points at rack densities of several hundred kilowatts and eventually a megawatt.

    A data hall is a thirty-year asset. Its electrical distribution, floor loading, and cooling plant are specified against a rack density assumption made years before the equipment exists. A facility designed around one hundred forty kilowatts per rack and commissioned in 2028 will be hosting hardware designed for considerably more, and retrofitting a live facility is expensive in a way that greenfield construction is not. That mismatch drives an observable behaviour: operators overbuild electrical and thermal capacity relative to the current generation, accepting stranded capital today to avoid stranded buildings later, which raises the cost per megawatt of everything being constructed and shows up in the construction cost escalation the sector has been reporting.

    That mismatch is the reason the financing question about how long a GPU earns cannot be separated from the building question. The silicon has a four-to-six-year argument attached. The rack architecture that houses it changes annually. The building is committed for decades. Three clocks, no synchronization, and the fastest one setting the specification for the slowest.

    The claims that do not hold up

    An audit, because hardware specifications generate more misleading comparisons than almost any other technical domain.

    More FLOPS means faster AI is wrong for the dominant workload. Decode-phase inference is memory-bandwidth-bound, and a compute comparison between accelerators frequently predicts the opposite of measured throughput.

    The NVL72 is a rack of seventy-two GPUs is technically accurate and misses what makes it different, which is that they present as one device inside a coherent memory domain.

    Bigger is always better ignores fit. A model that fits comfortably on a single accelerator gains nothing from a pooled seventy-two-GPU domain and pays for it, and the correct question is whether memory requirements exceed a single card.

    GPUs are the bottleneck was true in 2023. Advanced packaging capacity expanded, and the constraints moved to HBM allocation, electrical equipment, and grid connections.

    The performance gains are all from better chips understates the contribution of numerical precision and software, and the precision gains require the model to be quantized to capture them.

    HBM shortages will resolve with more fabs understates the DRAM line-sharing problem, the stacking yield curve, and the multi-year cycle to add capacity, which is the same dynamic that governs every specialty materials shortage.

    A rack is a rack is the assumption this whole subject exists to correct, since the difference between a networked cluster and a coherent domain of the same total capacity is a difference in what can run on it at all.

    Vendor benchmark comparisons are directly comparable is rarely true, since published figures specify cluster size, precision format, latency target, and sequence lengths that differ between the compared systems.

    What the rack is actually telling us

    Step back from the specifications and AI rack architecture is a single argument stated in copper and silicon: the models outgrew the chips, so the chips had to be assembled into something larger that still behaves like one chip.

    Every element of AI rack architecture follows from that. Unified memory addressing exists because a model needs one address space. The NVLink fabric exists because splitting a model across a network costs more than the model gains from being split. Liquid cooling exists because the density required to keep the interconnect short generates heat air cannot remove. Higher-voltage distribution exists because the currents involved otherwise waste too much copper. Four-bit arithmetic exists because halving memory traffic is worth more than doubling compute.

    None of those is an independent innovation. They are consequences of a memory bandwidth constraint that has been widening for forty years and became binding at the point where model size passed accelerator memory.

    Which produces the observation worth carrying into everything downstream. The scarce inputs in this industry are not the ones the coverage names. It is not transistors, which TSMC produces at extraordinary volume. It is stacked memory from three suppliers, advanced packaging capacity, printed circuit boards at the edge of manufacturability, and the electrical and thermal infrastructure to run any of it. Every one of those is a physical manufacturing constraint with a multi-year expansion cycle, sitting underneath a demand curve that changes annually.

    That asymmetry is the whole shape of the industry at present. The critical minerals sector spent two decades learning the same lesson, which is that a concentrated supplier of a specialized input captures a disproportionate share of the value created downstream, and that the downstream participants generally do not notice until the supplier exercises the position.

    The ten-lecture briefing on how AI data centers work runs the physics, the money, and the politics in the order the constraints actually bind, and the hardware sets the first one: the rack determines the thermal load, the thermal load determines the building, and the building determines everything anybody argues about afterward.

    Seventy-two chips pretending to be one, drawing the power of a small neighbourhood, cooled by liquid because air cannot carry the heat away, waiting on memory produced by three companies. That is the unit of computing now, and AI rack architecture did not arrive at three tons because anybody wanted a three-ton rack. It became the unit because the models stopped fitting, and everything downstream in this subject is a consequence of that single fact arriving faster than the buildings, the grids, and the supply chains underneath them could respond.

  • Data Center Cost Allocation: Five Bills, Five Different Payers

    Google paid roughly seventy-eight million dollars in property taxes to Caldwell County, North Carolina, and received about seventy-three million of it back under a rebate agreement. The county kept around five million.

    That is not a scandal and it is not a secret. It is a negotiated economic development agreement, approved in public session, of a type that hundreds of jurisdictions have signed. What it illustrates is that the question of who pays for a data center has an answer, that the answer is written down in specific instruments, and that almost nobody reads them before forming a view.

    Data center cost allocation is not one question. It is five, and they resolve through completely different mechanisms operating on completely different timescales. There is an electricity bill, a tax bill, a water bill, an infrastructure bill, and a bill that only arrives if the facility leaves early. Each has its own payer, its own paper, and its own failure mode, and conflating them is why the public argument produces so much heat and so little resolution.

    Data center cost allocation bill one: electricity

    The mechanism by which a data center raises somebody else’s electricity bill is indirect, which is why it took so long to become political.

    A large new load tightens regional supply. Capacity auctions clear higher. Every customer in the zone pays the higher clearing price. No transfer from a residential customer to a data center appears anywhere in the accounting, because none occurs. What occurs is a price effect distributed across a market.

    That indirectness is why data center cost allocation resisted regulation for as long as it did. A rate case is designed to allocate identifiable costs to identifiable customers, and a capacity price effect is neither. It took auctions clearing visibly short, and the resulting increases appearing on bills in the same news cycle as a project announcement, before commissions treated it as a rate design problem rather than a market outcome.

    The regulatory response has been the fastest-moving development in American utility rate design in decades. As of mid-2026, roughly twenty-three states had approved at least one large load tariff, a dedicated rate class for very large customers designed to assign the costs of serving them directly rather than spreading those costs across the general rate base, with several more proposals pending.

    The Database of Emerging Large-Load Tariffs, assembled by the Smart Electric Power Alliance and the North Carolina Clean Energy Technology Center, catalogs the resulting architecture, and a recognizable archetype has emerged from the filings. In the March 2026 public update, thirty-three of seventy-seven filings carried numeric minimum-bill requirements, averaging around eighty percent of contracted capacity, and thirty-seven included collateral requirements. That is a rate design archetype assembling itself in real time across dozens of jurisdictions, which is unusual, since utility rate structures normally change on a generational timescale and the regulatory apparatus was built for incremental load growth.

    The analysis of what that archetype looks like across filings identifies the components clearly enough to list. Minimum billing, fixing monthly payment at a percentage of contracted load regardless of consumption. Extended contract terms aligned with the life of the infrastructure being built. Collateral, in letters of credit or cash. Exit fees. And provisions governing contract modification and capacity reassignment.

    Virginia’s GS-5 rate class is the reference implementation, approved by the State Corporation Commission to take effect in January 2027. It applies to loads at or above twenty-five megawatts, requires fourteen-year contracts, and takes-or-pays a minimum of eighty-five percent of contracted transmission and distribution capacity and sixty percent of generation demand regardless of actual usage, with collateral reported at one and a half million dollars per megawatt. For a five-hundred-megawatt campus that is seven hundred and fifty million dollars posted before the first rack ships, which changes the capital structure of the project and pushes the financing question back onto the operator’s balance sheet.

    What the tariff fight is actually about

    The design parameters look technical and each one is a distributional decision, which is why proceedings that used to be uncontested now draw intervenors.

    Minimum demand percentage is the central battleground. Oregon’s proceeding is representative: staff and a coalition supported a ninety percent minimum, the utility proposed eighty percent arguing that peer utilities sit near there and that setting it too high could deter siting, and the data center coalition agreed with the utility while noting that no other customer class faces such a requirement at all.

    Every position in that dispute is defensible. A higher minimum shifts more risk onto the customer and protects other ratepayers. A lower minimum keeps the jurisdiction competitive and leaves more stranded asset exposure with the utility and therefore with everybody else. A higher minimum also has a second-order effect worth noting: it encourages customers to contract for less capacity than they might need, which improves the utility’s risk position and worsens its planning information, since the forecast it builds against becomes systematically conservative. And the observation that no other rate class faces a take-or-pay obligation is factually correct and is the strongest argument the industry has, though the response is that no other rate class arrives at five hundred megawatts requiring generation that would not otherwise be built.

    The eligibility threshold is its own quiet fight. Utah legislated large loads at one hundred megawatts and above. Virginia’s GS-5 applies at twenty-five. Minnesota directed its commission to set a threshold. Where the line falls determines which facilities are covered and creates an obvious incentive to design just underneath it or to split a campus into separately metered parcels, which is the standard behaviour around any regulatory threshold in an industrial context.

    Contract duration works the same way. Oregon staff recommended fifteen-year minimums for loads above twenty megawatts on the reasoning that stranded asset risk scales with load size. Longer contracts align customer commitment with the depreciation schedule of the assets built to serve them, which is precisely the point, and they also require a company to commit for three times the useful life of the hardware inside its building.

    Pennsylvania’s Public Utility Commission adopted a model framework in April 2026 establishing that interconnection upgrade costs are recovered directly from large load customers rather than from the general rate base, with deposits and collateral sufficient to cover upgrade costs so that projects which do not proceed do not strand those costs on other customers.

    Capacity reassignment provisions are the underrated innovation. Some tariffs permit a customer to reduce or reassign a portion of contracted capacity without penalty, require the utility to attempt reassignment beyond that threshold, and reduce or waive exit fees when the departing customer supplies a successor. That converts a binary default into a transferable position, which is a genuinely better instrument than either a hard lock-in or a free exit. One utility’s schedules permit reassigning or reducing up to twenty percent of contracted capacity without penalty under notice conditions, require the utility to attempt reassignment beyond that, and impose exit fees otherwise.

    Bill two: taxes, and the instruments that hide the number

    Property tax treatment is where the largest sums move and where the accounting is most opaque, and the opacity is structural rather than conspiratorial.

    The straightforward instrument is an abatement, reducing the tax bill by some percentage for a set term. Arkansas law permits up to sixty-five percent abatement for as long as thirty years on projects financed through industrial development revenue bonds, and PILOT agreements reached for Google projects in the state carry the maximum sixty-five percent for thirty years on both real and personal property.

    A payment in lieu of taxes agreement works differently and is easy to misread. The local government takes title to the property, which makes it public and therefore exempt, and leases it back to the company, which pays a negotiated fee instead of taxes. In one Ohio case a facility received a fifteen-year, seventy-five percent property tax abatement alongside a PILOT of five hundred thousand dollars annually.

    Industrial revenue bonds are the third structure and produce the largest numbers. Dona Ana County, New Mexico approved approximately one hundred sixty-five billion dollars in industrial revenue bonds for a data center project, with associated abatements described as undisclosed and running as long as thirty years.

    The aggregate figures where they exist are substantial. Virginia’s data center sales tax exemption reached roughly $1.6 billion annually. Georgia localities were estimated to lose $1.1 billion in 2026 and $1.4 billion in 2027 from state-awarded exemptions. Data centers owned by four large operators in Oregon received $616 million in property tax abatements between 2016 and 2025, with annual program costs rising several hundred percent across that period.

    The measurement problem is the part that should trouble everybody regardless of position. Arkansas does not track the impact of PILOT agreements on property tax collections at the state level. In several states the cost estimates surfaced only because a legislator requested them or a records request produced them. Accounting standards require governments using generally accepted principles to disclose tax abatements in their financial reports, and compliance is uneven.

    The case for the abatements, stated properly

    The critique is easier to write than the defense, so the defense deserves its strongest form.

    The counterfactual argument is the real one. If a facility would not have located in the jurisdiction without the abatement, then the abated revenue was never available to lose, and whatever the county collects, plus construction employment, plus utility revenue, plus any assessed value that does eventually appear, is a gain against a baseline of nothing. Virginia’s original exemption in 2008 carried a fiscal note estimating a forgone $2.8 million against a project the state was otherwise going to lose to North Carolina.

    The service-demand argument is also legitimate. A data center generates minimal traffic, few emergency calls, no students, and no demand on the largest line items in a county budget. Bartow County, Georgia’s policy statement makes exactly this case: substantial revenue against comparatively little service demand, with an explicit intent to increase homestead exemptions as the revenue arrives. Whether that intent survives a change of commissioner is a different question, and the history of resource-revenue windfalls being absorbed rather than distributed is not encouraging on that point.

    The depreciation point is the technical one that both sides underuse. Where equipment is assessed at a percentage of fair market value with statutory depreciation applied, the tax digest impact is a function of depreciated value rather than gross capital cost, which means the headline investment figure and the eventual assessment are very different numbers, and the assessment declines every year as the servers age. The equipment inside is also replaced on a cycle shorter than most abatement terms, which means the digest is a function of reinvestment as much as of the original build. That depreciation curve is also why the useful-life assumption fight in the financing layer has a fiscal consequence nobody discusses: a shorter economic life for the equipment means a faster-declining tax digest for the county that hosts it.

    The honest counter is the timing. During an abatement period running ten to thirty years, the community provides road maintenance, emergency response capacity, and utility infrastructure to a facility that is not yet contributing proportionally, and the service demands arrive before the revenue does, which is the same sequencing problem that afflicts any large industrial facility with a long ramp. Whether that sequencing is a subsidy or an investment depends on what happens at the end of the term, which nobody in the room when it is signed will still be in office to see.

    Bill three: water, and the rate structure underneath it

    Water is the smallest of the five bills in dollar terms and frequently the largest politically, and the reason is that the rate structure makes the cost invisible.

    Municipal water systems are largely fixed-cost businesses. Treatment plants, mains, and pumping stations cost what they cost regardless of throughput. A large new customer improves the utilization of that fixed base, which is genuinely good for the system’s economics and can lower unit costs for everybody on the network, and a volumetric rate that reflects average cost therefore undercharges relative to the capacity the customer requires at peak.

    Where the withdrawal is from groundwater, the cost is not on any bill at all, because a permit to withdraw is not a purchase. An aquifer drawn down faster than it recharges is a stock being depleted, and the depletion shows up as a cost to future users rather than to the current one. Prior appropriation states add a market layer, since a new entrant holds a junior right and can only obtain seniority by purchasing it, generally from an agricultural holder, which converts a public resource question into a private transaction the county has no formal role in. The same structure governs mineral and extraction rights and produces the same local ambivalence.

    Reclaimed water is the mitigation with the clearest economics, since it uses a resource that had no competing use and improves the utilization of treatment infrastructure. It also requires the distribution network to exist, which is a capital project somebody has to fund, and the funding question is a rate case. Purple pipe extensions are expensive per mile and only pencil where a large anchor customer justifies them, which means the data center is frequently the reason the reclaimed system exists at all, and that is a genuine public benefit that the water accounting rarely credits.

    Bill four: infrastructure, and who owns what afterward

    Substations, transmission upgrades, road improvements, and water main extensions are capital projects with thirty-to-forty-year cost recovery, built for a customer whose relationship may be much shorter.

    The Pennsylvania framework’s answer, now increasingly standard, is that interconnection upgrade costs are recovered directly from the large load customer with deposits and collateral sufficient to cover them. That resolves the funding question and leaves the ownership question, because the utility owns the substation afterward regardless of who paid, and the asset either serves a successor or does not. Where it does, the community has acquired grid capacity it did not pay for, which is the best case and is real. Where it does not, the community has acquired a substation serving nothing, which is the stranded infrastructure pattern that has emptied out industrial towns before.

    California’s proceeding surfaced the refund issue that follows. Where a customer funds infrastructure and the utility later serves others from it, some portion is conventionally refunded. The California commission provisionally found a maximum refund of seventy-five percent of total capital expenditure appropriate, reasoning that large load customers present unique stranded cost risks, already receive favourable energy rates, and require investment at a scale that justifies limiting refunds.

    Underneath all of it sits the equipment lead time problem, which changes the economics of any upgrade because a transformer ordered for one customer and delivered four years later may arrive for a project that no longer exists.

    Bill five: the stranded cost, and the one nobody has paid yet

    This is the bill that has not arrived, and it is the reason every tariff proceeding in the country now contains the phrase stranded asset risk.

    The structure is straightforward and uncomfortable. A utility builds generation, transmission, and substation capacity for a data center, financed over decades and recovered through rates. The customer signs for five, ten, or fourteen years. The silicon inside the building has an economic life of four to six. If the load leaves, reduces, or never fully materializes after the assets are built, the assets remain and somebody pays for them, and that somebody is the remaining ratepayers.

    The protections regulators are assembling against exactly this scenario are extensive and recent, and the historical precedent is why they are treating this seriously rather than theoretically. Utility commissions have been through stranded cost episodes before, in the aftermath of nuclear construction programs and again during restructuring, and the resolutions were expensive and politically brutal. The institutional memory is real, and it is why commission staff in these proceedings are frequently more conservative than either the utility or the intervenors.

    The protections now standard in tariffs are all addressed to this single risk. Minimum billing creates a revenue floor independent of consumption. Extended contract terms align commitment with asset life. Collateral covers unpaid obligations. Exit fees penalize early departure. Capacity reassignment creates an alternative to default. Each is an attempt to make a customer whose planning horizon is short behave like one whose horizon matches the infrastructure.

    Whether they are sufficient is unknown, because none has been tested by an actual departure at scale. A fourteen-year contract with collateral at one and a half million dollars per megawatt looks robust against a five-hundred-megawatt facility walking away, and considerably less robust against a general repricing in which many facilities reduce simultaneously and the utility cannot reassign capacity because nobody wants it.

    That is the tail risk in data center cost allocation, and it is correlated rather than idiosyncratic, which is exactly the property that makes collateral requirements less protective than they appear.

    The verification problem

    A structural difficulty runs underneath all five bills and deserves naming, because it limits what any of this analysis can establish.

    Cost shifting is extremely difficult to verify. Utility cost allocation methodologies are complex, contested, and jurisdiction-specific. Whether a particular tariff fully assigns incremental costs depends on assumptions about how those costs are measured, which are precisely the assumptions being litigated. Harvard researchers examining the question have noted that in many markets verification is close to impossible with publicly available information.

    The same applies to fiscal impact. Studies commissioned by industry find substantial net benefits, with one national assessment putting the sector’s contribution above two trillion dollars. Studies commissioned by opponents find substantial net costs. Both are typically methodologically defensible, because the result depends on the counterfactual assumed, and the counterfactual is unobservable.

    Non-disclosure agreements compound it. In several documented cases, commissioners approved abatement agreements while under NDA and without accompanying economic impact assessments or cost-benefit analyses. The same confidentiality practice governs water and power figures, which means a single project can present three separate unverifiable numbers to the same board. A decision made on information the decision-maker could not share and the public could not review is not necessarily a bad decision, and it is one nobody can audit. Legislative efforts to prohibit officials from signing such agreements have been introduced in several states and have generally stalled against the argument that confidentiality is required to compete for projects, which is an argument with real force and no way to test it.

    Which means the honest position on most specific claims in this area is that the number is contested and the methodology is where the argument actually lives. Anyone presenting a clean figure for what a data center costs or contributes is presenting a modeled result and usually not the assumptions. That is the same discipline the critical minerals literature requires of any demand forecast, and for the same reason: the number is downstream of a model, and the model is downstream of an interest.

    What a resident can actually check

    The information asymmetry is real and it is narrower than it looks, because most of these instruments generate a public record even when the negotiation did not.

    The abatement agreement itself is usually a matter of record. If a deal required a vote by a city council, county board, or industrial development authority, that vote appears in minutes, agenda packets, and resolutions, typically published or available by request, and those documents frequently contain the executed terms.

    Where the incentive was created by statute rather than negotiated case by case, the legislation and its fiscal note are public, and the agency administering it usually reports aggregate usage to the legislature.

    State open records laws reach the rest. Every state has one, and the executed incentive agreement between an economic development agency and an operator is generally a public record even where the negotiation was confidential.

    Utility filings are the most useful and least used source, and the technical constraints they document frequently explain project delays that get attributed to politics. A tariff proceeding is a public docket containing the utility’s cost justification, the intervenor testimony disputing it, and the commission’s reasoning, which together constitute a far more rigorous examination of the cost allocation question than any news coverage of it.

    And accounting standards require governments using generally accepted principles to disclose forgone revenue from tax abatements in their annual financial reports. Compliance varies, and where it exists the number is in the notes.

    None of that resolves the counterfactual problem. It does mean that the specific terms of a specific deal are usually knowable, and that most public argument about data center cost allocation proceeds without anybody having read them.

    The claims that do not hold up

    An audit, because the confident assertions run in both directions.

    Data center cost allocation is a solved problem in states with large load tariffs overstates instruments that mostly take effect in 2027 and have never been tested by a departure.

    Data centers pay nothing in taxes is false. Abatements are partial and time-limited in most jurisdictions, sales tax exemptions typically cover equipment rather than everything, and utility taxes and payroll taxes are unaffected.

    Data centers pay their own way is equally unsupported as a general claim, since it depends entirely on the specific instrument, the abatement term, and the tariff in force, all of which vary enormously by jurisdiction.

    Ratepayers are subsidizing data centers is the strongest form of a claim that is genuinely hard to verify, and the direction is plausible while the magnitude is contested.

    Large load tariffs solve the problem overstates instruments that are two years old, mostly untested, and varying widely in how much risk they actually transfer.

    Tax abatements are always a giveaway ignores the counterfactual question, which is the only question that matters and the one nobody can answer.

    The jobs justify the incentives is difficult to sustain at the ratios involved, and most serious economic development arguments have shifted to the tax base and utility revenue rather than employment. Construction employment is genuinely substantial and genuinely temporary, and the specialized trades involved are in national shortage, which means a project frequently imports its workforce rather than hiring locally.

    Data centers do not use public services understates road wear during construction, the heavy-haul permits required to move transformers and turbines, emergency response capability that must be maintained for a high-value facility, and the water and power infrastructure that is public in most jurisdictions.

    Communities can just say no is true and incomplete, since state-level exemptions frequently bind localities that had no vote on them, which is the specific grievance driving several state legislative fights. A county that never voted on a state sales tax exemption still absorbs the service demand of the facility the exemption attracted, and that mismatch between who granted the incentive and who bears the cost is the structural complaint underneath a great deal of otherwise inarticulate local anger.

    What data center cost allocation is actually telling us

    Assemble the five bills and the pattern is that each one is a mechanism for deciding who bears a risk, and the risks are what actually differ.

    The electricity bill allocates the risk that supply tightens. The tax bill allocates the risk that a facility does not deliver the promised base. The water bill allocates the risk that a resource depletes. The infrastructure bill allocates the risk that an asset built for one customer serves nobody. And the stranded cost bill allocates the risk that all of the above happen at once because the demand did not hold.

    That last correlation is the thing most of the instruments handle poorly. Collateral, exit fees, and minimum billing all protect against an individual customer failing. None of them protects against a sector-wide repricing in which many customers reduce simultaneously, capacity cannot be reassigned because demand has fallen everywhere, and the utility holds assets built for a load that no longer exists. Financial instruments designed for idiosyncratic risk perform badly against correlated risk, which is precisely the failure mode the asset-backed lending against depreciating hardware exhibits one layer up the capital stack, which is a lesson that gets relearned expensively about once a decade, most recently in commodity markets where every producer hedged against their own idiosyncratic risk and none against the cycle.

    Which suggests the useful question for anybody evaluating a specific project is not whether data centers pay their fair share, because that phrase does not designate anything checkable. It is: which instrument governs each of the five bills here, what does it assume, and what happens under it if the load reduces by half in year six.

    The ten-lecture briefing on how AI data centers work runs the physics, the money, and the politics in sequence because the allocation question sits downstream of all of them. The thermal density set the load, the load required infrastructure, the infrastructure required financing over decades, and the financing has to be recovered from somebody across a period longer than anybody involved can forecast. Every instrument described here is an attempt to write down, in advance, who that somebody is under conditions nobody can specify.

    A county kept five million dollars out of seventy-eight. Whether that was a good deal depends entirely on what the county would have collected from an empty field, and nobody in that room knew, and nobody knows now.