Background
Understanding data centres
Enough of the physical and economic picture to read the rest of this site properly: what these buildings contain, why their size is quoted in megawatts rather than square feet, why cooling them spends water, and what each number here is and is not claiming.
What is actually inside one
A data centre is a building whose purpose is to keep computers running without interruption. Strip away the scale and it is three systems wrapped around each other:
- The white space
- Rows of racks — steel frames, each holding servers, storage and networking gear. This is the part doing the work, and its electricity use is called the IT load.
- Power
- A connection to the transmission grid, transformers to step it down, uninterruptible power supplies with batteries to bridge the seconds before backup generators start, and usually diesel generators to carry the site through an outage. Redundancy is the product being sold.
- Cooling
- Almost every watt a server draws leaves as heat, and the heat has to go somewhere or the equipment fails. Chillers, air handlers, and often cooling towers move it outdoors. This is the second-largest consumer of electricity in the building and the reason water enters the story at all.
What varies enormously is scale. Across the 1,506 mapped buildings in this dataset with an outline drawn, the median covers 10,537 m² while the largest covers 124,329 m² — a spread of roughly 12 to one. That gap is the reason Helios weights its power allocation by building floor area instead of dividing the national total evenly across facilities; a flat per-facility figure would be wrong by about an order of magnitude at both ends of that range.
Those are buildings. The same OpenStreetMap tags are also used for the land a campus sits on, and the largest such parcel covers 3.2 km² — 306 times the median building. Counting that area as though it were floor space is what the allocation used to do, and it sent 82% of the national total to geometry that is not a building. Helios now weights buildings only, and says so wherever a region has parcels it therefore cannot estimate.
Why the unit is the megawatt
Ask how big a data centre is and the answer comes back in megawatts, not floor area. That is not jargon — it reflects what is actually scarce. Floor space is straightforward to build. A grid connection able to deliver hundreds of megawatts continuously is not, and in several regions the waiting list to get one now runs for years. Power is the binding constraint, so power is the unit of account.
The distinction that trips people up is power versus energy. A megawatt (MW) is a rate — how fast electricity is being used at this instant. A megawatt-hour (MWh) is a quantity — a megawatt sustained for an hour. A terawatt-hour (TWh) is a million of those.
For a human sense of scale: an average US home uses roughly 10,500 kWh a year, which spread over 8,760 hours is a continuous 1.20 kW. One megawatt is therefore about 834 homes worth of average electricity demand. A single large campus drawing 300 MW is in the range of a small city. Treat this as an illustration, not a measurement: it compares averages, and a neighbourhood's peak demand behaves quite differently from a data centre's flat one.
PUE, and why the building uses more than the computers
The industry's standard efficiency measure is PUE, power usage effectiveness: total facility electricity divided by the IT load alone.
A PUE of 1.0 is the unreachable floor — every watt reaching the computers and nothing spent on cooling, conversion losses or lighting. A PUE of 1.5 means that for every kilowatt of computing, another half-kilowatt goes to running the building. The gap between an efficient hyperscale facility and an older enterprise room is large, and it is mostly cooling.
Two things to carry with you. PUE is almost always self-reported and rarely audited. And it is a ratio, so it improves when the computers work harder — a site can report a better PUE while consuming considerably more electricity in absolute terms. Nothing on this site is derived from a PUE figure; it is described here because you will meet it everywhere else.
Why cooling spends water
The cheapest way to reject a large amount of heat is to evaporate water. A cooling tower does exactly that: warm water is trickled through moving air, some of it evaporates, and the evaporation carries the heat away. The water that evaporates is consumed — it does not return to the source.
The alternative, a closed-loop or air-cooled design, consumes almost no water, but rejecting the same heat without evaporation takes noticeably more electricity. So the industry faces a direct trade: water or power. Which is chosen depends on climate, local water price and politics — which is why the same operator builds differently in Oregon and in Arizona. The efficiency measure here is WUE, water usage effectiveness, in litres per kilowatt-hour of IT load.
LBNL puts direct water consumption by US data centres at 17.4 billion gallons in 2023. “Direct” matters: it counts water evaporated on site and excludes the water consumed generating the electricity in the first place, which is substantially larger and depends entirely on how the local grid is powered. The water figures on this site are the direct kind, because that is what the source reports.
Why they cluster so tightly
One county in this dataset holds more mapped data centres than most states. That is not an artefact of the data; the industry genuinely concentrates, for reasons that compound:
- Fibre follows fibre. Long-haul routes and internet exchanges were laid where earlier ones already ran. Northern Virginia's density traces back to the region hosting one of the internet's earliest major exchange points; networks converged there, and networks are what a data centre is selling access to.
- Power availability and price. Substation capacity and cheap generation attract siting more than almost anything else, which is why the map lights up along particular transmission corridors.
- Tax treatment. Many states exempt data-centre equipment from sales tax. On a build where the servers cost more than the building, that changes the arithmetic decisively.
- Latency to users. Distance costs milliseconds, so anything serving interactive traffic wants to be near population.
- Everyone else being there. Once a cluster exists, it has the trained electricians, the permitting precedent, the peering partners and the supply chain. Each new build makes the next one easier.
The consequence for reading this site: national totals hide almost everything interesting. The regions view is where the concentration becomes visible.
How to read every number here
Helios keeps three kinds of claim strictly apart, and labels which is which wherever a figure appears.
| Claim | Class | What it rests on |
|---|---|---|
| This facility is at this coordinate | Reported | An OpenStreetMap contributor mapped it and tagged it. |
| It first appeared on the map in this month | observed | OpenStreetMap's edit history records when the element began matching the data-centre filter. |
| It was built in that month | never claimed | OpenStreetMap carries no construction dates. Not one facility in this dataset has one. |
| US data centres used 192 TWh | Reported | Published by Lawrence Berkeley National Laboratory. |
| This county accounts for N megawatts | Inferred | Its mapped footprint's share of that national total. Not a meter reading, and an upper bound. |
| 2030 will reach 649 TWh | Predicted | A published scenario, quoted with its range, not a forecast Helios makes. |
Why the megawatt figures are upper bounds: the national total is spread across only the facilities that have been mapped. Every real data centre nobody has mapped has its consumption silently handed to the ones that have been. The allocation is exact by construction — the state shares re-sum to 21,918 MW — but exact arithmetic on an incomplete denominator is still an over-estimate per facility.
What this site cannot tell you
- Whether a region truly has no data centres. A county showing zero has not been shown to have none — it may simply have no one mapping it. OpenStreetMap coverage is contributor-driven and uneven.
- How complete the 1,853 is. There is no authoritative public count of US data centres to check it against. The coverage rate is not withheld here; nobody knows it.
- When anything was built. The growth curve is a record of mapping activity. Its near-zero start before 2017 is a tagging convention being adopted, not an empty country, which is why that stretch is drawn hatched rather than cropped away.
- That a removal means a demolition. An element leaves this dataset when it stops matching the filter — which happens when a contributor retags it just as readily as when a building comes down.
- What any individual facility actually draws. Per-site metered power and water are not public. Everything here is an allocated share.
The methodology page states how each figure is produced and lists the known defects in full.
Glossary
- Colocation
- A facility renting space, power and cooling to many customers who bring their own servers.
- Hyperscale
- A very large facility operated by a single company for its own cloud or platform — the buildings that dominate the footprint figures here.
- IT load
- Electricity drawn by the computing equipment itself, excluding the building.
- kW / MW / GW
- Units of power, each a thousand times the last. Rates, not amounts.
- MWh / TWh
- Units of energy: power sustained over time. 1 TWh is one million MWh.
- PUE
- Total facility energy ÷ IT energy. Lower is better; 1.0 is unreachable.
- WUE
- Litres of water consumed per kWh of IT energy.
- Footprint
- Ground area of the mapped building outline, in square metres, computed on the ellipsoid rather than from degrees.
- FIPS
- The federal numeric code identifying a US county — how facilities are matched to regions here.
- Interconnection queue
- The waiting list to connect a large new load or generator to the grid. Often the real schedule constraint on a project.