Nearly all the electricity that goes into a server comes out again as heat. A single rack of ordinary servers can put out as much heat as several domestic ovens running flat out, and racks built for AI hardware can produce several times that. The room has to carry it away continuously or the equipment shuts itself down to protect itself.
That second point is the one that surprises people. Power failures get the attention, but a cooling failure can be quicker and nastier. When the batteries and generators are working, the electrical side carries on regardless. When the chillers stop, a full hall can gain several degrees in a few minutes, and servers begin to throttle, then to power off, long before anyone has finished reading the alarm.
This piece covers how a hosting data hall moves that heat, how efficiency is measured, what changes when the racks get dense, and what a customer can reasonably ask about it.
Hot aisle and cold aisle
The standard layout is hot aisle and cold aisle. Racks face each other across a cold aisle, so servers draw in cool air from the front; the heated air blows out of the back into a hot aisle, which is enclosed or ducted so it does not mix with the cool air. That air returns to cooling units that chill it again, using refrigeration, chilled water, or, when the weather allows, outside air.
The reason for the fuss is mixing. If hot exhaust air curls over the top of a rack and gets drawn into a server above, that server runs warmer than it should, and the cooling plant has to be colder to compensate. Everything done at the room level is about keeping the two streams apart: blanking panels in empty rack spaces, brush strips where cables pass through the floor, doors on the ends of aisles, and a roof over the cold or hot aisle. Walk into a well-run hall and the cold aisle feels like a draughty corridor while the hot aisle feels like standing behind a hairdryer.
What does the cooling
Inside the room, the units that chill the air are called CRACs (computer room air conditioners, which contain their own compressor) or CRAHs (air handlers, which take chilled water from elsewhere). Outside or on the roof sit the chillers or dry coolers that get rid of the heat. Between the two there is a loop of water, sometimes mixed with glycol, and pumps that must themselves be on the protected power supply, or the best chiller in the world is decoration.
Where the climate allows, operators use free cooling: when the outside air or water is cold enough, the compressors are switched off and the outdoor temperature does the work. A site in a cool region may run that way for much of the year; a site in a hot region may hardly ever. Some halls draw in filtered outside air directly, others keep the loops separate and only use the outdoors to cool the water.
Servers are more tolerant than people assume. Industry guidance has long recommended supply air somewhere in the low-to-mid twenties Celsius, and many operators have raised their set points over the years because every degree of unnecessary chilling costs electricity. A hall that feels cold is not necessarily better run.
Measuring efficiency
The industry measures efficiency with PUE, power usage effectiveness: total energy used by the facility divided by the energy used by the IT equipment alone. A value of 2.0 means half the electricity goes to overhead. Good modern sites report figures around 1.1 to 1.4, though the number depends heavily on climate and on how it is measured, so treat comparisons between operators with some caution.
A worked example with illustrative figures: a hall draws 1,000 kW of IT load. Cooling, power conversion losses and lighting add 300 kW. Total is 1,300 kW, so PUE is 1.3. Add a second hall with older chillers that burn 800 kW on the same IT load and the figure is 1.8. Neither number says anything about how the electricity was generated, or about how well the servers use their share of it, which is a separate question.
Where the measurement boundary sits matters too. One operator counts only the cooling plant for the hall; another includes offices, security and a canteen. Annual averages are more representative than a best-day reading taken in winter.
Dense racks and liquid cooling
Liquid cooling, where coolant runs to or around the chips themselves, is increasingly used for high-density racks such as those holding GPUs. For ordinary web hosting, air is still the norm.
The arithmetic explains why. Air is a poor carrier of heat: it takes a lot of it, moving fast, to remove a few kilowatts. Racks of conventional servers have commonly sat in the region of 5 to 10 kW, which air handles well. Racks of accelerator hardware can run many times that, and at some point the fans cannot move enough air through the chassis. Water carries far more heat per litre, so cold plates sitting directly on the processors, or rear-door heat exchangers on the rack, take over. Some sites immerse equipment in a non-conductive fluid. All of these bring plumbing into the hall, which is why the operators who adopt them think hard about leak detection.
When cooling fails
A worked example, with illustrative numbers. A hall has four cooling units and needs three of them (N+1). One fails, the fourth takes the load, and nobody outside the building hears about it. Now suppose a pump on the shared chilled-water loop trips. All four units lose their supply together, which is exactly the sort of shared dependency that redundancy on paper does not cover.
In the first minutes the cold aisle warms, because the stock of chilled air is used up. Within ten or fifteen minutes the inlet temperatures on the upper servers pass their limits, fans ramp to full speed (the hall suddenly gets very loud), and processors lower their clock speeds. Shortly afterwards the hottest machines shut down on their own. Staff meanwhile open doors, bring in portable units, shed non-essential load, and restart the pump or switch to a standby loop. Recovery is slower than the failure, because a hall full of hot metal takes a long while to cool back down, and servers that crashed hard may need file system checks.
What a customer sees is a slow site first, then a dead one, then a pile of servers coming back one by one. This is why chilled-water loops are usually built with two independent paths, and why the pumps and controls sit on protected power.
Commands worth running
You cannot measure a provider's air, but you can read what they publish. Look for a stated PUE and whether it is an annual figure. Look for the cooling redundancy on the facility page, often written N+1 or 2N: N+1 means one spare unit beyond what the load needs. Look at past incidents on the status page for words such as "thermal" or "cooling". On a dedicated server, the operating system itself can tell you the story from your side:
sensors
smartctl -a /dev/sda | grep -i temperature
dmesg | grep -i -E "thermal|throttl"
Readings that creep upward over a week, or throttling messages that line up with one afternoon, are worth a ticket with the times attached.
Loose ends
Does a hotter data centre mean worse hosting?
Not necessarily. Warmer set points are deliberate and within what the hardware is rated for. Rising temperature during an incident is the problem, not the number on an ordinary day.
Why do some providers talk about water use?
Evaporative cooling uses water to remove heat cheaply, which suits some climates and not others. Dry coolers use none but need more electricity. It is a trade-off, and a site should say which it uses.
Will my shared hosting be affected by a cooling fault?
Yes, in the same way as everyone else's equipment in that hall. This is one reason a second location for backups is worth having.