Behind the Rack / Power

Power

BEHIND THE RACK

5 min read · 1,208 words

A data centre is, at heart, a very large and very well-organised electrical installation. The servers get the attention, but most of the engineering budget goes on getting clean power to them and keeping it there when something upstream misbehaves. Power arrives from the grid, often over two independent feeds, and passes through transformers and switchgear into the building. Before it reaches a server, it goes through a UPS, an uninterruptible power supply, which is essentially a large bank of batteries that smooths out dips and bridges the gap if the grid fails.

Batteries are not meant to run the building for long. They hold on for a few minutes, enough for diesel or gas generators to start up, stabilise and take over. In a well-run facility these generators are tested regularly under load, because a generator that has never been started in anger is a hope, not a plan.

Redundancy is described in shorthand, and the shorthand is worth learning because hosting companies use it in sales material. N is the capacity you need. N+1 means one spare unit of each critical component, so any single failure is covered. 2N means two complete, independent systems, each able to carry the whole load. Uptime Institute's tier ratings, from I to IV, grade facilities on this kind of design. A higher tier is more resilient, and more expensive, and it is paid for through your hosting plan, noticed or not.

From the street to the rack

Follow the electricity in. A large facility is usually fed at medium voltage, and a transformer steps it down to the voltage the building distributes internally. Switchgear sits after that: heavy breaker panels that route power, isolate faults and decide which source (grid or generator) is feeding the building at any moment. From there power flows through the UPS, then to power distribution units (PDUs) on the floor, and finally to strips mounted in each rack that the servers plug into.

Every stage is a place where something can fail, which is why the interesting question is never "is there a UPS" but "what happens when this particular box needs maintenance". A design that cannot survive routine maintenance without a shutdown is not really redundant.

Gridtwo feedsTransformersteps downSwitchgearpicks sourceUPSbatteriesPDUfloor panelRackserversGeneratordiesel or gasStandby source, starts when the grid fails
The power path in a typical facility: the generator joins at the switchgear, and the UPS sits between the switchgear and everything the customers can see.

What the UPS actually does

A UPS does two jobs. In normal operation it conditions the supply, absorbing spikes and sags so that servers see something close to a clean sine wave. When the grid drops out, it carries the load from its batteries with no gap, which matters because a server power supply can only ride through a few tens of milliseconds on its own.

The battery runtime is deliberately short: the UPS is a bridge covering the seconds to minutes before the generators accept load. Batteries wear out, typically over several years, and a UPS with tired batteries will look healthy on a dashboard right up to the moment it is asked to do its job. Good operators test batteries and replace modules on a schedule.

Generators and the tests that matter

When the grid goes, the switchgear signals the generators to start. Each one has to crank, reach speed and voltage, and then the transfer gear closes it onto the building. That takes somewhere around ten to thirty seconds in a healthy setup, which is the gap the UPS covers.

The phrase to look for is "tested under load". Starting a generator with nothing attached proves the engine runs. It does not prove that the transfer switch works, that the cooling plant restarts in the right order, or that the fuel filter has not clogged. Operators who take this seriously run periodic tests where the building actually transfers off the grid and runs on generators for a while.

UPS batteriesGenerator carries the loadBack on the grid0 sgrid lost~30 sgenerator on loadhours latergrid stable, transfer backIllustrative timings, not to scale
A grid failure as the building sees it: the batteries bridge only the short gap, then the generators do the long work.

N, N+1 and 2N in practice

Take a facility that needs three UPS modules to carry its load. That load is N. With N+1 it installs four, so any one can fail or be taken out for service and the other three still carry everything. With 2N it installs two entire systems, each with its own modules, switchgear and distribution, and each capable of carrying the whole load by itself.

DesignSurvivesWeak pointRelative cost
NNothing; any fault is an outageEvery componentLowest
N+1One component failing or being servicedShared paths and a second fault during repairModerate
2NAn entire path failing or being servicedOperator error, shared fuel or coolingHigh

The catch with N+1 is that the modules often share an output panel, which stays a single point of failure. With 2N the paths are physically separate, which is why servers in such a hall have two power supplies, one on the "A" feed and one on the "B" feed. If an entire side goes down, each server carries on from its other supply.

N+1: three needed, four fittedUPSUPSUPSspareOne shared outputSurvives a module failureShared panel remains a risk2N: two full, separate systemsSystem ASystem BServer, two PSUsSurvives a whole side failingCosts roughly double
N+1 protects against a component failing; 2N protects against a whole power path failing or being taken down for maintenance.

What the tier ratings tell you

Uptime Institute's tiers are a classification of design, not a measure of how well a particular building is run. In rough terms, Tier I is a basic site with no redundancy to speak of. Tier II adds redundant components. Tier III is concurrently maintainable, meaning any single piece of the power or cooling path can be taken out for work without shutting down the IT load. Tier IV is fault tolerant, so a single failure of any kind does not affect the load.

Two cautions. First, a provider saying "Tier III design" or "Tier III equivalent" is making a claim about its own design, which is different from a facility that has been formally certified. Second, the tier says nothing about the other things that cause outages: staff following a wrong procedure, a firmware bug in a transfer switch, or a network problem that has nothing to do with electricity. Most real-world incidents involve a chain of small things, and a good tier rating removes some of the links.

What this means for your site

On shared hosting you rarely choose the power design; you take whatever the provider's data centre has. For a small business site the better investment is usually not a higher tier but what happens after power returns. A machine that comes back with a corrupt database, or a service that does not start by itself, turns a ten-second blip into an hour of work. Check that services are enabled at boot, and ask the provider what counts as downtime in their terms.

Things people ask

Does a UPS mean my server never goes down?

No. It covers the gap until generators take over. If a UPS module fails, a breaker trips or the batteries are worn, the load can still drop.

Why do some servers have two power cables?

So that each cable can go to a different feed. If one side of the power system fails or is serviced, the machine keeps running from the other.

NextCooling

More from Behind the Rack

Behind the Rack

Cooling

Nearly all the electricity that goes into a server comes out again as heat. A single rack of ordinary servers...

Behind the Rack

Maintenance windows and change control

Many outages begin with a planned change. Serious operators therefore treat changes carefully: peer review...

Behind the Rack

Why redundancy still fails

Spare parts do not prevent every outage, because many outages are not caused by parts. Three examples that...