Behind the Rack / Remote hands and hardware swaps

Remote hands and hardware swaps

BEHIND THE RACK

3 min read · 713 words

When a disk fails in a server hundreds of kilometres away, somebody has to walk to the rack. Data centre staff provide a service often called remote hands: swapping disks and memory, reseating cables, rebooting machines, labelling equipment.

Hosting companies keep spare parts on site and have processes for replacing components without interrupting service. In well-designed systems, a failed drive in a redundant array is replaced while the server runs. You may never hear about it.

What remote hands covers

The name is literal: your hands, in someone else's building. A typical request is small and physical. Press the power button on a machine that has hung. Look at the front panel and read out which light is amber. Swap a failed drive for a spare. Plug a cable into the second network port, or move it to another switch. Put a label on a lead so the next person knows where it goes.

Larger jobs exist too, such as racking a new server or moving one between cabinets, but they are usually scheduled and priced differently. For a hosting company running its own hardware, the same work is done by its own staff or by the facility's technicians under a contract, and the quality of that process shows up in how invisible failures are to customers.

A drive swap, step by step

Here is what happens, in order, when a drive in a redundant array dies.

  1. Monitoring notices. The RAID controller or software marks the drive as failed, and an alert is raised before any customer sees an error.
  2. An engineer confirms from a distance, using the machine's logs and the drive's SMART data, that it really is the drive and not a loose cable.
  3. A ticket goes to the on-site team with the server's rack position and the slot number, and a spare of the right model is taken from the shelf.
  4. The technician identifies the right slot, often by lighting a locate LED, so there is no guessing between two similar machines.
  5. The failed drive is pulled and the new one inserted. The array starts rebuilding automatically while the server keeps serving traffic.
  6. The engineer watches the rebuild finish, then closes the ticket and sends the old drive for secure disposal.

The step people underestimate is the fourth. Pulling the wrong drive from a degraded array is a classic way to turn a minor fault into a major one, which is why locate lights and a second pair of eyes matter.

Alertdrive failedDiagnoseremotelyTicketon-site teamSwap driveserver runsRebuild, verifyclose ticketCustomers see nothing if the array has redundancy
The life of a drive failure: most of the steps are remote, and only one needs a person at the rack.

Spares and processes

A provider with a good process keeps a stock of the parts that fail most often: drives, memory modules, power supplies, fans and network cards. Hot-swappable parts (drives, power supplies, fans) can be changed with the machine running. Memory and most cards cannot, which means a planned short outage, usually agreed with the customer in advance and scheduled for a quiet hour.

Standardising on a few server models matters more than it sounds. If every machine uses the same drive caddy and the same memory type, one shelf of spares covers the fleet and a technician rarely needs to look anything up.

Reaching the machine without being there

A lot of problems never need a person. Servers carry a small independent computer, a baseboard management controller, reachable over a separate network port. Through it, an engineer can power a machine off and on, read hardware sensors, see the console as if a screen were attached and even mount a disk image to reinstall the operating system. The standard protocol is IPMI, and vendors wrap their own web interfaces around it.

Server not respondingManagement portstill answers?Reset, read console,reinstall remotelySend technician:swap, reseat, replaceyesno
Remote tools settle most cases; a human at the rack is the last step, not the first.

Checking it yourself

On a server you rent or own, find out what you can see before you need it. sudo smartctl -H /dev/sda reports a drive's health summary, cat /proc/mdstat shows array state, and sudo ipmitool sdr lists sensors if the machine has a management controller. Ask your provider what the remote hands response time is, whether it is included or billed per request, and whether they will confirm the slot with you before pulling a drive.

PreviousFire detection and suppressionNextMonitoring: sensors everywhere

More from Behind the Rack

Behind the Rack

Power

A data centre is, at heart, a very large and very well-organised electrical installation. Power arrives from...

Behind the Rack

Physical security

A facility that stores other people's websites and customer data is serious about who can get in. Typical...

Behind the Rack

Peering, transit and why your host's network matters

A hosting provider reaches the rest of the internet in two ways. It buys transit from larger networks, which...