When a disk fails in a server hundreds of kilometres away, somebody has to walk to the rack. Data centre staff provide a service often called remote hands: swapping disks and memory, reseating cables, rebooting machines, labelling equipment.
Hosting companies keep spare parts on site and have processes for replacing components without interrupting service. In well-designed systems, a failed drive in a redundant array is replaced while the server runs. You may never hear about it.
What remote hands covers
The name is literal: your hands, in someone else's building. A typical request is small and physical. Press the power button on a machine that has hung. Look at the front panel and read out which light is amber. Swap a failed drive for a spare. Plug a cable into the second network port, or move it to another switch. Put a label on a lead so the next person knows where it goes.
Larger jobs exist too, such as racking a new server or moving one between cabinets, but they are usually scheduled and priced differently. For a hosting company running its own hardware, the same work is done by its own staff or by the facility's technicians under a contract, and the quality of that process shows up in how invisible failures are to customers.
A drive swap, step by step
Here is what happens, in order, when a drive in a redundant array dies.
- Monitoring notices. The RAID controller or software marks the drive as failed, and an alert is raised before any customer sees an error.
- An engineer confirms from a distance, using the machine's logs and the drive's SMART data, that it really is the drive and not a loose cable.
- A ticket goes to the on-site team with the server's rack position and the slot number, and a spare of the right model is taken from the shelf.
- The technician identifies the right slot, often by lighting a locate LED, so there is no guessing between two similar machines.
- The failed drive is pulled and the new one inserted. The array starts rebuilding automatically while the server keeps serving traffic.
- The engineer watches the rebuild finish, then closes the ticket and sends the old drive for secure disposal.
The step people underestimate is the fourth. Pulling the wrong drive from a degraded array is a classic way to turn a minor fault into a major one, which is why locate lights and a second pair of eyes matter.
Spares and processes
A provider with a good process keeps a stock of the parts that fail most often: drives, memory modules, power supplies, fans and network cards. Hot-swappable parts (drives, power supplies, fans) can be changed with the machine running. Memory and most cards cannot, which means a planned short outage, usually agreed with the customer in advance and scheduled for a quiet hour.
Standardising on a few server models matters more than it sounds. If every machine uses the same drive caddy and the same memory type, one shelf of spares covers the fleet and a technician rarely needs to look anything up.
Reaching the machine without being there
A lot of problems never need a person. Servers carry a small independent computer, a baseboard management controller, reachable over a separate network port. Through it, an engineer can power a machine off and on, read hardware sensors, see the console as if a screen were attached and even mount a disk image to reinstall the operating system. The standard protocol is IPMI, and vendors wrap their own web interfaces around it.
Checking it yourself
On a server you rent or own, find out what you can see before you need it. sudo smartctl -H /dev/sda reports a drive's health summary, cat /proc/mdstat shows array state, and sudo ipmitool sdr lists sensors if the machine has a management controller. Ask your provider what the remote hands response time is, whether it is included or billed per request, and whether they will confirm the slot with you before pulling a drive.