Backup & recovery

Recovery runbooks: What the on-call admin needs

What a recovery runbook must name, which decisions need a second person, and how to test it on someone new.

6 min read
Dark data center aisle at night with a single lit workbench and racks of servers behind it

Recovery runbooks are usually written by the person who built the system, and that person is not the one who will be woken at three in the morning. The document reads correctly to its author because it assumes everything the author already knows: which console, which account, which of the two similarly named servers, and which failure this procedure is for. Handed to a colleague under pressure, the same document stalls at the first sentence that begins with log in.

The gap is not effort. Most runbooks are long. The gap is that they describe the recovery as a sequence of actions rather than as a set of facts a stranger would need to look up, and looking things up is exactly what is hardest at three in the morning with a director on the phone.

A runbook that works under those conditions reads more like a reference card than a narrative. It names things. It states where to find what it does not name. It is explicit about the two or three points where a person has to decide rather than execute.

Names and addresses rather than instructions to navigate

The most common failure is the sentence log into the backup console. It assumes the reader knows the address, which account to use, and that this console is the right one when the environment has a backup server, a storage management interface and a hypervisor manager that all look similar at speed.

What belongs in the document instead is the address, the account name, and one sentence about what that system does. Not the password, which lives elsewhere, but enough that someone can tell whether they are on the right screen. The same applies to hostnames. An estate with several similarly named servers needs the full name and the address written out, since the difference between two of them is often a single character and the consequence of choosing wrong is restoring over something live.

Storage endpoints deserve the same treatment. A repository is reached by a name, a port and a bucket, and any of the three can be wrong in a way that produces a confusing error rather than an obvious one. Writing them down costs a line and saves an hour.

Where credentials come from and who can open them

Credentials are the part most often left implicit, because the author has them. The runbook does not need the secrets, but it does need to say which vault or safe holds them, who can open it, and what to do when the usual method depends on the systems that are down.

That last case is the one to write for. If the password manager authenticates against a directory that is part of the incident, the recovery stops before it starts. The document should name the offline path: a sealed envelope, an account that does not depend on the directory, a second administrator who can approve access. The same reasoning applies to encryption material, since a recovery copy of the encryption keys is only useful if the runbook says where it is and who may retrieve it.

Decision points and who to wake

A runbook that only contains steps implies every step is safe to perform. Some are not. The document should separate the actions an on-call admin can take alone from the ones that need somebody else awake, and it should say who that somebody is by name and role, with a phone number that does not depend on the corporate systems.

Decision pointWho should decideWhat the runbook must already state
Restore in place or to an isolated locationThe application owner, not the on-call adminNamed owner and deputy, with contact details held outside the affected systems
Which restore point to useOn-call admin with the incident leadWhere verified restore points are recorded and how a clean one is identified
Whether this is a failure or an intrusionThe incident lead, on a defined thresholdThe signals that change the classification and who is called when they appear
Whether to restore over live dataThe service owner, in writingWhat is overwritten and whether the current state can be captured first
Whether to lift a retention holdTwo named people acting togetherWho holds that right, and that it is never exercised by one person alone

The second row is where most time is lost. Choosing a restore point under pressure is a judgment call about when the damage started, and identifying a clean recovery point is much faster if the runbook already points at the log, the alert history and the last verified restore.

The order in which systems come back

Restore order is the part a stranger cannot reconstruct. Directory services, name resolution, certificate services, the virtualization management layer and the backup infrastructure itself all have dependencies that are obvious to the person who built them and invisible in a list of machines sorted by name.

The runbook should carry a short dependency order with the reason attached, since reasons survive changes better than lists do. Authentication before applications that use it. Name resolution before anything addressed by name. The backup server before the machines it restores, if the backup server itself was lost. Where a database has its own log based recovery, that step belongs after the machine is running and not merged into it.

Capacity and time belong here too. Restoring fifty machines is not fifty times one restore, because the constraint moves to the shared path between repository and hypervisor. Planning for a large scale recovery is what turns a restore order into a realistic sequence with expected durations.

Where ARTESCA fits

ARTESCA is object storage software used as a backup target, deployed on infrastructure the customer runs. For runbook purposes that means the details a restore needs are local facts: the endpoint name, the bucket, the credentials for the account the backup software uses, and the network path between the repository and the servers being restored.

Where immutability is configured through S3 Object Lock, the runbook should state what the retention settings are and what they prevent. A locked restore point cannot be deleted or altered during an incident, which is the point, and it also means that any expectation of cleaning up space mid recovery is misplaced. Whether anyone can lift governance retention, and who, is a fact worth recording before it is needed.

Because deployment is on customer infrastructure, physical access and operational access are also the customer's. If recovery depends on someone reaching a rack or a management network, that person and that path belong in the document alongside the software steps.

Testing a runbook by giving it away

The only reliable test is to hand the document to someone who has never used the system and watch. The author does not explain, does not correct, and does not answer questions except to note them. Every question asked is a defect in the document, and the list of questions is the revision.

That exercise is cheap compared with a full drill and finds a different class of problem. A recovery drill run as though production were down tests whether the plan works. Handing the runbook to a stranger tests whether the plan can be followed by the person who will actually be holding it.

A rhythm that holds up is a read through by a different person each quarter, a rewrite of whatever they stumbled on, and a review after any change to authentication, addressing or the backup repository. During a real event, the on-call admin should record the start time, the restore points used, each decision and who approved it, and the elapsed time per system. That record is the next revision of the runbook, and it is worth more than the one written from memory a week later.

Try ARTESCA free

Immutable object storage that scales from 20TB to petabytes. Deploy a working cluster in under an hour.

Start a free test drive