Cyber resilience

Backup encryption keys: Who holds the recovery copy?

Where backup encryption keys live across software, storage and hardware, and what an escrow copy must contain.

7 min read
Dark server room with a sealed vault door beside racks, thin cyan light tracing a path from the door to storage

Encrypted backups behave like any other encrypted data. Without the key they are a large set of unreadable blocks, and no support contract changes that. The risk is rarely that an organization failed to encrypt. It is that the key, the passphrase or the keystore that unlocks the backup sits inside the same environment the backup exists to recover, and nobody discovers this until that environment is gone.

Encryption is also the one control that fails silently. A misconfigured retention policy shows up in a capacity report. A broken replication job raises an alert. A key that exists in exactly one place produces no signal at all, since everything works right up to the moment that place stops existing. The useful test is not whether backups are encrypted, it is whether a named person can produce the key material on a day when the estate is unavailable.

Where keys actually live

Most environments hold key material in three or four separate layers without a single inventory covering them. The backup application encrypts job data using a passphrase or a certificate, and stores what it needs in its own configuration database or keystore on the backup server. That store is protected by the application, which means it is usable only by that installation, and in several products a restored copy of the database on a fresh server will not open the old keystore without an explicit export taken beforehand.

Underneath that, the storage layer may encrypt at rest on its own terms, with keys held by the platform or by an external key manager reached over KMIP. Underneath that again, self encrypting drives and tape drives hold authentication keys paired to a controller or a key manager, and a chassis or controller replacement without those keys turns healthy media into noise. None of these layers knows about the others, and a restore can be blocked by any one of them.

Finally there is material that is not strictly a key but is required to use one: a product version that can read the format, a license file, and the record of which passphrase applies to which job and date range. A recovered passphrase still stalls a restore if three were in rotation over two years and nothing recorded which restore points belonged to which.

The dependency loop that makes keys unrecoverable

The common failure is circular rather than careless. The key manager runs as a virtual machine on the cluster it protects. The password manager holding the backup passphrase is itself a virtual machine protected by the backup system. The documented recovery procedure lives on an intranet page served by a server inside the affected estate. Access to any of these is gated by directory authentication, and the directory is the thing that is down.

Each decision is reasonable in isolation and invisible in normal operation, since the dependency only ever resolves in one direction during a working day. Drawing it explicitly, from key back through every system required to reach it, is a short exercise that tends to produce an uncomfortable diagram.

This is the same reasoning that makes an offline path necessary for administrative access generally, and it is worth resolving alongside separating backup administration from everyday accounts. An escrow copy that can only be opened by an identity in the compromised directory is not an escrow copy, it is a second instance of the same dependency.

What an escrow copy has to contain

An escrow copy is not a copy of a password. It is the complete set of material required to turn ciphertext back into usable data without help from the environment being recovered. Written out per layer, the list is usually longer than expected.

Key materialWhere it normally livesWhat the escrow copy must carry
Backup job encryption passphraseConfiguration database on the backup serverThe passphrase, plus the jobs and date ranges it covers
Application keystore or certificateProtected store tied to that specific installationAn export file, the password that opens it, and the product version
Storage side encryption keysThe storage platform or an external key managerKey identifiers, the manager's own master credential, and the unseal procedure
Self encrypting drives and tapeDrive controllers paired with a key managerPer device authentication keys and the steps to import them into replacement hardware
Remote or secondary tier keysA separate account or external key serviceThe account recovery path and an identity that works without the local directory
Access to the escrow itselfPassword manager, often a protected virtual machineAn offline retrieval route with named holders and a stated time to open

Two fields belong with every entry and are frequently omitted. The first is a date, so that anyone reading the escrow knows how stale it might be relative to the last key rotation. The second is a fingerprint or key identifier, which allows the escrow copy to be matched against the live key without the escrow being opened. That single field is what makes routine verification possible.

Testing the escrow without exposing it

The objection to testing an escrow copy is sound. Opening a sealed envelope and reading a passphrase aloud in a conference room turns a controlled secret into a widely known one. The way through is to separate the two things being tested, since they fail for different reasons.

The first is whether the retrieval path works. That is tested without touching the content at all: a named holder is asked to produce the sealed item, under the same conditions the real event would impose, and the elapsed time is recorded. If retrieval requires two people, a safe in another building, or a bank on a weekday morning, the drill surfaces that in an afternoon rather than during an outage. Building this into a drill that assumes production is unavailable keeps the dependency assumptions honest.

The second is whether the content is correct. Where the material has a fingerprint or key identifier, correctness is checked by comparing that value against the live system, which proves the escrow matches without revealing it. Where it does not, the test is a restore performed in an isolated environment by someone using only the escrow material. Afterward the material is rotated and resealed, which makes the exposure acceptable because the exposed value is retired. This belongs in the regular restore testing schedule rather than being handled as a one off project.

Where ARTESCA fits

ARTESCA is object storage software used as a backup target through an S3 compatible API, with immutability provided by S3 Object Lock. When a backup application encrypts data before writing it, the resulting objects are opaque to the storage, and the key involved is held entirely on the backup side. Object Lock protects those objects from deletion or overwrite, which is a separate property from whether anyone can still read them.

Because ARTESCA is deployed on infrastructure the customer runs, key management and the boundaries around it stay with the customer rather than a provider. That matters for escrow because the platform administrative credential, any storage side encryption settings and the access keys issued to backup clients are held locally, so they can be captured into the same escrow record as the backup application's own material.

What to put on a schedule and what to write down

Maintain one document listing every layer capable of blocking a restore, the key material each requires, the named holder of the escrow copy, the retrieval route that does not depend on the directory, and the date each entry was last verified. Keep a copy in printed form or on removable media held with the escrow itself, since a recovery document stored only on the estate under recovery repeats the original mistake.

On a quarterly rhythm, verify fingerprints against live keys and confirm that the retrieval path still works with the people currently employed. On an annual rhythm, run the isolated restore using escrow material only, then rotate and reseal. Whenever a passphrase changes, a key manager is replaced, or a backup server is rebuilt, treat the escrow as stale until it has been reissued, because a rebuilt console is exactly the situation where an outdated keystore export is discovered to be worthless.

Try ARTESCA free

Immutable object storage that scales from 20TB to petabytes. Deploy a working cluster in under an hour.

Start a free test drive