Backup & recovery

Restore testing: How to choose representative workloads

How to pick the few workloads whose restore results stand for the rest of the estate, and when to rotate them.

7 min read
A small group of brightly lit teal blocks selected from rows of dimmer blocks in a dark data center

Restore testing usually fails at the first step, and the first step is not scheduling and not tooling. It is the choice of what to restore. A team picks the file server because it is easy, or the domain controller because it is important, and then treats a successful restore of either as evidence that the estate can be recovered. Those two results say very little about the other two hundred machines, and nothing at all about the ones that differ most.

The instinct is to test more machines, which costs more and still does not solve the problem. A restore test is only useful if its result transfers to workloads that were not tested, and transfer depends on similarity along the dimensions that govern restore behavior. Importance to the business is not one of them. Two equally critical machines can restore in completely different ways, and two unimportant ones can be the only honest proxies for the hardest case in the estate.

Selection is therefore a sampling problem. The question is which small set of workloads covers the range of restore behaviors present, so that a result from inside the set is a reasonable prediction for anything outside it. The wider program around that set is treated separately in the recovery testing program as a whole.

Why criticality is the wrong sampling axis

Most test lists are built from a tiering exercise. The tier one systems are listed, someone picks two or three, and those become the restore tests. That logic is sound for deciding what gets recovered first in an incident and close to useless for deciding what gets tested, because criticality is a property of the business while restore duration is a property of the data and the storage path.

The practical result is a test set clustered in one corner of the estate. Tier one systems tend to be similar: moderately sized, application consistent, on the fastest source storage, backed up by the best tuned jobs. Restoring three of them exercises one behavior three times. The workloads that will actually surprise the team are further down the list, absent from the sample precisely because they matter less.

A sample built for coverage looks unbalanced against a tier chart, and it should. Some members will be machines nobody would rush to recover. Their value is that they sit at the edge of a dimension nothing else in the sample reaches.

The axes that change how a restore behaves

Five dimensions account for most of the variation in how a restore proceeds. They are worth naming separately because a workload can be extreme on one and unremarkable on the others, and a sample that mixes them up tests the same combination repeatedly.

AxisWhy it changes restore behaviorHow to sample it
Restored sizeLarge restores are dominated by sustained throughput; small ones are dominated by fixed overhead that does not shrinkOne workload near the largest in the estate and one well below average, not two mid sized ones
File count and densityMillions of small files change the work from streaming blocks to metadata operations, whatever the total sizeThe densest file system present, chosen by object or file count rather than by capacity
Application consistencyDatabase and mail systems restore into a state that still needs replay, attach or consistency work before service returnsAt least one workload whose recovery continues after the data lands
Chain depth and restore point ageA point deep in an incremental chain requires more of the chain to be read than the most recent oneA restore from an old point on a long chain, not only from last night
Source and target storage typeReading from an object repository, a local disk repository or a copy at another site produces different read patterns and pathsOne workload per repository type in use, including any capacity tier

The last axis is the one most often collapsed. A workload tested from the local repository says nothing about the same workload restored from a remote copy, since the data path differs end to end.

Covering the axes without testing everything

Full coverage of five axes is not required, and attempting it produces a test list nobody completes. What is required is that each axis has a workload near each end of it somewhere in the set, and that no single workload carries three extremes at once, since a failure then cannot be attributed to any of them.

A workable set is usually four to seven workloads. One very large machine, one very small one, one file system with an extreme file count, one application consistent system, one restore from a deep chain, and one drawn from each repository type. Several roles can be filled by the same machine where the combination occurs naturally, which is how the list stays short.

Two cases deserve their own members rather than being folded in. Databases behave differently enough that recovery beyond the machine level is a separate behavior, not a harder file restore. And a single file pulled from a large machine exercises a path a whole machine restore never touches, which is why file level and machine level restores both belong in the set.

When a representative stops representing

A chosen workload represents its neighbors only while it still resembles them. That resemblance decays quietly, and the test keeps passing while it does, which is the failure mode worth guarding against here.

The common causes are structural rather than dramatic. The representative was moved to different source storage while its neighbors were not. Its backup job was split, shortening its chain. A retention change lengthened everyone else's chain but not its own. Data grew on the workloads it stands for while it stayed flat. In each case the sample still tests successfully and no longer stands for anything.

A review of the sample against the estate catches this. The useful comparison is not whether the test passed, it is whether the representative is still near the edge of the axis it was chosen for, and whether anything new has appeared beyond that edge. Newly onboarded systems are the most frequent cause, since they arrive outside the range the set was built to cover.

Where ARTESCA fits

ARTESCA is object storage software used as a backup target, deployed on infrastructure the customer runs and presented to the backup software through the S3 API. When it holds part of an estate's restore points, it is one of the storage types on the source axis, which means at least one member of the sample should be restored from it rather than from a local repository alone.

Because restore points can be written under S3 Object Lock, the set should include at least one workload whose chosen point is still inside its immutability period, since that is the state most points are in when needed. A test that only reads unlocked older data exercises an untypical condition.

Where ARTESCA sits behind a scale out repository as a capacity tier, the same workload has two restore paths depending on which tier holds the selected point. Treating those as one case understates the range, so the sample should identify which tier each chosen point comes from.

What to record about the chosen set

The selection deserves a written record, because the reasoning behind it is what makes the result transferable, and that reasoning is invisible in a list of machine names. For each member, the record should hold the axis it represents, the value that put it at that edge, and roughly how many workloads it stands for.

Rotation is best handled by axis rather than by machine. Replacing one member at a time, keeping its axis filled, keeps the set stable while preventing the tests from becoming a routine that only proves one path works. A member in place through several equipment or retention changes is the first candidate for replacement.

The set should be revisited whenever something could move a workload along an axis: a new application onboarded, a repository added, a job split, a retention extended, or a migration of source storage. Those events also drive planning for a large scale recovery, and the same inventory serves both, which keeps selection from becoming maintenance nobody owns.

Try ARTESCA free

Immutable object storage that scales from 20TB to petabytes. Deploy a working cluster in under an hour.

Start a free test drive