A backup console showing a long run of green checkmarks proves that data was written somewhere. It does not prove the data can be read back, that an application can start from it, or that the whole exercise fits inside the recovery window the business has been promised.
Ransomware has changed the shape of the problem. A recovery used to mean bringing back one server or one deleted folder. It now often means bringing back an entire tier of systems at once, into an environment that cannot be trusted, while the people who know the infrastructure best are locked out of their usual tools.
This article describes a practical restore testing program: three levels of test, a starting cadence for each, the measurements that actually predict recovery time, how to record results against RTO and RPO, and the reasons drills most often fail.
Recovery testing is the deliberate act of restoring data or systems from backup into a controlled location and confirming that the result is complete, consistent and usable. It is distinct from backup job monitoring, which only reports that the write phase completed, and from automated verification features that boot a snapshot or compare checksums. Those are confidence checks on one copy, not a rehearsal of the process a team would follow under pressure.
Recovery time is dominated by things a backup job never exercises: the throughput of the restore path, the order in which dependent systems must come back, the availability of credentials and DNS, and the time it takes people to make decisions. A restore drill measures all of them.
A workable program tests small things often, medium things regularly and large things rarely but seriously.
The lightest test restores an individual file, mailbox item, database table or S3 object from a randomly selected backup point and compares it to a known reference. The goal is catalog integrity: does the backup application know where the data is, can it retrieve it, and does the content match. These checks are cheap and can be automated. They should rotate across workloads and backup ages so that older restore points, including those in an immutable tier, get exercised.
The middle level restores a complete application, typically a database, a file server or a business system with its dependencies, into an isolated network segment. Success is not that the virtual machines power on. Success is that the application owner logs in, runs an agreed set of validation queries or transactions, and signs off. This is where latent problems surface, such as a restored database that boots but cannot reach its authentication source.
The heaviest test restores an entire tier of systems in parallel into a clean room: a freshly built, isolated environment with its own identity, networking and management tooling, assumed free of any attacker presence. It is the only test that measures aggregate restore throughput under realistic load, exposes ordering problems between systems, and rehearses the incident response plan. It should be run as an exercise with a scenario, a scribe and a debrief.
Cadence should follow the criticality of the workload and the cost of the test, and it should be adjusted from evidence. If a Level 3 rehearsal uncovers a dependency nobody documented, the next one should come sooner.
Regulated organizations may have prescribed minimums. The cadences above are not a standard, only a defensible place to begin.
A drill that ends with “it worked” has wasted most of its value. Five measurements turn a test into data that can be compared against agreed recovery objectives.
Measured in terabytes per hour across the whole restore path, not the peak of a single stream. Suppose an organization has 200 TB in its tier 1 systems and a stated RTO of 24 hours. It needs a sustained rate above roughly 8.5 TB per hour, with headroom, or the RTO is fiction regardless of how well the backups were taken.
Elapsed time from the decision to restore until the first tier 1 system is validated. This captures the fixed overhead of standing up the clean room and recovering identity, which no throughput figure reveals.
Elapsed time until the last system in the tier passes validation. This is the headline result of a Level 3 exercise.
The percentage of restored items that passed validation on the first attempt. Anything below 100 percent is a finding.
Object lock retention should be tested, not assumed. An operator using a standard, non-administrative identity attempts to delete or overwrite a locked restore point and records the refusal. The same attempt should be made with the backup application’s own service account, since that is the credential an attacker is most likely to obtain. If either succeeds, the backups are not protected.
| Test level | Scope | Starting cadence | Key metrics | Owner |
|---|---|---|---|---|
| Level 1: spot check | Single file, object or table from rotating backup points and ages | Weekly (tier 1), monthly (tier 2), quarterly (tier 3); automated | Pass rate; retrieval time; immutability refusal recorded | Backup administrator |
| Level 2: application restore | One application with dependencies, isolated segment, owner validation | Quarterly (tier 1), twice yearly (tier 2), annually (tier 3) | Time to validated application; pass rate; missing dependencies | Backup administrator with application owner sign-off |
| Level 3: mass restore rehearsal | Entire tier in parallel into a clean room, run as a scenario with debrief | Annually and after major infrastructure change | Sustained throughput; time to first usable system; time to full tier | Infrastructure lead with security and business continuity |
Each drill should produce a short record a non-specialist can read six months later: date, scenario, systems in scope, the restore point selected and its age (the achieved RPO for that test), the elapsed times above (the achieved RTO), pass rate, immutability check result, and every finding with an owner and due date.
The useful comparison is achieved against agreed. A tier 1 application with a 4 hour RTO that took 7 hours to validate is a documented gap: either the RTO changes, the infrastructure changes, or the business accepts the risk in writing. Tracking results across drills also shows whether restore throughput is keeping pace with data growth; the earlier post on backup repository sizing covers the capacity side.
The technical restore of data is rarely the part that breaks. The recurring causes are:
Scality ARTESCA is S3 object storage designed as a backup target, with S3 object lock immutability enforced by the storage rather than by the backup application. Because the immutable copies are online rather than on removable media or in an offline vault, they can be selected as restore points and read back directly through the backup application, which makes Level 1 and Level 2 tests against protected copies a routine operation.
ARTESCA is validated with the major backup applications, so drills use the same console, catalog and workflows operators already know. The immutability check described above can be carried out as a plain S3 delete attempt against a locked object, with the refusal recorded as evidence. The security and cyber resilience page describes the controls in more detail.
The practical takeaway: schedule the first Level 2 restore this quarter, measure the five numbers, and write down the gap between what was achieved and what was promised.