Backup & recovery

Backup success vs. recovery success: What to verify

A green job proves data was written. What each layer of verification proves, what it misses, and how often to run it.

7 min read
Dark data center corridor with nested glowing archways lit in cyan along a single reflective path

Every job in the console is green, and has been for months. Then a server is restored after a storage failure and the application that depends on it will not start, because the database it uses was captured mid-transaction and the backup software had no way to know. Nothing failed. Nothing was reported. The green status was accurate about what it measured, and what it measured was narrower than the team assumed.

A successful job means the software read the blocks it was asked to read, wrote them where it was told, and encountered no error it recognizes. That is a statement about a write operation. It is not a statement about whether the data can be read back, whether it is internally consistent, or whether an application will accept it.

That gap is not closed by a better backup product. It is closed by verification, which comes in layers costing very different amounts of time and producing very different levels of assurance.

What a green job actually proves

A completed job proves the source was reachable, a snapshot was taken and released, data moved across the network, and the target acknowledged the writes. With per-object checksums it also proves the bytes written match the bytes read at the time of writing. These are real guarantees and they eliminate a category of failure, since a job that cannot reach its source or target says so.

What it cannot prove is anything about the state of the data inside those blocks. If a guest was quiesced improperly, the backup holds a crash-consistent image of a running database, which is valid data and may or may not replay cleanly. If an application writes outside the protection scope, the backup is complete for what it covers and silently incomplete for what the application needs. If the restore path has a problem, an unreachable key, a missing catalog entry, a corrupted index, none of that surfaces during a backup.

There is also a timing gap. A job verifies data at the moment it was written. Whatever happens afterward, in storage, during a copy to a second location, or across a retention period measured in years, falls outside the status the job reported.

The layers of verification and what each one costs

The cheapest layer is job status, already described. It costs nothing because it happens anyway. Its value lies in noticing the absence of a job rather than confirming the presence of good data, which makes a weekly look at which jobs did not run more useful than a look at which went green.

The second layer is automated health checking. Most backup platforms can read a restore point back and verify its checksums or internal structure without restoring anything. This catches storage-level corruption and finds a chain whose dependencies are broken. It is cheap in effort since it is scheduled once, though it consumes repository read capacity and can run into production hours if pointed at everything at once.

The third layer is a mount and file read. The restore point is presented as a filesystem and a person or a script opens something from it. This proves the chain rehydrates, the key is usable, the catalog resolves, and the filesystem inside the image is intact. It takes minutes per workload and covers most of the distance between a checksum and a real recovery.

The fourth layer is a full restore to an isolated location, which proves the data path works end to end and gives the only honest answer to how long a restore takes. The fifth is application-level validation: starting the restored workload and having someone who understands it confirm that it works, the data is current, and dependent systems can reach it.

What each layer leaves untested

Choosing between layers is easier when what each one does not cover is written down rather than assumed.

Verification layerWhat it provesWhat it still does not prove
Job statusThe job ran, data moved, the target acknowledged writesThat the data is consistent, complete in scope, or readable later
Checksum or health checkStored blocks are intact and the chain resolvesThat an application can use the contents, or that the key is available
Mount and file readThe chain rehydrates, keys work, the filesystem is readableRestore duration at full size, or application consistency
Full restoreThe whole data path works and the real duration is knownThat the workload functions, or that dependencies resolve
Application validationThe workload runs and the data is usable by its ownerThat other workloads recover in the same conditions or timeframe

Each layer proves the previous one plus a narrow addition, and each addition costs roughly an order of magnitude more time than the one below. That is why a program running only the top layer runs it rarely, and one running only the bottom layer learns nothing new.

Deciding which layers run and how often

A workable pattern is frequency inversely proportional to cost. Job status is reviewed continuously, with attention on jobs that did not run. Health checks run on a schedule the repository can absorb, often weekly and staggered so they do not all land together. Mount and file read happens monthly, taking a different workload each time. Full restores and application validation happen quarterly or twice a year, against workloads chosen because they matter.

Which workloads get the expensive layers is harder, since testing everything is impossible and testing the same easy virtual machine every quarter proves nothing new. Choosing representative workloads means covering the axes that actually change restore behavior rather than the convenient ones, and rotating the set so a long-untested system does not go unnoticed.

Databases deserve a separate decision. An image backup of a database server may restore perfectly and still leave the database needing recovery work, and point-in-time requirements are usually finer than a nightly image provides. Those workloads need validation at the database layer, and the person who can confirm it is rarely the person who runs the backup.

The same applies to granularity. Most real recovery requests are for a few files, not a whole machine, so a program that only tests whole-machine restores is testing the rarer case. The difference between file restores and machine restores shows up in duration and in which parts of the stack are exercised.

Where ARTESCA fits

ARTESCA is object storage software used as a backup target, deployed on infrastructure the customer runs, presenting an S3-compatible API and supporting immutability through S3 Object Lock. In verification terms it is responsible for the durability and integrity of stored objects, and for serving them back when a health check or a restore asks for them.

That covers the lower verification layers and no more. Object integrity is a storage property. Application consistency is a property of how the backup was taken, and usability is a property of the application. A team using immutable object storage still has to test restores, because immutability guarantees the data cannot be altered, not that the data was correct when it arrived.

Since the storage runs on the customer's own infrastructure, the read capacity that health checks and test restores consume belongs to the same team. That makes cadence a local question about when spare read capacity exists rather than a question about transfer charges.

What to run on a schedule and what to record

The decision to record is which layer applies to which workload and at what interval, written once where the whole team can read it. A tier where job status and health checks suffice, a tier that gets a monthly mount, and a small tier that gets a full restore with application validation is enough structure for most mid-sized estates.

Each exercise should leave behind the workload tested, the restore point used, the layer performed, the elapsed time, and anything that had to be looked up or fixed during the attempt. That last field is the valuable one, since the step someone had to search for during a calm test is the step that stalls a real recovery. Over a year these records answer the question a green dashboard cannot, which is not whether backups ran but whether recovery has been demonstrated recently enough to be believed.

Try ARTESCA free

Immutable object storage that scales from 20TB to petabytes. Deploy a working cluster in under an hour.

Start a free test drive