ARTESCA Blog | Backup, recovery and cyber resilience

Restore speed: How to find the slowest step

Written by Joshua Silvia | Sep 20, 2026, 3:34:43 AM

A restore expected to take two hours takes seven. The team watching has no way to tell whether the delay sits in the repository, the network, the backup server or the host receiving the data. The conclusion recorded afterward is that backup storage is slow, and the next budget cycle carries a line item that will not fix the problem. A restore is not one operation. It is at least six, and one of them involves reading from storage.

The phases have completely different remedies. Some are settings that can change before the next test. Some are structural and need different hardware or a different chain layout. Treating the restore as a single number makes it impossible to tell which kind of problem is in front of the team.

What a restore is actually made of

The first phase is the read from the repository. The backup software requests the blocks that make up the restore point, and the repository serves them. On object storage this is a sequence of object reads, and the size and number of those objects influences it more than the raw capability of the media.

The second and third phases are decrypt and decompress. Both are processor work, and both happen on whichever component the software designates, usually a proxy or the backup server itself. A restore running on an undersized virtual machine with two cores will be limited here regardless of what the repository can deliver, and nothing about the symptom points at the processor.

The fourth phase is chain rehydration. A restore point that is an increment cannot be served alone. The software reads the full it depends on and every increment between, then assembles the blocks in order. A point twenty increments deep reads far more data than a synthetic full, though both produce the same virtual machine. This is the phase most often mistaken for slow storage, because the storage is busy, just serving more data than anyone expected.

The fifth phase is transport, moving assembled data to the target host or datastore. The sixth is the write at the target, which lands on production storage that may already be serving users. The seventh, easy to forget, is registration and boot: creating the machine object, attaching disks, starting the guest, and waiting while the operating system and its applications come up. On a database server this phase can outlast all the others together.

Why the repository is rarely the answer

Storage gets blamed because it is the only component in the path that is visibly a storage product. The repository owns one phase, and in most slow restores it is not the largest. The larger candidates are chain depth, processor limits on the proxy doing decompression, and the write path into production.

There is also a measurement artifact worth knowing about. Many restore progress indicators report data written, not data read. When the software is rehydrating a long chain, progress appears to stall while a great deal of reading is happening. The operator sees a frozen percentage and concludes the repository has stopped responding, when it is working hard on data that will never appear in the progress figure.

Timing each phase with instrumentation that already exists

Almost none of this requires new tooling. The backup software logs restore session events with timestamps, usually including when the session started, when data transfer began, and when the virtual machine was registered and powered on. Subtracting those timestamps gives phase boundaries accurate enough to find the dominant one.

The rest comes from the hosts. Processor utilization on the proxy answers the decrypt and decompress question in one look, since a core pinned for the duration is the answer. Interface counters on the transport path show whether the network was saturated or idle. Latency and queue depth on the destination datastore show whether the write phase was the constraint. Gathered during a scheduled test rather than an incident, these take minutes.

Restore phaseWhat usually dominates itWhere the fix lives
Read from repositoryObject size and request concurrency, not media speedConfiguration first, then hardware
Decrypt and decompressProcessor capacity on the proxy or backup serverHardware, or moving the role to a larger host
Chain rehydrationHow many increments sit between the point and its fullConfiguration: chain type and full frequency
TransportGateway placement and the path between componentsConfiguration, sometimes network hardware
Write to targetProduction storage contention at the destinationHardware, or choosing a quieter destination
Register and bootGuest startup, service dependencies, application recoveryNeither, mostly application design and sequencing

The last row is the one teams underestimate. A file server registers and boots in a minute. A database that must replay a transaction log takes as long as the log demands, and no storage decision changes that.

Which phases configuration fixes and which need hardware

Chain rehydration is the clearest configuration lever. A restore from a point that sits far from its full reads the full plus every increment in between, so the frequency of fulls sets an upper bound on rehydration work. Shortening the chain shortens the restore and costs capacity, which is the trade that has to be made deliberately rather than inherited from whatever the job was created with.

Transport is usually configuration as well. Where the component that reads from the object repository runs, relative to both the repository and the destination, determines how many times data crosses the network. Gateway placement that was reasonable when the repository was local can become the dominant cost after the repository moves.

Decrypt and decompress are genuinely hardware. If the proxy is processor bound during restores, more cores is the remedy, and it is a cheap one compared with replacing storage. Encryption adds work here whether or not it is needed, so a team should know where encryption happens in its own data path, and separately should know that the key has to be available at restore time or the phase never starts at all.

Registration and boot are neither. They are a property of the workload. The realistic response is to avoid the phase where possible, by running the machine directly from the backup while it migrates in the background, which trades a fast return to service against reduced performance during the move. That has its own consequences after the guest is up and is not free.

Where ARTESCA fits

ARTESCA is object storage software used as a backup target, presenting an S3-compatible API and supporting immutability through S3 Object Lock, deployed on infrastructure the customer runs. In restore terms it serves the first phase: responding to object read requests for the blocks a restore point needs.

Because it sits on the customer's own network, transport is a local question rather than an internet one, and the path between repository, proxy and destination is under the same team's control. That matters mainly for large restores, where a path crossing a constrained link once per block becomes the dominant phase.

What ARTESCA does not influence is the rest of the list. Chain depth is set by the backup software, decompression runs on the proxy, and boot time belongs to the guest. A restore measured end to end against an on-premises object target will still spend most of its time outside the repository in most environments.

What to time and what to record

A quarterly restore test is worth more when someone records phase boundaries rather than a single duration. The useful record holds the restore point chosen and its distance from the nearest full, the timestamps at session start, first data transfer, transfer complete and guest login, plus proxy processor utilization and destination datastore latency during the run.

Repeat that with two workloads that differ deliberately, one small and shallow in its chain, one large and deep. The comparison between them isolates rehydration cost better than any single test does, and it produces a number the team can quote when someone asks how long a restore takes.

Finally, write down which phase was largest and what was done about it. Restore performance work accumulates slowly and is easy to forget between incidents, and a short history of dominant phases separates a considered purchase from a repeat of the same guess.