Backup & recovery

Mass VM recovery: How to plan for competing restores

Why fifty concurrent restores are limited by shared bottlenecks, how to measure aggregate throughput, and how to order recovery.

7 min read
Dark data center with many light paths converging from storage racks into one narrow bright channel in cyan

A single virtual machine restores in twenty minutes during a test, so fifty should take somewhere between twenty minutes and a long afternoon, depending on how many run at once. That arithmetic is the most common planning error in recovery work. Fifty concurrent restores do not behave like fifty independent ones, because they share a repository, a network path, a set of production datastores and a hypervisor that schedules their tasks.

The usual response is to measure one restore carefully and multiply. The result looks defensible, gets written into a plan, and is wrong by a wide margin. Sometimes it is optimistic because it assumed perfect parallelism, sometimes pessimistic because it assumed none. Neither number came from anything measured under contention.

Why fifty restores are not fifty times one restore

A restore is a pipeline. Data is read from the repository, decompressed and decrypted by a data mover, pushed across a network, and written into production storage, where the hypervisor registers and starts the machine. Run one restore and each stage has the whole pipeline to itself. Run fifty and every stage is divided among them, with the narrowest stage setting the pace for all.

That changes the shape of the estimate. The total is not the sum of individual restore times, nor the longest individual time. It is roughly the volume of data to be moved divided by the aggregate throughput of the narrowest shared stage, plus the fixed overhead each machine incurs on its own. The per machine figure matters far less than the aggregate figure under load, and the aggregate figure is the one nobody has.

Two effects make it worse than a simple division. Restores rehydrate data, so what lands on production storage is larger than what was read from the repository, and the write side can become the constraint even where the read side looked ample. And parallelism has a ceiling: past a certain number of concurrent tasks, aggregate throughput stops rising and queueing begins, so more concurrency adds latency and timeout risk without adding speed.

Where the contention actually shows up

Repository read is the first shared stage. Aggregate read throughput, the number of concurrent streams the repository will serve, and for object targets the request concurrency and object size in use all cap what can be pulled at once. That ceiling is usually discovered during the first real event, because single restore tests never approach it.

The network path is the second, and the most frequently overlooked, since it may be a single uplink shared with whatever production traffic still exists. Datastore write is the third, and during a mass recovery it does something it never does in normal operation: accepting sustained writes across every volume at once while also serving the machines that already came back.

The fourth is the hypervisor itself. Platforms limit concurrent provisioning, registration and disk operations per host and per datastore, and backup software adds task slots, proxy counts and per repository limits on top. The effective concurrency is the smallest of several numbers rather than whichever one was configured most recently, so concurrency settings on the backup side deserve to be read together with the hypervisor limits.

Shared stageWhat limits itWhat to measure before the day
Repository readAggregate read rate and concurrent stream or request limitsThroughput with five restores running, then ten, then twenty
Network pathThe narrowest link between repository and hosts, shared with productionUplink capacity on that path and what else uses it during an event
Data moverProcessor time for decompression and decryptionUtilization on the proxy or gateway at peak concurrency
Datastore writeSustained write across all volumes while recovered machines runWrite throughput with several restores landing at once
Hypervisor tasksConcurrent operation limits per host and per datastoreThe configured limits, and which one is lowest in practice
Target capacityFree space on the destination, including rehydrated sizeFree space today against the full restored footprint

Finding the concurrency that actually helps

The measurement that settles all of this is short and rarely done. Restore five machines simultaneously to an isolated location and record aggregate throughput. Repeat with ten, then twenty. Aggregate throughput rises, flattens, and eventually falls as queueing and retries take over. The number at the flattening point is the estate's real parallel capacity, and it is the number that turns a recovery plan into a schedule.

With that figure and the total restored footprint, the wall clock estimate for a full recovery follows, and so does the more useful answer: how much can be back within four hours, within eight, within a day. Those are the questions asked during an incident, and a single machine test cannot answer them.

The same measurement exposes which stage plateaued first, which is the only honest basis for spending money on the problem. A repository ceiling, a shared uplink and a datastore write limit call for three different purchases, and the one chosen without measurement is the repository.

Building the order from dependencies rather than importance

Recovery order is almost always written by business importance, which produces a sequence that cannot execute. The finance application is first on the list, so it is restored first, and it fails to start because its database server is thirtieth and directory services are not back at all. Importance says what matters. Dependency says what can start.

The dependency layer comes first in every estate: directory and DNS, a time source, certificate services, address assignment, and the backup infrastructure itself if it is being recovered too. Then data services, then application servers, then anything that talks to them. Within that structure, importance decides the order among peers rather than the order overall.

Building the map takes a conversation rather than a tool. For each service, the question is what must already be running for it to start cleanly, and the answer comes from the people who operate it. That map is what makes a drill with production assumed down useful, since a drill that restores machines in isolation never tests whether the order works.

Boot behavior belongs in the same plan. Fifty machines starting at once produce a storm of reads on storage that is already absorbing restore writes, so staggering start-up is worth as much as staggering the restores themselves.

Where ARTESCA fits

ARTESCA is object storage software used as a backup target, and in a mass recovery it occupies the repository read stage of the pipeline. Restores read from it concurrently, so what matters is its aggregate read behavior under many simultaneous streams rather than the speed of any one restore.

Because it runs on infrastructure the customer operates, the path between it and the hypervisor hosts is a local network that can be sized, measured and changed, and a full site recovery does not meet a metered circuit on the way. That differs from designs where a full site restore pulls data back from a remote service and the same exercise carries both a transfer time and a retrieval cost.

The storage layer does not remove the other constraints. Network path, datastore write capability and hypervisor task limits are properties of the environment, and they will set the pace of a mass recovery regardless of which target holds the backups.

What to fix in advance and what to decide on the day

Fix in advance the things that cannot be changed under pressure: the measured aggregate restore throughput, the dependency ordered recovery list, free capacity at the destination including rehydrated size, credentials and encryption keys reachable without the systems being restored, and the concurrency limits on both the backup side and the hypervisor side written down in one place.

Leave for the day the decisions that depend on the event: which machines are deferred, whether any are restored with reduced processor and memory to fit more in parallel, which restore point each group uses, and whether some services run directly from backup storage to shorten time to service, accepting that a machine running from the repository still has to be migrated back later.

Review the aggregate throughput measurement whenever the estate changes materially, and record the date beside the number. A figure measured two hardware refreshes ago is worse than no figure, since it will be trusted.

Try ARTESCA free

Immutable object storage that scales from 20TB to petabytes. Deploy a working cluster in under an hour.

Start a free test drive