Instant VM recovery is measured in the demo by how quickly a machine answers a ping. That number is real, and it is also the least interesting part of the operation. Once the machine is running, it is running from the backup repository, and every read it issues is served by storage that was sized for sequential backup traffic. The recovery is not finished. It has barely started.
The mechanism is simple to describe. The backup file is presented as a datastore or a virtual disk, and the machine is powered on directly against it. Nothing is copied first, which is why the boot is fast. It is also why the machine now depends on two components that were never in the production data path: the repository, and a redirect area holding every write it makes while it lives there.
Teams that have tested only the boot step discover the rest during an incident. What decides whether the recovery holds is how long the repository can sustain random production reads, what happens when the write cache fills, how migration back behaves, and how many machines can run this way at once.
What the machine is actually running on
A backup repository is built for a workload that looks nothing like a running server. Backup jobs write large sequential streams inside a window, and restore jobs read them back the same way. A live virtual machine issues small scattered reads with no locality, continuously, for as long as it is powered on. Most repositories were sized for capacity and sequential throughput, not for concurrent random reads at low latency.
The effect depends on the media underneath. On a deduplicating appliance, rehydration on read can dominate, and the machine boots quickly and then feels slow in a way that is hard to attribute. On object storage, each guest read becomes a request against a block that may have to be assembled from a chain of increments, so the limiting figure is not bandwidth but requests per second and the latency of each one.
The useful test is therefore not a boot test. It is a boot followed by the application's own workload, with someone watching latency inside the guest rather than progress in the backup console. A machine that boots in two minutes and then takes forty seconds to open a record has not been recovered in any sense the business recognizes.
Where the writes land
The backup file is read-only, so writes from the running machine have to go somewhere else. Products handle this with a redirect target: a differencing disk, a cache volume, or a scratch area on production storage. Every block written, and every block rewritten, accumulates there for as long as the machine runs in this state.
Two properties of that area decide the outcome. The first is where it sits. Placed on the repository, it makes the repository serve reads and writes at once. Placed on production storage, it makes the recovery depend on the storage that may have just failed. The second is size. The cache grows with write activity rather than with the size of the machine, so a quiet file server can run for a week while a transactional workload fills a generous cache in hours.
Behavior at the limit varies by product and configuration and is worth confirming in the deployment at hand. In some arrangements the machine pauses. In others the session fails and the machine has to be started again from a restore point, losing everything written since the boot.
| Signal | What it usually means | What to check |
|---|---|---|
| Guest latency climbs while repository throughput stays low | A random read pattern the repository cannot serve, not a bandwidth limit | Request latency and queue depth at the repository |
| Redirect cache grows faster than expected | A write heavy workload, or an application rewriting the same blocks | Free space against the current growth rate, and behavior when it fills |
| Migration back stalls or slows overnight | Migration competing with the backup window on the same repository | Job schedules overlapping the migration, and repository task slots in use |
| Machine boots but an application service does not start | The restore point predates a configuration change, or a dependency is down | Guest application logs, and the order dependent machines were recovered in |
| Second and third machines much slower than the first | Repository saturation rather than a per machine problem | Total concurrent sessions against what the plan assumed |
The last row is the one that ruins plans, since the constraint is shared and it is reached without warning.
Migration back is the real recovery
Running from the repository is a temporary state that has to be exited deliberately. The exit is a migration that copies the machine's disks to production storage while it keeps running, then performs a short cutover in which the redirected writes are merged. The machine is briefly suspended at that moment, so the cutover is a scheduled event.
Duration is governed by two rates that have nothing to do with boot time: how fast the repository can read the full contents of the disks, and how fast production storage can absorb them. A machine that booted in ninety seconds may take many hours to migrate, since the migration is functionally a full restore running underneath a live service.
Migration also competes with the backup window. One that starts in the afternoon and is still running when jobs begin will slow down, and the jobs will slow with it. Where the repository sits behind a performance tier, it is worth knowing which tier the machine is reading from, since the two behave differently enough to change the plan.
How many machines can run this way at once
Instant recovery is designed for a small number of important machines, not for a site. Each one consumes a recovery session, repository task slots, cache space and hypervisor resources, and the repository runs out first. The practical limit is rarely a documented maximum. It is the point at which guest latency across all running machines becomes unacceptable, and it has to be found by testing.
That limit separates instant recovery from a recovery plan. Bringing four machines up to restore one service is a good use of the feature. Bringing forty up to restore a site is a different exercise, and it is worth planning mass recovery as its own scenario rather than assuming the instant path scales into it.
Where ARTESCA fits
ARTESCA is object storage software used as a backup target, presenting an S3 compatible API and supporting immutability through S3 Object Lock. When it holds the backup files behind a running recovery session, the guest read pattern arrives as object requests, so the characteristics that matter are concurrency and per request latency rather than the sequential throughput that governs ingest.
Whether a given backup product can run instant recovery directly from an object repository, or requires the restore point to be staged on a local performance tier first, varies by product and version. That is a question to settle in the deployment at hand, with a documented test.
Because ARTESCA runs on infrastructure the customer owns, the repository sits inside the customer's own network. The link between the hypervisor hosts and the repository is part of the recovery design and belongs in the same test.
What to settle before the next outage
Three numbers belong in the runbook for every machine that is a candidate for this treatment: the measured time to boot it from the repository, the measured time to migrate it back, and the size and location of the redirect cache. Only the first is usually known, and it is the one that matters least.
A quarterly exercise keeps those numbers honest. Recover a representative machine, run its real workload for an hour, watch guest latency rather than job status, then migrate it back and record the elapsed time. Repeat with three machines at once at least once a year. Treating this as part of an ongoing restore test program keeps it from being rediscovered during an incident.
Write down the exit criteria too. Decide in advance how long a machine may run from the repository before migration becomes mandatory, who authorizes the cutover window, and what happens if the cache reaches its limit first.
