A team that has just restored an entire virtual machine without difficulty is often surprised when a request for one deleted spreadsheet takes longer. The repository is the same, the backup is the same, and the restore is smaller by several orders of magnitude. The difference is not a fault or a misconfiguration. A whole machine and a single file ask the storage underneath for completely different things, and only one of those requests is shaped the way backup storage is built to serve.
The assumption behind the surprise is that restore time scales with restored size. It does for large restores, where throughput dominates and everything else is noise. It does not for small ones, where the fixed cost of locating data is the whole job.
The consequence is practical. A recovery plan that quotes one restore rate for the environment is wrong in both directions: pessimistic about full machine recovery and badly optimistic about the single file requests that make up most of the ticket volume.
Restoring a full machine is a large, mostly sequential read of a known set of blocks. The backup software already knows which blocks make up the restore point, in what order, and where they live. It can queue many reads at once, read ahead, and keep the network and the target datastore busy from the first second to the last.
That is the access pattern backup storage is designed around. Object storage rewards it, since large objects read in parallel streams use the available bandwidth efficiently and per request overhead is amortized across a lot of data. The restore becomes a throughput problem, and those respond predictably to more parallelism, more network, or a faster target.
Chain depth still matters, since blocks may be spread across a full and the increments after it, but the reads remain large and planned in advance. Whether a point sits at the end of a long chain or just after a fresh full is the tradeoff described in the choice between full and incremental recovery.
A file level restore starts somewhere else entirely. Before any file data moves, the backup software has to present the contents of the restore point as a file system. Depending on the product and the guest, that means loading a catalog or index, mounting the restore point, and having a file system driver read the guest's own metadata: partition table, file system superblock, directory structure, and then the extent records for the file itself.
Every one of those steps is a small read at an unpredictable offset. The directory holding the file might be a few kilobytes in the middle of a large object. The file's extents may point at blocks written on five different dates, meaning five increments in the chain and reads scattered across separate objects. None of it is sequential and little can be predicted far enough ahead to prefetch.
That is why the file's own size barely enters the calculation. Recovering a small document and recovering one a hundred times larger involve almost the same locating work, and the transfer finishes in moments either way. It is also why a slow file restore rarely gets faster by adding bandwidth.
Breaking a file level restore into stages makes the dominant cost visible, and it is usually not the stage anyone watches.
| Stage | What is happening | What makes it slow |
|---|---|---|
| Index or catalog load | The backup server locates the restore point and, where indexing is enabled, loads the file listing for it | Machines with very high file counts, or an index that was never built and has to be read from the backup instead |
| Mount of the restore point | The restore point is presented as a readable volume, often through a helper appliance or a mount server | Mount server placement, available resources, and how far it sits from the repository on the network |
| Guest file system walk | Partition, file system metadata and directory entries are read to find the file and its block list | Deep directory trees, large directories, and metadata scattered across the chain |
| Block reads across the chain | The file's blocks are fetched from whichever full and increments hold the most recent version of each | Chain depth, small request sizes, and per request latency on the repository |
| Copy to the destination | The recovered file is written to the original location or an alternate one | Permissions, the destination file server, and whether the original path still exists |
Four of those five stages are unaffected by the size of the file. That is the whole explanation for the case that looks absurd on paper, where a full machine finishes while a single document is still being located.
Several configuration choices move file level restore times considerably, and none are the ones that help a full machine restore. Guest file system indexing is the largest. When the file listing is captured at backup time, the search is answered from the backup server rather than the repository, and the mount and walk are shortened. Indexing costs backup window time and some space.
Mount server placement is second. A mount or helper appliance reaching the repository over a congested or high latency link turns each of thousands of small metadata reads into a slow operation, and the effect compounds because file system logic serializes them. Placing the mount function close to the repository in network terms is often the change with the largest effect.
Chain structure is third. A file restore from a point deep in a long incremental chain touches more objects than one taken shortly after a full. This is closely related to how backup chains are laid out on object storage, and to the more general question of which step in a restore is actually the slowest, which is rarely the step being optimized.
ARTESCA is object storage software used as a backup target, deployed on infrastructure the customer runs and reached through the S3 API. Both restore types read from the same repository, but they generate different request profiles against it: a machine restore issues a modest number of large reads, while a file restore issues many small reads at unrelated offsets, often as ranged requests inside larger objects.
Because the repository sits on the customer's own network, the path between the mount or helper function and the storage is under the customer's control, which is the part of a file level restore most exposed to latency. Where that function runs, and how many hops it takes to reach the repository, affects file restores far more than it affects machine restores.
Where restore points are held under S3 Object Lock, the lock governs deletion and modification rather than reading, so an immutable point is read like any other. Immutability is why a point still exists when needed, not a factor in how long either restore takes.
Two restore rates belong in the recovery documentation, not one: a throughput figure for whole machine recovery, and a typical elapsed time for a single file request, each recorded from a real restore in the environment rather than estimated. They are different numbers with different causes, and quoting one for the other is where expectations break.
The configuration decisions are worth revisiting as a pair. Which machines have guest file system indexing enabled, and whether that list still matches where file requests actually come from. Where the mount or helper function runs relative to each repository. Which restore points a file request is likely to reach, given the chain and retention in place. For urgent cases involving a whole machine, whether running a machine directly from the backup is available is a separate path with its own constraints.
It is also worth recording which of the two cases the environment sees most. In most estates the file request is routine and the machine restore is rare, yet effort goes to the rare one because it is the one that gets rehearsed. The ordinary request deserves the same attention, since it is the one the team will answer this week.