Storage economics

Cloud vs. on-prem backup: Price the recovery scenario

How to compare cloud and on-prem backup on the cost of reading data back, not the cost of storing it.

7 min read
Two converging teal light paths crossing a dark data center aisle between server racks toward glowing storage blocks

Cloud and on-premises backup storage are almost always compared on what it costs to keep data: a price per terabyte per month against a capital purchase spread over a few years. That comparison is tractable, which is why it gets made. It also describes the quiet years. The difference between the two models shows up on the day the data has to be read back, and that day is the one the platform exists for.

The asymmetry is structural rather than a matter of pricing policy. Writing backups is a steady, low-urgency flow. Reading them back during an incident is a single large burst under time pressure, and the two models charge for that burst in different currencies. One bills it, the other constrains it with hardware already bought.

Pricing the recovery means defining the event first. Without a stated scenario the comparison drifts into generalities, since a single file restore and a full-site rebuild sit at opposite ends of the same price list.

Define the event before pricing it

A usable comparison starts with one written scenario, specific enough that both vendors quote the same thing. The minimum content is what is lost, what has to come back, in what order, and by when. A site-level loss of the virtualization estate is the common choice, since it exercises the whole path at once.

From that scenario come three numbers, none of which is the repository size. The first is the volume that actually has to move, meaning the current restore points for the machines in scope rather than all retained history. The second is the deadline, as the time by which the first workloads must be serving and the time by which the rest must follow. The third is the destination, since data has to land somewhere with room for it.

Those three are enough to make the models comparable. Work on recovering many machines at once rather than one is where the ordering and the destination usually get settled, and without that ordering the recovery has no shape to price.

What a full-site restore requires in each model

On owned infrastructure the recovery cost has mostly already been paid. The read happens across a local network at whatever rate the platform, the fabric and the restore target support, and the limiting factor is capability rather than billing. What can still be spent is time, and time turns into cost through the outage rather than an invoice.

In a cloud repository the same read is metered. Bytes leaving the provider are charged, requests are charged, and the rate at which bytes can leave may be shaped by the access tier the data was written to. A colder tier chosen to reduce the storing cost often carries a retrieval charge, a minimum storage duration, or a delay before the first byte is available, and those conditions apply whether or not they were considered when the tier was chosen.

A hybrid case is worth pricing separately, where the repository is in the cloud and the recovery target is on-premises. That combination pays both the metered read and the transit time across a link usually sized for the daily backup flow rather than for a single large recovery.

The line items that differ

Most of the difference is concentrated in a handful of items. Setting them side by side shows which are invoices and which are constraints, since a comparison that counts only invoices favors the model whose limits are physical.

Element of the recoveryIn a cloud repositoryOn owned infrastructure
Reading the bytes backMetered, usually per unit of data leaving the providerAlready paid for, limited by platform and network capability
Requests and operationsBilled per operation, and a restore issues a large number of themNot billed, but concurrency limits still apply
Access tier of the dataColder tiers may add a retrieval charge and a delay before data is availableUniform access, unless a tiering policy was configured deliberately
UrgencyFaster retrieval options are typically priced above standard onesCannot be bought at the moment of need, only designed in beforehand
Landing the dataTarget capacity and its own charges, plus transit if the target is elsewhereCapacity on existing hosts, which must have been reserved
Failed or repeated attemptsCharged again, since a partial read still moved bytesCosts time only

The repeated attempt row converts an operational risk into a financial one. A restore that stalls and is retried has moved data twice, and a recovery under pressure is not the moment to be counting attempts.

The concurrency row matters in the other direction. Owned platforms do not bill for parallel reads, but streams are finite, and the practical recovery rate is set by whichever component saturates first rather than by the headline capability of any one. Analysis of which step in the restore path is actually the slowest turns that into a number.

Urgency, first byte and sustained rate

Two timing assumptions decide whether a recovery plan is real. Time to first byte is how long after the request data starts arriving, and sustained rate is how quickly it continues once started. A plan built on one without the other is incomplete in a way that only shows during a test.

Cloud models expose these as product choices. Standard access usually starts quickly, archival tiers may not, and expedited retrieval exists as a priced option. The consequence is that urgency is purchasable, which means it has to be budgeted in advance or authorized during an incident by somebody able to approve unplanned spending.

On owned infrastructure the same two properties are fixed by design decisions made at purchase. Time to first byte is short. Sustained rate is whatever the platform and the network deliver, and it does not improve because the situation is urgent. Both models therefore need a stated assumption and a test, since an untested sustained rate is a guess. A program of measuring restores on a regular cadence is the only source of a defensible number for either model.

Where ARTESCA fits

ARTESCA is object storage software used as a backup target, presenting an S3-compatible API and deployed on infrastructure the customer runs, in capacities from roughly 50 TB upward into the petabyte range. Because the platform is owned and local, reads during a recovery are not metered events, and the recovery rate is a property of the hardware, the network and the backup software configuration rather than of a service tier.

It also supports S3 Object Lock, so an immutable copy can be held on the same platform the recovery reads from. That places the read path and the protection mechanism in one location, removing the transit component from the hybrid case above.

What remains to be established locally is the sustained rate, since it depends on the specific deployment. It is measurable in a proof of concept using the organization's own backup software and workloads, and that result is the figure the recovery plan should carry.

What to ask each vendor

The comparison becomes honest when both sides answer the same scenario. For a cloud service, the questions are which charges apply to reading the stated volume, whether the access tier adds a retrieval fee or a delay, how operations are counted, what a faster retrieval option costs, and whether any minimum storage duration applies to data deleted after a restore. Each answer should be traceable to the provider's published structure rather than to a summary, and the method for assembling those into an estimate of a full-site restore is worth documenting.

For an on-premises platform the questions are about capability: the sustained restore rate observed with the backup software in use, how it changes with concurrent streams, what the network between repository and recovery target supports, and how quickly staging capacity can be added. None of these are invoices, which is precisely why they get skipped.

What belongs in writing is one page per model: the scenario, the volume, the assumed first byte and sustained rate, the resulting time to complete, the charges that apply, and the date each assumption was last tested. Reviewing it yearly, and after any change in tier, retention or estate size, keeps a decision made once from quietly becoming wrong.

Try ARTESCA free

Immutable object storage that scales from 20TB to petabytes. Deploy a working cluster in under an hour.

Start a free test drive