Storage economics

Usable backup capacity: Compare quotes on equal terms

Raw, usable and effective capacity are different numbers. How to restate competing backup storage quotes in one common unit.

7 min read
Three differently sized stacks of glowing teal storage blocks aligned to one measuring line in a dark server room

Three quotes for the same backup repository arrive with three different capacity figures, and none of them mean the same thing. One states raw capacity, one states usable capacity after a protection scheme the quote does not name, and one states an effective figure that already assumes a data reduction ratio. Compared as printed, the cheapest per terabyte is usually the one that counted most generously, not the one that delivers most.

The comparison problem is separate from the sizing problem. How much capacity the environment actually needs comes from protected data volume, change rate and retention, and that calculation is set out in backup repository sizing. What follows assumes the requirement is already known and deals only with restating competing quotes in a single unit so that cost per unit means something.

This matters more for backup than for primary storage because backup data has already been through one or two reduction stages before it reaches the target. A ratio measured on unreduced data does not transfer to a repository holding compressed, deduplicated backup files.

Three numbers that all get called capacity

Raw capacity is the sum of the drive labels. It is the largest number available, it is rarely what anyone gets to use, and it is the figure most likely to appear in a headline price per terabyte. It is still worth collecting, since it is the only unambiguous figure and every other number traces back to it.

Usable capacity is what remains after the data protection scheme, formatting overhead, reserved space and any spare capacity held for rebuilds. The gap between raw and usable is set almost entirely by the protection scheme chosen, and that choice is a design decision rather than a property of the hardware. Two vendors quoting the same drives can quote materially different usable figures because they assumed different schemes.

Effective capacity is usable capacity multiplied by an assumed data reduction ratio. It is the most useful number when the assumption holds and the most misleading when it does not, since the multiplier applies to the whole figure and magnifies any error in it. It is a forecast rather than a specification, and belongs in the comparison labeled as one.

The protection scheme behind the usable figure

A usable figure without a named scheme is not comparable to anything. The specific questions are which erasure coding or replication scheme produces the number, how many drive or node failures it tolerates without data loss, and whether the scheme changes as the system grows, since some designs become more efficient at larger node counts and others do not.

Rebuild reserve belongs in the same conversation. Some quotes hold capacity back so a failed drive or node can be rebuilt without the system filling, and some assume that reserve comes out of the customer's headroom. That difference is invisible unless it is asked about directly.

Resilience is a dimension of the comparison rather than a footnote to it. A quote that reaches a larger usable figure by tolerating fewer failures is not cheaper, it is different, and recording tolerated failures beside the usable figure keeps that visible later.

Where data reduction is counted

The question to settle for each quote is whether the stated figure sits before or after deduplication and compression, and whose deduplication is being counted. Backup software commonly compresses and deduplicates before data leaves the backup server, and a storage target then claims its own ratio on top of that. Counting both is double counting, and it is the single most common way a comparison goes wrong.

Ratio claims deserve their own scrutiny, since the data type a ratio was measured on determines whether it transfers at all. The reasoning behind deduplication ratio claims applies directly here. For comparison purposes the safest treatment is to strip every assumed ratio out, compare the quotes on usable capacity, and then reintroduce reduction as a separate, clearly labeled assumption applied equally to all of them.

What the quote statesWhat it may leave outQuestion to put to the vendor
Raw capacityProtection overhead, formatting, rebuild reserveWhat usable capacity remains once the proposed scheme is applied
Usable capacityThe scheme that produced it and the failures it toleratesThe exact scheme, its failure tolerance, and whether it changes with growth
Effective capacityWhether the ratio was measured on data like thisThe ratio assumed, the data it came from, and the usable figure beneath it
Capacity after reductionWhether two ratios have been multiplied togetherWhether backup software reduction is counted again at the target
Licensed capacityWhich quantity the license metersWhether the license counts stored data, written data or protected front end data
Immutable copyLocked versions retained beyond job retentionWhether the figure includes versions the lock holds past their expiry

The immutable copy and the unit that survives

An immutable copy is where quotes most often diverge without saying so. A locked object cannot be deleted before its retention period expires, which means versions accumulate beyond what job retention alone would suggest, and the effective lock is frequently longer than the number typed into the backup job. How those settings compound is covered in the discussion of immutability and retention settings, and the capacity consequence has to be inside the comparison rather than discovered afterward.

The practical step is to ask each vendor whether the quoted figure assumes immutability is enabled, and if so, at what retention period. A quote sized for a mutable repository and one sized for a locked repository are not competing for the same job, even when both name the same capacity.

Once those questions are answered, the common unit is usable terabytes at a stated failure tolerance, with immutability assumed on, before any data reduction. Every quote can be restated in that unit, cost per unit follows, and reduction and license terms sit beside it as separately labeled assumptions rather than being folded into the headline. The wider set of questions that belong in a formal comparison is covered in procurement comparison.

Where ARTESCA fits

ARTESCA is object storage software used as a backup target, deployed on infrastructure the customer runs, in a range of roughly 50 TB to 8.5 PB. Because the software and the hardware are separate line items, the usable capacity question resolves against the server and drive configuration proposed, and that configuration can be inspected directly rather than inferred from a product family datasheet.

Immutability is provided through S3 Object Lock against an S3 compatible API, so the retained version behavior that drives capacity is the standard object lock behavior the backup software already expects. That makes the immutability assumption in a quote something that can be stated explicitly and checked in a test, rather than a vendor specific qualifier.

For comparison work the useful consequence is that the questions in the table above have addressable answers: a named protection scheme, a drive count, a rebuild reserve and a stated lock period. Where an environment is genuinely multi site or far beyond that capacity range, RING covers that territory.

Writing the comparison down

Build the comparison as one sheet with a row per quote and a column for each normalized quantity: raw capacity, usable capacity, the named protection scheme, tolerated failures, rebuild reserve, whether immutability is assumed, the lock period assumed, any data reduction ratio claimed, and the license metering basis. Quotes that cannot fill a cell get a blank rather than a guess, and a blank is itself a finding.

Send the same written questions to every vendor and keep the answers with the sheet. Verbal clarification does not survive the months between shortlist and purchase order, and whoever defends the decision later is often not the person who heard it.

Finally, record the unit the decision was made in. When the next expansion is quoted, possibly by a different account team, comparing it against the original means comparing like with like, and that is only possible if the original unit was written down rather than reconstructed from memory.

Try ARTESCA free

Immutable object storage that scales from 20TB to petabytes. Deploy a working cluster in under an hour.

Start a free test drive