ARTESCA Blog | Backup, recovery and cyber resilience

Backup deduplication ratios: What to verify in a quote

Written by Joshua Silvia | Sep 20, 2026, 3:59:11 AM

A deduplication ratio in a quote reads like a property of the product, the way rack units and drive slots are properties. It is not. It is the result of one algorithm run against one particular set of backup data, held for one particular length of time. Change the workload mix or the retention and the number changes. A ratio quoted without its measurement conditions commits the vendor to nothing.

The obvious response is to discount the claim and buy extra capacity to cover the gap. That works until the budget is fixed, which it usually is. The team still has to decide how much raw capacity sits behind a given amount of protected data, and the ratio is the only thing joining those two figures. Assume too little reduction and scarce money goes into shelves nobody needed. Assume too much and the repository fills partway through a three year purchase. The useful move is not to argue about plausibility but to make the claim testable.

A ratio is an outcome, not a specification

Deduplication finds repeated content and stores one copy of it. How much it finds comes from the data, not the product. Two estates of identical size running identical software report different ratios because one holds many near identical virtual machines built from a common template and the other unrelated application servers. The engine did the same work. The raw material differed.

Retention matters as much as content. Most of the repetition in a backup set is not within one backup, it is across successive backups of the same machines over time. A repository holding a few weeks of restore points has fewer copies of the same operating system files to collapse than one holding a year of monthly and yearly points. Hence a result that surprises people: extending retention often improves the reported ratio while increasing the capacity consumed. The ratio moved in the flattering direction and the invoice moved in the other one.

What a quoted ratio has to disclose to mean anything

A claim becomes assessable once five things are attached to it. The first is the workload mix, described by what the machines do rather than how many there are, since virtual desktops, file servers, database servers and media repositories reduce very differently and the proportions decide the blended result.

The second is the retention period and the restore point structure behind it, since a ratio measured across a year of weekly, monthly and yearly points is not comparable to one measured across a fortnight of dailies. The third is the chain type, because forever forward incremental, periodic active fulls and synthetic fulls each present a different pattern of repeated blocks to whatever is deduplicating them.

The fourth is whether the backup software already reduced the data upstream. Most backup applications deduplicate within a job and compress before the stream leaves the proxy, and the repetition they removed is repetition the storage layer will never see. The fifth is whether compression is folded into the figure. A combined reduction number is legitimate when labeled as one and misleading when called a deduplication ratio, since the two mechanisms move independently. Knowing how compression behaves across different workload types explains why a combined figure shifts even when deduplication does not.

Why one estate produces different numbers at different layers

One set of backups will honestly report several different ratios at the same moment, because each layer measures a different denominator. None is lying. They answer different questions, and a quote rarely says which one it answered.

Where the number comes fromWhat it is actually comparingWhat has to be stated alongside it
Backup job report in the backup consoleSource data read against data written for that jobWhether compression is included, and whether the job is a full or an incremental
Repository or appliance dashboardLogical data presented to the repository against physical capacity usedWhether the backup software had already reduced the stream before it arrived
Storage layer beneath the repositoryBlocks written against blocks stored after its own reduction passBlock size, and whether the scope is per volume or global across the system
Front end terabytes on the quoteProtected data against purchased raw capacityWhether protection overhead, headroom and locked restore points are inside the figure

The last row causes most of the disputes after delivery. Front end terabytes describe what is protected. Raw capacity describes what was bought. Between them sit deduplication, compression, the data protection scheme, reclaim behavior and immutability, and a comparison leaving any of those implicit is not a comparison. The same problem appears whenever two quotes sit side by side, so comparing usable capacity across competing quotes has to begin by agreeing what each number counts.

Structuring a claim so it can be tested rather than argued

A ratio that arrives as marketing cannot be enforced. One that arrives with a defined input, a defined measurement point and a defined date can be. The conversion takes four sentences in the purchase paperwork.

State the input: the protected set the claim applies to, by workload class and front end terabytes, measured on the estate being quoted rather than a reference estate. State the retention: the restore point schedule, including long term points, and the chain type the jobs will use. State the measurement point: which console reports the figure, what it counts, and on what day of the retention cycle it is read, since a ratio taken the morning after an active full differs from one taken a week later.

Then state the consequence. If the measured figure falls short once a full retention cycle has completed, what happens next. Additional capacity at the original unit price is one answer. A revised sizing at no charge is another. Nothing at all is also an answer, and it is better to know that before signing. The aim is to turn a sales figure into a commitment that survives the handover, in the way that sizing a repository from measured change rate turns a growth assumption into something auditable.

Where ARTESCA fits

ARTESCA is object storage software used as a backup target, deployed on infrastructure the customer runs, in the range of roughly 50 TB to 8.5 PB. In most deployments it sits behind a backup application that has already deduplicated within its jobs and compressed the stream before sending it. The reduction credit belongs to the backup software, and sizing an object repository on top of a ratio that application produced counts the same saving twice.

Capacity planning for an object backup target is therefore better done on the post reduction stream: the bytes that arrive and are stored, measured over a full retention cycle, rather than front end terabytes multiplied by a datasheet ratio. Both the backup software and the repository report that figure.

Immutability changes the arithmetic separately. Restore points held under S3 Object Lock occupy capacity until the lock expires, whatever job retention says, so the stored footprint of an immutable repository reflects the longest lock applied rather than the shortest retention configured. That is a capacity commitment, and it belongs beside the reduction assumption.

What to record before the purchase closes

Three things belong in writing, and all are available without special tooling. The first is a measured baseline from the current environment: front end terabytes by workload class, the stored footprint the repository reports, and the date and point in the retention cycle at which both were read. That baseline is the only honest input to a model, and it costs an afternoon.

The second is the assumption sheet behind the quote, in the vendor's own words: workload mix, retention, chain type, measurement layer, and whether compression is included. If those cannot be supplied, that is itself a finding.

The third is a recheck date, set one full retention cycle after go live and after the first long term restore points exist, since a ratio read before them describes a repository that does not yet exist. Read the same figure from the same console, compare it against the assumption sheet, and record the variance whether or not it is favorable. Any change to chain format, retention depth or workload mix resets that measurement, so the recheck belongs on the calendar.