Storage economics

Backup compression: Why savings vary by workload

Which workloads compress, which do not, where compression belongs in the path, and how to estimate a mixed estate.

7 min read
Dark data center racks with streams of glowing blocks, some shrinking sharply and others passing through unchanged

Compression is one of the few things in a backup platform that behaves predictably, as long as the data type is known. The trouble is that a backup set is never one data type. It is a mixture of virtual machines, databases, file shares, mailboxes and whatever media somebody put on a share three years ago, and one reduction factor applied to that mixture produces a capacity plan that is confidently wrong.

The usual shortcut is to take the factor a vendor or sizing tool offers, multiply the protected estate by it, and move on. That is defensible when the estate resembles the one the factor came from. It stops being defensible the moment a large share of images, a busy database, or a set of encrypted volumes enters the protected set, since those behave nothing like each other.

What compression is actually removing

Compression looks for redundancy inside a block of data and encodes it more efficiently: repeated byte sequences, long runs of the same value, predictable structure. Text has a great deal of that. A log file repeats timestamps, hostnames and message formats endlessly. A database page repeats field layouts and padding. An empty region of a virtual disk is a long run of zeros, the easiest case there is.

What compression cannot do is remove redundancy that has already been removed. A JPEG, a video file, a ZIP archive and a modern installer package have all been through this once. Little structure is left for a second pass to find, so running one costs CPU time and returns close to nothing. Encrypted data is the extreme case: good encryption produces output statistically indistinguishable from random, and random data does not compress. That is not a limitation of an implementation, it is what the mathematics allows.

This is where compression and deduplication part company. Deduplication looks across blocks and across time for repeated content, while compression works within a block on structure. The two can move in opposite directions on the same data, which is why a combined data reduction figure in a quote hides more than it explains.

Which workloads give back space and which do not

Sorting the protected estate by data type rather than by machine is the step most sizing exercises skip. A file server is not a category. One holding scanned documents and one holding design renders and video behave very differently, and the machine name reveals neither.

Data in the protected setWhy it behaves this wayWhat to check before assuming a factor
Images, audio, video, archives, installer packagesAlready compressed at creation, so little structure remains for a second passHow much of the protected capacity these file types account for, by extension
Encrypted volumes and encrypted application dataCiphertext is statistically close to random and does not compressWhich machines use volume encryption, and whether the backup runs above or below it
Database files and transaction logsRegular page structure, padding and repeated field layoutsWhether the database engine is already compressing pages or backups itself
Unallocated and zeroed regions inside virtual disksLong runs of identical values, handled cheaplyWhether the backup reads free space at all, and whether disks are thin or thick
Text, logs, mail and office documentsHigh internal repetition in the text, though modern office formats are zippedThe split between plain text and already compressed office and mail formats

The fourth row makes platforms look better than they are. A freshly provisioned virtual machine with a large thin disk is mostly empty space, and empty space compresses spectacularly. The reported factor for that machine says almost nothing about how it will behave once the application has run for a year. Sizing from a lab build therefore flatters every number in the model.

Where in the path compression happens

Compression can occur in several places on the way from a production disk to a backup repository, and in a typical deployment more than one is switched on. The application may compress its own backups. The backup software compresses in the proxy before the stream leaves. A deduplicating appliance compresses after its own reduction pass. The storage beneath may compress again, and an inline compressing filesystem will attempt it on data that arrived already compressed.

Doing it twice is not harmful, it is wasted work. The second pass reads the output of the first, finds almost no remaining redundancy, and either returns a negligible saving or stores the block as it is. The cost is real: CPU cycles on whichever component did the work, and on a backup proxy those cycles compete with the job throughput that determines whether the backup window closes on time.

The practical rule is to compress once, as early in the path as the architecture allows, and to know where that happens. Early compression also shrinks what crosses the network, which matters more than the storage saving when the repository sits at another site. It does not follow that every layer offering the feature should have it enabled.

Estimating a mixed estate instead of applying one factor

A workable estimate treats the estate as a set of buckets rather than a single number. Group the protected data by behavior: already compressed content, encrypted content, database content, general file and operating system content, and empty space inside virtual disks. Capacity per bucket comes from file server reports, the hypervisor and the database inventory that already exist.

Then apply a separate expectation to each bucket and add the results, rather than one blended factor to the total. The arithmetic is trivial, and the value is visibility rather than precision: it shows which bucket dominates. In most mid sized estates one or two buckets carry the entire result, and a model that makes that obvious survives a change in the estate far better than a blended figure does.

Databases deserve their own line because the answer depends on configuration rather than data. An engine that compresses its own backups hands the backup software something already reduced, and the credit belongs upstream. That is one reason database protection is planned separately from virtual machine protection.

Where ARTESCA fits

ARTESCA is object storage software used as a backup target, deployed on infrastructure the customer runs, at scales from roughly 50 TB to 8.5 PB. Data reaching it through a backup application has generally been compressed already, in the proxy, before it was sent. The bytes that arrive are the bytes that have to be stored, which makes the stored object footprint the right input for capacity planning.

That has a direct consequence for sizing. Measuring the volume of objects written over a full retention cycle answers the capacity question without any factor at all, since the reduction has already happened by the time the data is visible to the repository. Front end terabytes multiplied by an assumed factor is an estimate. The stored footprint is a measurement.

It also affects where compression settings belong in a deployment discussion. On an object repository the controls that matter sit in the backup software and in how the job writes, which is part of planning an object storage repository before the first job runs rather than something tuned at the storage layer afterward.

What to measure and when

Two measurements taken quarterly keep the model honest. The first is the bucket inventory: capacity by data type across the estate, refreshed from the same reports each time so the numbers stay comparable. Large shifts arrive quietly, through a new application, a video archive moved onto a protected share, or encryption enabled fleet wide.

The second is the stored footprint against the front end terabytes for the same period, read from the backup software and the repository on the same day. The relationship between them is the blended factor the estate is actually delivering, and comparing it against the model shows drift before it becomes a capacity problem.

Two events should trigger a recheck outside that schedule. One is any change to encryption, since it can remove most of the compressible material in a workload overnight. The other is onboarding a new application, where the sensible move is to record its behavior separately for its first full retention cycle before folding it into the blended figure, writing down what was measured rather than what was expected.

Try ARTESCA free

Immutable object storage that scales from 20TB to petabytes. Deploy a working cluster in under an hour.

Start a free test drive