Compression is one of the few things in a backup platform that behaves predictably, as long as the data type is known. The trouble is that a backup set is never one data type. It is a mixture of virtual machines, databases, file shares, mailboxes and whatever media somebody put on a share three years ago, and one reduction factor applied to that mixture produces a capacity plan that is confidently wrong.
The usual shortcut is to take the factor a vendor or sizing tool offers, multiply the protected estate by it, and move on. That is defensible when the estate resembles the one the factor came from. It stops being defensible the moment a large share of images, a busy database, or a set of encrypted volumes enters the protected set, since those behave nothing like each other.
What compression is actually removing
Compression looks for redundancy inside a block of data and encodes it more efficiently: repeated byte sequences, long runs of the same value, predictable structure. Text has a great deal of that. A log file repeats timestamps, hostnames and message formats endlessly. A database page repeats field layouts and padding. An empty region of a virtual disk is a long run of zeros, the easiest case there is.
What compression cannot do is remove redundancy that has already been removed. A JPEG, a video file, a ZIP archive and a modern installer package have all been through this once. Little structure is left for a second pass to find, so running one costs CPU time and returns close to nothing. Encrypted data is the extreme case: good encryption produces output statistically indistinguishable from random, and random data does not compress. That is not a limitation of an implementation, it is what the mathematics allows.
This is where compression and deduplication part company. Deduplication looks across blocks and across time for repeated content, while compression works within a block on structure. The two can move in opposite directions on the same data, which is why a combined data reduction figure in a quote hides more than it explains.
Which workloads give back space and which do not
Sorting the protected estate by data type rather than by machine is the step most sizing exercises skip. A file server is not a category. One holding scanned documents and one holding design renders and video behave very differently, and the machine name reveals neither.
| Data in the protected set | Why it behaves this way | What to check before assuming a factor |
|---|---|---|
| Images, audio, video, archives, installer packages | Already compressed at creation, so little structure remains for a second pass | How much of the protected capacity these file types account for, by extension |
| Encrypted volumes and encrypted application data | Ciphertext is statistically close to random and does not compress | Which machines use volume encryption, and whether the backup runs above or below it |
| Database files and transaction logs | Regular page structure, padding and repeated field layouts | Whether the database engine is already compressing pages or backups itself |
| Unallocated and zeroed regions inside virtual disks | Long runs of identical values, handled cheaply | Whether the backup reads free space at all, and whether disks are thin or thick |
| Text, logs, mail and office documents | High internal repetition in the text, though modern office formats are zipped | The split between plain text and already compressed office and mail formats |
The fourth row makes platforms look better than they are. A freshly provisioned virtual machine with a large thin disk is mostly empty space, and empty space compresses spectacularly. The reported factor for that machine says almost nothing about how it will behave once the application has run for a year. Sizing from a lab build therefore flatters every number in the model.
Where in the path compression happens
Compression can occur in several places on the way from a production disk to a backup repository, and in a typical deployment more than one is switched on. The application may compress its own backups. The backup software compresses in the proxy before the stream leaves. A deduplicating appliance compresses after its own reduction pass. The storage beneath may compress again, and an inline compressing filesystem will attempt it on data that arrived already compressed.
Doing it twice is not harmful, it is wasted work. The second pass reads the output of the first, finds almost no remaining redundancy, and either returns a negligible saving or stores the block as it is. The cost is real: CPU cycles on whichever component did the work, and on a backup proxy those cycles compete with the job throughput that determines whether the backup window closes on time.
The practical rule is to compress once, as early in the path as the architecture allows, and to know where that happens. Early compression also shrinks what crosses the network, which matters more than the storage saving when the repository sits at another site. It does not follow that every layer offering the feature should have it enabled.
Estimating a mixed estate instead of applying one factor
A workable estimate treats the estate as a set of buckets rather than a single number. Group the protected data by behavior: already compressed content, encrypted content, database content, general file and operating system content, and empty space inside virtual disks. Capacity per bucket comes from file server reports, the hypervisor and the database inventory that already exist.
Then apply a separate expectation to each bucket and add the results, rather than one blended factor to the total. The arithmetic is trivial, and the value is visibility rather than precision: it shows which bucket dominates. In most mid sized estates one or two buckets carry the entire result, and a model that makes that obvious survives a change in the estate far better than a blended figure does.
Databases deserve their own line because the answer depends on configuration rather than data. An engine that compresses its own backups hands the backup software something already reduced, and the credit belongs upstream. That is one reason database protection is planned separately from virtual machine protection.
Where ARTESCA fits
ARTESCA is object storage software used as a backup target, deployed on infrastructure the customer runs, at scales from roughly 50 TB to 8.5 PB. Data reaching it through a backup application has generally been compressed already, in the proxy, before it was sent. The bytes that arrive are the bytes that have to be stored, which makes the stored object footprint the right input for capacity planning.
That has a direct consequence for sizing. Measuring the volume of objects written over a full retention cycle answers the capacity question without any factor at all, since the reduction has already happened by the time the data is visible to the repository. Front end terabytes multiplied by an assumed factor is an estimate. The stored footprint is a measurement.
It also affects where compression settings belong in a deployment discussion. On an object repository the controls that matter sit in the backup software and in how the job writes, which is part of planning an object storage repository before the first job runs rather than something tuned at the storage layer afterward.
What to measure and when
Two measurements taken quarterly keep the model honest. The first is the bucket inventory: capacity by data type across the estate, refreshed from the same reports each time so the numbers stay comparable. Large shifts arrive quietly, through a new application, a video archive moved onto a protected share, or encryption enabled fleet wide.
The second is the stored footprint against the front end terabytes for the same period, read from the backup software and the repository on the same day. The relationship between them is the blended factor the estate is actually delivering, and comparing it against the model shows drift before it becomes a capacity problem.
Two events should trigger a recheck outside that schedule. One is any change to encryption, since it can remove most of the compressible material in a workload overnight. The other is onboarding a new application, where the sensible move is to record its behavior separately for its first full retention cycle before folding it into the blended figure, writing down what was measured rather than what was expected.
