Headroom on a backup repository is usually a round proportion, carried over from the last platform or from whatever the previous administrator wrote down. It survives because nothing tests it until the week it fails. The figure has a specific job, though. It has to absorb the next full, a retention change nobody costed, a restore staged back onto the same platform, the capacity a rebuild consumes, and the weeks between deciding to buy and racking the result.
A proportion is not wrong so much as untethered. The same fraction on two repositories represents two different amounts of absorbable work, because what has to be absorbed is measured in terabytes of particular events rather than in proportion to what is already stored. A large repository with flat growth and a smaller one about to take on a new application need different buffers.
Headroom is also spent by things that are not backup jobs. Restores, maintenance, a failed node and the lag between expiry and reclaim all draw on the same free space, and none of them appear in a sizing calculation built from change rate and retention.
Why a proportion is the wrong starting point
Sizing produces a requirement: this much capacity holds this many restore points for this long. Headroom is a separate question, about absorption rather than storage. The right unit is the set of demands the repository has to survive without an emergency, expressed in terabytes rather than as a share of a number that was itself an estimate.
Proportions also drift. As a platform grows, a fixed share grows with it, which sounds prudent but usually overshoots, since the events that consume headroom do not scale with total capacity. A synthetic full scales with the protected set. A rebuild scales with one failure domain. Neither tracks how much history is kept.
The demands already on the calendar
Two of the largest claims on headroom are known in advance and can be read out of the existing configuration. The first is the next full. Whether it is active or synthetic, it writes a volume close to the size of the protected set, and it coexists with the chain it is replacing until merge or expiry releases the older data. On object storage the mechanics differ, since new blocks are written and old ones become eligible for deletion rather than being overwritten in place, and how a backup chain behaves on an object repository determines how long the overlap lasts.
The second is any retention change agreed but not fully landed. Extending monthly retention consumes nothing on the day the setting changes. It adds restore points month by month until the new depth is reached, so a change made last quarter may still have most of its consumption ahead of it, and that outstanding portion is committed spending absent from the used figure.
Immutability lengthens both effects. A restore point under a lock occupies space until the lock expires whatever the job's retention says, so the overlap is governed by the longest lock applied rather than the shortest retention configured.
The demands that are not backup jobs
The remaining claims are the ones a sizing model leaves out. Each has a different size and a different way of being estimated, and separating them turns headroom from a guess into an arithmetic result.
| Demand on headroom | Why it is missed | How to size it |
|---|---|---|
| The next active or synthetic full | Sizing counts steady state, not the moment two chains coexist | Written size of the last full of the same type, plus anything added to the job since |
| A retention change already approved | Its consumption arrives over months, so the used figure understates it | Restore points still to be added, multiplied by the average size of one |
| A restore staged back onto the same platform | Restores are treated as reads, not writes | The largest recovery the runbook contemplates, plus any temporary copy the process creates |
| Rebuild after a node or drive failure | Protection overhead is assumed static | Capacity needed to restore the data protection scheme after losing one failure domain |
| Reclaim that has not completed | The console reports space as released before the storage returns it | Observed lag between expiry and free space rising, expressed in days of growth |
| Procurement and installation lead time | Capacity is treated as available on the day it is approved | Growth per day multiplied by quote, approval, delivery and rack time |
The restore row is the one most often disputed and most often real. Recovering a large file server or a set of virtual machines usually means writing data somewhere before it is usable, and if the only platform with room is the backup platform, that space has to exist in advance.
The rebuild row is decided by the data protection scheme and the failure domain, not by the backup software. Losing a node means re-establishing the configured level of protection somewhere, and without spare capacity the system either cannot, or does so at the cost of space the next job needs.
What happens when a repository runs genuinely full
Running out of space does not produce a clean stop. Jobs fail partway through, leaving incomplete restore points that still occupy capacity until they are cleaned up, and the cleanup itself sometimes needs room to work. A merge that cannot complete leaves a chain in a state neither old nor new.
The recovery position degrades before the alerts get loud. Missed jobs mean the newest usable restore point ages, and a retention policy that cannot write new points stops aging out old ones in the normal sequence, so the retained set no longer matches what was designed.
Immutability narrows the emergency exits. The usual reflex, deleting older restore points to make room, is unavailable when those points are under a lock, and that is correct behavior rather than a fault. What remains is reducing what is protected, moving data elsewhere, or adding capacity, all slower than the situation allows. That asymmetry is the argument for counting lead time explicitly, and why alerting on days of runway rather than space used changes the decision.
Where ARTESCA fits
ARTESCA is object storage software used as a backup target, deployed on infrastructure the customer runs, in capacities from roughly 50 TB upward into the petabyte range. It is where retention and immutability settings made in the backup console turn into consumed capacity, so a headroom figure has to be validated against the storage view rather than the console view alone.
Because it supports S3 Object Lock, capacity held under a lock is not available for reclaim until the lock expires, whatever a job's retention setting reports. The overlap component of headroom therefore follows the longest lock currently in force, and the rebuild component follows the configured data protection scheme and the size of one failure domain.
Capacity is added by expanding the deployment rather than replacing it, which changes what the lead time component contains. The relevant duration becomes quote, delivery and the time for new hardware to join the cluster, a figure normally available from the reseller who supplied the platform.
Deriving the number and keeping it current
The derivation is a sum of the demands rather than a maximum, since several can coincide. A node failure during a monthly full week is not unusual. A workable model adds the largest full, the outstanding portion of any approved retention change, one rebuild and the lead time allowance, then treats the restore allowance as either included or explicitly excluded with a written reason. Sizing work on how much repository capacity the environment requires supplies the baseline this sits on.
It needs revisiting when the inputs move, which in practice means quarterly and after any change to job membership, retention, immutability duration or the protection scheme. What belongs in writing is the figure in terabytes, the date it was derived and the inputs that produced it, so the next person can see which assumption broke.
The figure belongs in the maintenance plan too, since several operations need free space to run. Work such as storage side maintenance on a live backup platform competes for the same capacity, and headroom already spent by growth leaves no room to perform it.
