ARTESCA Blog | Backup, recovery and cyber resilience

Backup capacity alerts: When should you add storage?

Written by Joshua Silvia | Sep 20, 2026, 3:37:51 AM

Most backup repositories are monitored with a single threshold, an alert at eighty percent full and a louder one at ninety. That threshold fires late, fires often, and says nothing useful about when storage has to be ordered. A repository can sit at seventy percent for months and then cross into trouble in nine days, because backup consumption is not a steady line and space that has expired is not the same thing as space that has been returned.

Moving the threshold fails the same way whichever direction it moves. A level says where the repository is, not how fast it is moving or what the next week of jobs is about to write. Procurement runs in weeks, and a level based alert warns in identical language for a repository with six months of runway and one with eleven days.

The second problem is reclaim. Deleting a restore point in the console does not return capacity on the storage holding it. Retention expiry, chain merges, garbage collection and any remaining immutability period sit between the delete and the free space, and the repository keeps filling while the console reports that space should already be available.

Why the fill line moves in steps

Backup consumption is periodic rather than linear. Daily incrementals write a small and fairly predictable amount. Then a synthetic or active full lands and writes something much closer to the size of the protected set, and a monthly restore point lands on top of that and is held for a year. Used capacity against time is a staircase, not a slope, and the risky moment is always the step.

The flat part is where a level threshold gets evaluated, which is why it misreads the situation. A repository at seventy eight percent on Thursday looks comfortable. If the monthly full runs on Saturday and the previous monthly cannot be released until its immutability period ends, that repository is past its threshold on Sunday with no action possible in between.

The number worth knowing is therefore the size of the next scheduled full rather than the current percentage. It is derivable from the last full of the same type, adjusted for anything added to the job since.

Days of headroom instead of space used

A trajectory alert answers a different question. At the current rate of consumption, including the next scheduled full, how many days remain before the repository can no longer complete a job? That number is comparable to a lead time. If replacement capacity takes six weeks to quote, ship and rack, an alert at forty five days of headroom is actionable and one at ninety percent full is late.

Calculating it needs no special tooling. Average daily growth should be taken over a complete retention cycle rather than the last week, since a week may contain no monthly full. The expected size of the next full is added, and capacity certain to be released is subtracted, where certain means locks already expired rather than retention merely scheduled. Free space divided by that rate gives headroom in days.

This also turns a one time sizing exercise into an operational control. Work on growth and churn across a purchase cycle sets the buffer at purchase. The headroom number reports each month whether reality has stayed inside it.

Signals that arrive before the capacity problem

A few measurable changes reliably show up ahead of a free space crisis, and each has a different remedy. Separating them matters because most are configuration effects correctable in days, while only one requires buying disks.

SignalWhat it usually meansWhat to check
Daily growth rises while the protected data set is flatChange rate increased, often from a new application, a database reindex, or disk encryption on endpointsWritten size per job over two retention cycles, and recent changes inside the guest workloads
The reduction ratio fallsNew data compresses or deduplicates worse than the existing set, typically images, video or encrypted volumesRatio reported per job rather than per repository, and what was added to each job recently
Free space does not rise after restore points expireReclaim is pending, through chain merge, garbage collection, or an immutability period still runningWhether deleted restore points are still counted on the storage side, and when the oldest lock expires
A retention setting changed recentlyThe full effect has not landed, since extra copies accumulate over monthsJob retention history, and how many weekly, monthly and yearly points are now scheduled to be kept
One job's session size jumps once and stays highThe job now includes a machine or volume it did not include beforeJob membership, and whether an exclusion was removed or a template changed

Three of those five are visible only in per job data. A repository level dashboard averages them away, which is why a capacity review that looks only at the total runs a month behind the cause.

Reduction effectiveness as a leading indicator

Every capacity forecast assumes a reduction ratio, and that ratio is an average over the data currently stored rather than a property of the storage. When the mix changes the ratio changes, and a forecast built on last quarter's average understates next quarter's consumption. A file server full of photographs, or database volumes with transparent encryption, moves the effective ratio with no change to the backup configuration.

Tracking the ratio per job makes that legible. A repository wide number is dominated by the largest jobs and hides the one behaving differently. Vendor claims about deduplication ratios are averages too, which is why the measured number from the environment at hand is the only one worth forecasting from.

Retention changes deserve the same treatment. Keeping twelve monthly restore points instead of six does not double capacity on the day the setting changes. It adds one restore point a month for six months, precisely the pattern a level based alert notices last. Tracking the restore points a retention policy actually produces makes the increase visible at the change rather than at the ceiling.

Where ARTESCA fits

ARTESCA is object storage software used as a backup target, deployed on infrastructure the customer runs, in capacities from roughly 50 TB upward into the petabyte range. As an S3 target it is where the backup software's retention and immutability decisions turn into consumed capacity, so the console view and the storage view of the same repository have to be read together during a capacity review.

Because it supports S3 Object Lock, restore points written under a lock occupy space until that lock expires, whatever the job's retention setting says. A capacity trajectory therefore has to account for the longest lock currently applied rather than the shortest retention configured.

Capacity is added by growing the deployment rather than replacing it, which changes what the lead time is made of. The relevant duration becomes how long hardware takes to arrive and join the cluster, not how long a platform migration takes, and that figure is usually obtainable from the reseller.

What to review on a schedule

A monthly record of four numbers per repository is enough to build a trajectory: total used on the storage side, average daily growth across the last complete retention cycle, the written size of the most recent full of each type, and the reduction ratio per job. Four months of that history makes the next step predictable.

The alert itself belongs on days of headroom, with the threshold derived from the procurement lead time the organization actually experiences rather than a round number. Quote to delivery time comes from the reseller, internal approval time has to be added, and a margin of one retention cycle sits on top. A second, lower alert separates the moment to open a conversation from the moment the plan has failed.

Alongside the alert belongs a written note of what happens when it fires: who is contacted, which capacity increment is the standard next step, and which temporary measures are permitted. A capacity alert with no documented response produces the same outcome as no alert, one week later. Comparing the monthly number against the headroom the design was meant to carry closes the loop between the purchase decision and operational reality.