Backup & recovery

Backup retention: Match recovery points to real needs

Retention worked backward from the recoveries an organization actually performs, and the GFS tiers that reconcile them.

7 min read
Receding rows of glowing data blocks in a dark hall, dense in front and thinning into the distance

Retention is one of the few backup settings that is almost never decided. It arrives with the first job, copied from a previous product or left at a vendor default, and then it stays. Thirty days, ninety days, seven years. The number gets defended long after anyone remembers what it was meant to cover, and it is usually the wrong shape for every recovery the organization will actually perform.

The reason is that a single retention figure answers a single question: how far back the copies reach. Real recoveries ask a second question at the same time, which is how precisely a restore can land once it gets there. Those two are independent, and one number cannot express both.

A more useful approach is to work backward from the recoveries an organization actually performs. There are roughly four of them, they are discovered on very different timescales, and each one needs a different granularity of restore point.

Four recoveries, four different shapes

The first is the accidental change. Someone deletes a folder, overwrites a file, or runs an update against the wrong database. It is noticed within hours, usually by the person who did it. Granularity matters here, not depth. Last night's restore point may already be too old if six hours of work sits on top of it.

The second is corruption that nothing reports. A storage fault, a failed upgrade, or a bad batch job damages data quietly, and it surfaces days or weeks later when someone opens a record that no longer makes sense. Depth matters here, and so does the ability to test several candidate points, since nobody knows when the damage began.

The third is an intrusion found after dwell time. Access began weeks before anyone noticed, so the most recent points are suspect and the question becomes which point predates the compromise. This is the scenario where retention and immutability have to hold together, because the attacker had time to reach the backup system too, and where identifying a point that can be trusted is the real work.

The fourth is not an incident at all. A legal request or an audit asks for data as it stood on a particular date months or years ago. Granularity barely matters. What matters is that a point exists for that period, that it can be located by date, and that it is still readable.

Why a single number serves none of them well

Set retention at thirty days and the first two scenarios are covered while the third is uncertain and the fourth is impossible. Set it at seven years at the same daily granularity and the capacity bill is absurd, and the first scenario is no better served than before. The scenarios pull in different directions.

Recovery scenarioWhen it is usually discoveredWhat the restore point has to provide
Accidental deletion or overwriteWithin hoursFine granularity across the last few days, closer than one point per night for some workloads
Silent corruptionDays to weeksDaily points deep enough to predate the fault, and several candidates to test
Intrusion found after dwell timeWeeks to monthsWeekly and monthly points, held immutable, with a known last known good date
Departure or deliberate deletionWeeksPoints nobody with production access can alter or remove early
Legal, audit or regulatory requestMonths to yearsA yearly or monthly point locatable by date, still readable on current software

The pattern is in the middle column. The later a scenario is discovered, the less granularity it needs and the more depth and protection it needs. A flat policy applies one granularity everywhere, so it either overspends at the old end or underserves the recent one.

Grandfather father son as the reconciling mechanism

The GFS scheme exists precisely for this. Rather than one retention figure it keeps several tiers with different densities: every daily point for a short window, one point per week for a longer one, one per month beyond that, and one per year at the far end. Each tier is tuned to the scenario it serves.

The mechanics are worth understanding, since GFS points are not copies made separately. In most products a weekly, monthly or yearly point is an ordinary restore point flagged for long term keeping and therefore exempted from the normal retention count. That flag is applied when the point is created, which means the schedule decides forever afterward which points survive. A monthly point flagged on the first Saturday cannot be moved to the last day of the month in hindsight.

That has a practical consequence for audit and legal work. If the requirement is data as it stood at a month end, a monthly point anchored to a weekend that falls mid month will not satisfy it, and nobody discovers this until the request arrives. The anchor day of each tier deserves a deliberate decision rather than a default.

What each tier costs and where the cost lands

Every tier added is capacity that never ages out during the window it covers. Daily points in a deduplicated or incremental chain are comparatively cheap, since consecutive days share most of their blocks. Yearly points are the opposite, since a point kept for years shares very little with current data and eventually holds a complete copy of a machine nobody remembers.

Growth from long tiers is slow and therefore invisible until it is not. A yearly tier retaining seven points reaches full size seven years after it was configured, long after the person who configured it moved on. Modeling what one more year of retention actually adds before adding it is far easier than reversing the decision later.

Immutability compounds this. When long term points are locked, the lock period is usually derived from the retention setting rather than set independently, and the effective lock can run longer than the figure typed into the job. The interaction between job retention, GFS flags and lock duration is product specific and worth confirming against the settings that actually produce the lock, because capacity cannot be reclaimed early once objects are under it.

Where ARTESCA fits

ARTESCA is object storage software used as a backup target, with an S3 compatible API and immutability through S3 Object Lock. Retention tiers are defined in the backup software, and the objects that make up each retained point arrive at the repository carrying a retain until date derived from those settings.

What the repository enforces is the lock, not the policy. It cannot tell a daily point from a yearly one, and it will refuse a delete request for any object whose date has not passed regardless of which tier the backup product believes it belongs to. That is the intended behavior, and it is also why a retention change made in the software does not shorten protection already applied. What happens at the far end of that period is covered in what a repository does when a lock expires.

Because ARTESCA runs on infrastructure the customer operates, the retain until dates, the bucket defaults and the identities allowed to change them are under the customer's own administration, and belong in the same review as the retention policy.

What to write down and when to revisit it

Retention deserves a one page record, written once and reviewed annually. It should name each tier, the anchor day or time that tier is flagged on, the number of points kept, the scenario the tier exists to serve, and who asked for it. The last field is the one that saves the most argument later, since the yearly tier usually exists because of a requirement nobody can now cite.

Two checks keep the record honest. Once a quarter, list the actual restore points for one protected workload and confirm the tiers present match the tiers described, since a job edited during a capacity squeeze will not announce that it stopped flagging monthlies. Once a year, restore from the oldest retained point rather than assuming it works, because reading a point from four years ago depends on software upgraded several times since.

Retention also changes when the business does. A new application, regulator or contract moves the requirement, and the policy belongs under review at that moment rather than at the next renewal.

Try ARTESCA free

Immutable object storage that scales from 20TB to petabytes. Deploy a working cluster in under an hour.

Start a free test drive