Cyber resilience

Backup retention locks: What happens when they expire?

What object storage does when a retention lock lapses, why backup and bucket retention disagree, and how capacity returns.

7 min read
Rows of glowing storage blocks in a dark data center, one row dimming as its seal releases

A retention lock is a date. When that date passes, the protection that made the copy immutable ends, quietly and without ceremony. Most teams configure retention once, watch the first objects land under lock, and never see the other end of the cycle. Expiry is where assumptions about capacity, cleanup and restore dependencies are all tested at the same time, usually with nobody watching.

The common expectation is that expiry is a kind of scheduled deletion, and that storage frees itself as retention lapses. That is not what the mechanism does. Expiry removes a restriction. It does not remove data, and it notifies nobody.

What follows is what happens at that boundary, why two systems hold different opinions about how long a backup should be kept, and what the capacity graph does after a large batch of objects comes out of retention at once.

What the storage does when the date passes

Each protected object version carries a retain until date. Until that timestamp a delete request against that version fails. After it, the same request succeeds. Nothing else changes. The object is not moved, not marked, not reported, and no event is raised by the lock itself. An object whose retention expired a year ago looks exactly like one that expired this morning.

Deletion happens only because something asks for it, normally the backup software's retention job, a lifecycle rule on the bucket, or a person running a cleanup. If none of those exist, objects sit at full size indefinitely with their protection gone, which is the worst combination of cost and risk.

Versioning adds a step that surprises people. Deleting an object in a versioned bucket normally writes a delete marker and leaves the previous version in place, still consuming capacity. Space returns only when noncurrent versions are themselves expired, usually by a lifecycle rule written for that purpose. A repository can look cleaned up in the backup console and unchanged on the capacity graph.

Why the software and the bucket disagree about retention

Backup software calculates retention against restore points and the dependencies between them. A weekly full with daily incrementals, or a forever forward chain with synthetic fulls, often forces the software to keep data longer than the nominal policy so that a restore point stays usable. Immutability settings usually add the length of the current block generation to the configured number of days, so the lock applied to an object runs longer than the figure printed in the job settings.

Object storage has no view of any of that. It sees individual objects with individual dates. When the backup software sets retention per object at upload, the two systems stay roughly aligned. When a bucket default retention is also configured, every object carries that period too, and the longer value governs. A default intended as a safety net becomes the binding constraint for the whole repository.

The disagreement runs in two directions. Software retention shorter than the lock means cleanup requests fail and space is not returned on schedule. Software retention longer than the lock means objects are unprotected while the software still depends on them. Neither is visible from one console, so the retention figure in the job and the retain until date on a sample of objects both need reading. Whether such a mistake can be corrected at all depends on the retention mode the bucket was created with.

An expired lock on an object a job still needs

Expiry has no effect on readability. A restore from an object whose lock has lapsed works exactly as it did the day before, since the lock governed deletion and never governed access. The risk is not that the data becomes unusable at expiry, it is that the data becomes removable while something still depends on it.

SignalWhat it usually meansWhat to check
Capacity flat for weeks after a large expiry dateNothing is issuing deletes, or only delete markers were writtenWhether noncurrent version expiration exists, and whether the retention job ran
Repository grows faster than the policy predictsObjects are locked longer than the software expectsRetain until dates on sample objects against the bucket default retention
Delete errors in the backup job logSoftware retention is shorter than a lock still in forceWhether the software retries later or abandons the object permanently
Restore of an older point fails after a cleanupA dependent object was removed by something the software does not trackLifecycle rules on the bucket and every identity holding delete rights
Objects no software claims to ownJobs or repositories were removed while objects were lockedA bucket inventory listing compared against the backup catalog

The dangerous case is the fourth row. A lifecycle rule written months earlier to control growth does not know which objects belong to a chain. Once the lock expires that rule is free to act, and the deletion it performs is silent and successful. The failure surfaces during a restore of an older point, which is a good argument for testing restores from more than the newest point.

What capacity does after a large expiry wave

Large expiry waves are common because repositories are filled in bursts. A seeding run, a migration, or the first full backups after a new job was created all write a large batch of objects within a few days, and those objects leave retention within a few days of each other as well.

Capacity then falls in steps rather than smoothly. The backup retention job runs on its own schedule and marks objects for deletion. Lifecycle evaluation on the bucket runs on a schedule of its own, often daily, and expires noncurrent versions. Space accounting can lag both. Reclaim takes days rather than minutes, which is worth confirming once rather than discovering during a capacity alert.

The more damaging pattern is the wave that frees nothing. If the backup software no longer references the objects, because a job was deleted or a repository was removed while the data was still locked, no cleanup will ever be issued for them. They stay at full size forever. Finding them requires comparing a listing of the bucket against the backup catalog, which is not a routine task and is the reason orphaned data usually appears as an unexplained gap in the sizing model for the repository.

Where ARTESCA fits

ARTESCA is object storage software used as a backup target, with immutability provided through S3 Object Lock. Retention behaves as the S3 API specifies, so an object version becomes deletable once its retain until date has passed and remains in place until a delete request arrives, whether from the backup software or from a lifecycle rule configured on the bucket.

Because versioning underpins Object Lock, capacity planning has to account for noncurrent versions as well as current ones. Reclaiming space after expiry depends on lifecycle configuration that expires noncurrent versions, not on the expiry of the lock by itself. The same applies to the guarantees the lock provides while it is in force, which end precisely when the date does.

Since the system runs on infrastructure the customer operates, reclaimed capacity returns to the customer's own hardware, and the timing is something the team can observe directly rather than infer from a bill.

What to watch on a schedule

Once, deliberately, watch a single object through expiry. Note its retain until date, attempt a delete before that date and confirm the failure, then attempt it afterward and confirm the success. That takes minutes and replaces assumptions with observed behavior, including which log entries appear and which do not.

Monthly, compare the number of objects in the repository against the number of restore points the backup software believes it holds. A widening gap is the earliest visible sign of objects that expired and were never cleaned up, or of objects the software has forgotten it wrote.

Keep a short written record of every lifecycle rule on every backup bucket, what it matches and when it was created, and review it whenever retention changes in the backup software. Put the date of the next large expiry wave on the operations calendar with the capacity that should return, so a wave which frees nothing is noticed within weeks rather than at the next purchase.

Try ARTESCA free

Immutable object storage that scales from 20TB to petabytes. Deploy a working cluster in under an hour.

Start a free test drive