Backup & recovery

Backup migration: How to preserve older restore points

Copy, run in parallel or restore and reprotect: how chains and retention locks decide what happens to backup history.

7 min read
Two racks of glowing storage blocks in a dark hall linked by light paths carrying data between them

Moving to a new backup repository is mostly a decision about history. The new target is straightforward to configure and the first full backup lands without drama. The question that stalls the project is what happens to the two or three years of restore points already sitting on the old system, some of them under a retention lock that will not lift for months. Three options exist, none of them free, and the date on the project plan usually assumes the cheapest one.

The options are to copy the history to the new repository, to keep the old system running in parallel until its retention expires, or to restore what matters and protect it again from the new target. Each trades money, time and risk differently, and the constraints that decide between them are technical rather than commercial.

What each of the three options costs

Copying everything is the simplest idea and the heaviest operation. Every object in the history is read from the old system and written to the new one, over a network also carrying the nightly jobs. The duration is set by the slower of source read and target write, and the work has to be paced so it does not push the backup window past morning.

Running both in parallel avoids the copy entirely. The new repository takes new backups, the old one serves restores until its last restore point expires, and nothing moves. The cost is two systems supported, powered and licensed for as long as the longest retention, and restores from before the cutover following a different procedure from those after it.

Restore and reprotect is the narrowest option. A small number of important systems are restored from the old repository and backed up afresh to the new one. It does not preserve history, since the new restore points reflect the state at migration rather than their original dates, which matters where old backups are held for a legal or regulatory reason.

Why chain dependencies constrain a copy

Restore points on object storage are not independent files. An incremental point depends on the blocks written before it, and a synthetic full is assembled from existing blocks rather than written whole. Copying a subset of objects therefore does not produce a subset of restore points. It produces restore points that appear in a list and fail when opened, because part of the chain was left behind.

This is why a migration belongs to the backup software rather than a storage level copy tool, wherever the software supports it. The software knows which objects belong to which chain and updates its catalog as it goes. A bucket to bucket copy moves bytes correctly and leaves the catalog pointing at the old location, which is an unpleasant thing to discover during an incident.

The way chains are laid out varies between products and versions, so the arrangement in the deployment at hand is worth confirming before planning around it. How backup chains are structured on object storage determines how much of the history has to move together, and whether the oldest points can be dropped independently or only with everything that depends on them.

Immutability and the date the old system can be switched off

Retention locks are the reason decommissioning dates slip. An object under compliance mode retention cannot be deleted before its retain until date by any account, including the one that wrote it, and copying that object to a new repository does not shorten the original. The old system holds protected data until the last date passes, whatever the project plan says.

Governance mode differs in that a sufficiently privileged account can lift the retention, but that is an exception with an audit trail attached, not a routine migration step. Planning that assumes it will be used is planning around a control that exists to resist exactly that.

ConstraintWhy it moves the cutover dateWhat to check before committing to a date
Chain dependencyA restore point is unusable without the blocks behind itWhere each chain starts and how much must move as one unit
Compliance mode retentionNothing and nobody can remove the object earlyThe latest retain until date across all objects, not the policy value
Long term restore pointsYearly points extend the tail long past the daily onesThe oldest yearly point held and the date it expires
Read load at the sourceA full copy reads every object once, alongside nightly jobsSustained source read rate and how many nights the copy needs
Catalog and metadataCopied objects without catalog entries are unusable dataWhether the software supports a repository move that keeps the catalog

The second row is the one to measure rather than assume. Policies change over the life of a repository, so the longest remaining lock is often older and longer than the current setting suggests, and what happens when a retention lock expires is worth understanding before the date is put in a plan.

Choosing deliberately rather than by default

Most estates end up with a mixture, and that is a reasonable outcome when it is chosen rather than arrived at. Recent history, where restores actually come from, is copied. Older history under lock stays where it is until it expires. A handful of systems with a long legal obligation are restored and reprotected so the obligation no longer depends on the old platform.

The decision is easier once the question becomes how often each part of the history is used. Restores overwhelmingly come from the last few weeks. Everything older is insurance, and insurance can sit on a system in maintenance mode more cheaply than it can be moved. The counterargument is support life, since hardware out of support is a poor place to keep the only copy of anything.

The same reasoning applies when the migration is really a consolidation of several repositories into one, where consolidating backup storage raises the same chain and retention questions across more than one source at a time.

Where ARTESCA fits

ARTESCA is object storage software used as a backup target, presenting an S3 compatible API, with immutability available through S3 Object Lock. In a migration it is usually the destination, which means the objects arriving carry retention settings determined by the new repository configuration rather than inherited from the old one. Retention on migrated data is therefore set deliberately at configuration time.

Because ARTESCA runs on infrastructure the customer operates, the copy path between old and new repositories is internal, and its throughput can be measured in advance rather than estimated. The same applies to parallel running, since both systems sit in the customer's own facilities and the cost of keeping the old one is power, rack space and support rather than a metered charge.

Migration mechanics belong to the backup software, not to the storage. Which moves it supports and how the catalog is preserved vary by product and version, so the questions to ask before moving a repository should be settled against the specific deployment.

What to verify before the old system is retired

The test that matters is a restore, not an object count. A migration is complete when a restore point from before the cutover has been restored from the new repository, end to end, without reference to the old system. Anything less confirms that data is present, which is not the same as usable.

A workable sample is one restore point from the oldest period retained, one from the middle, one from the most recent month, and at least one sitting on a synthetic full rather than a plain incremental. Each should be restored to an isolated location and opened, with the duration recorded, since restore time from migrated data often differs from restore time at the source.

What to write down is the map. Which restore points moved and which did not, the date each remaining lock expires, the procedure for restoring from the old system while it still exists, and the date it can genuinely be powered off. That document is what stops the old repository from quietly becoming permanent, and it is the first thing the on-call admin will look for when a restore request reaches back past the cutover.

Try ARTESCA free

Immutable object storage that scales from 20TB to petabytes. Deploy a working cluster in under an hour.

Start a free test drive