ARTESCA Blog | Backup, recovery and cyber resilience

Backup copy jobs: How far behind can you afford to be?

Written by Joshua Silvia | Sep 20, 2026, 3:42:52 AM

Most backup estates keep a second copy somewhere: another site, a separate repository, an object target. The copy jobs report success, the dashboard stays green, and everyone assumes the second copy is roughly current. It never is. It sits behind the primary by some amount of time, sometimes a few hours and sometimes several days, and that amount is almost never written down. The useful question is how far behind the team can afford to be.

The reflex answer is that the copy job runs nightly, so the second copy is at most a day old. That describes the schedule, not the outcome. A copy job cannot begin until the restore point it depends on is finished and closed. It has to move data across whatever link it has, inside whatever window it has been permitted, and it has to complete before the next one queues behind it. Every one of those can slip without raising an error.

The gap matters because the second copy is reached for exactly when the first is unusable. At that moment the recovery point is not the one on the schedule. It is whatever the secondary actually holds.

What actually sets the lag

A copy job is a dependent job. It reads a restore point that already exists and writes it elsewhere, so its clock starts when the source job finishes rather than when the copy job is scheduled. A group whose backup window has been eroding for months has a second copy that is late by the same margin plus its own runtime, every night, without anything on the copy side having changed.

The second constraint is the permitted window. Copy jobs are usually confined to hours when they will not compete with production traffic or with the backup jobs themselves. If the volume to move exceeds what the link carries in those hours, the job pauses and resumes the next night. It has not failed. It simply has not finished.

Bandwidth is the third, and it is governed by change rate rather than total protected size. An estate with low daily churn copies comfortably over a modest link. The same estate after a database reorganization or a migration between datastores produces a delta several times larger, and the copy stream does not expand to match it.

The fourth is chain dependency. An incremental point cannot be copied before its parents have arrived, so one gap early in a chain stalls everything behind it however much bandwidth is free. The fifth is throttling: rate limits set once during a link upgrade and left in place afterward.

Why the gap grows quietly instead of failing loudly

Backup software is deliberately resilient about copy operations, and that is what hides the problem. A copy that cannot complete in its window is treated as incomplete rather than failed. A point that no longer exists at the source is skipped. A retry that succeeds on the third attempt reports success. None of that produces a red status.

Reporting compounds it. Most consoles show the last run state and the count of restore points at the target, both of which look healthy while the newest point there is four days old. A count answers how many, not how recent.

There is also a retention interaction that surprises teams. When the secondary is near its capacity or retention ceiling, some configurations stop accepting new points rather than removing old ones, particularly where tiering policies or immutability periods prevent deletion. The chain then freezes at a fixed date while jobs report normal activity.

Measuring the real age of the newest secondary restore point

The measurement worth having is one number per protected group: the wall clock age, right now, of the newest restore point on the secondary that could actually be recovered. Not job status, not point count, not an average.

Usable is doing work in that sentence. A restore point whose chain is incomplete, whose encryption key is not available at the recovery location, or which sits in an archive tier with a retrieval delay is not available at the speed the number implies. A twelve hour lag with a four hour retrieval time is a sixteen hour recovery point.

Most platforms expose this through a report or an API listing restore points per repository with creation timestamps. A scheduled query recording the newest timestamp per job per target, once a day, builds a history within a month that makes a growing lag visible before it becomes an incident.

SignalWhat it usually meansWhat to check
Jobs green, newest secondary point days oldRunning but not completing in the windowTransfer volume against window length, pause and resume events
Lag stable midweek, larger after weekendsWeekend activity inflates the deltaWhat else runs Friday to Sunday, chain type on those days
Lag grew suddenly and stayedA one-time event moved the baselineData read against data transferred, recent infrastructure changes
One job far behind, the rest currentA stalled chain or one oversized workloadChain integrity at the target, per machine transfer sizes, retries
Point count normal, all dates oldNew points are not being acceptedFree space, retention settings, immutability periods, tier behavior

Deciding the acceptable figure from the scenarios that would use the copy

Acceptable lag cannot be derived from the recovery point objective alone, because that objective is usually written against the primary copy. The number has to come from the scenarios in which the second copy is the one being restored, and those are specific enough to list. Loss of the primary repository. Loss of the site. Restore points on site that cannot be trusted after a compromise.

Each implies a different tolerance. A hardware failure at the primary repository is survivable with a day of lag, since the event is discrete and the data before it is intact. After a compromise, the age of the newest copy matters less than whether the copy set reaches far enough back to hold a restore point taken before the intrusion. A site loss is strictest, since there the lag is the data loss.

Three numbers written down, one per scenario, turn an unbounded worry into a design constraint. Cutting lag means more bandwidth, a longer window or less data in flight, and those are budget conversations that go better with a target attached. The count of copies is the part of the rule everyone remembers. The currency of each copy decides what a recovery costs.

Where ARTESCA fits

ARTESCA is object storage software used as a backup target, commonly as the second copy in a design where the first lands on faster local storage. It presents an S3 compatible API, so copy jobs write to it as they would to any other object repository. Immutability is available through S3 Object Lock, applied per object for a defined period.

Two properties bear on lag. Because it runs on infrastructure the customer operates, the path between the primary repository and the second copy is a network the team controls and can measure. And because retention is expressed as object lock periods rather than as a pool that fills, capacity and retention stay separate questions with separate signals.

Neither removes the need to measure. Copy lag is a property of the job, the window and the link rather than of the target.

What to measure on a schedule and what to write down

Record the newest secondary restore point age, per protected group, once a day, and keep the history. One daily timestamp per job answers the question that matters during an incident without anyone having to go looking.

Alert on age rather than on job status. A threshold somewhat above the expected lag, with a second higher one that escalates, catches the drift that job failures never report. Both belong to the protected group rather than to the estate, since a file server and a finance database do not deserve the same number.

Write down the agreed acceptable lag per scenario, the measured lag today, and the date of the last verified restore performed from the secondary rather than the primary. Those three facts on one page, reviewed quarterly, keep the second copy from quietly becoming decorative.