Backup & recovery

Backup windows: Why jobs stop finishing overnight

Backup windows erode phase by phase. How to measure where the night goes and what to change when it no longer fits.

7 min read
Dark data center aisle with server racks and glowing cyan light paths tracing data movement between them

A backup job that used to finish by 4 a.m. now finishes at 6:30. Nobody changed anything. The next month it runs into the start of business, and the virtualization team is asked why the hosts are busy at 9 a.m. Backup windows rarely break in a single night. They erode a few minutes at a time, until one added workload or one longer chain pushes the finish past the point anyone can ignore.

The reflex is to treat this as a throughput problem and price faster disks for the repository. Sometimes that is correct. More often the job is not spending its night writing. It is waiting on a snapshot to commit, walking a change tracking map that was invalidated, or reading cold blocks from a datastore that also serves users.

Total job duration hides all of it. Nine reported hours tell nobody which were work and which were waiting. Until the night is broken into phases, any fix is a guess.

Where the overnight hours actually go

A backup job is a sequence of distinct operations, and each stops scaling for a different reason. The first is the source read. A proxy reads blocks from production storage, and that read competes with whatever else is running. A datastore comfortable during the day becomes the constraint at 2 a.m. if a database maintenance plan, a storage rebalance or a replication task was scheduled into the same hours without anyone comparing calendars.

The second is change block tracking. Incremental backups depend on the hypervisor or agent reporting which blocks changed since the last run. When that tracking is reset, and it is reset by storage migrations, certain snapshot operations, host crashes and some upgrades, the next job has no map and reads everything. The job type still says incremental. The duration looks like a full. This is the most common cause of a window that blows out one night and returns to normal the next.

The third is snapshot lifecycle. Creating a snapshot is fast. Committing it is not, since writes that accumulated during the backup must be consolidated into the base disk, and production notices because that work lands on the datastore users hit. The fourth is target write, the phase everyone assumes is the problem. The fifth is catalog and metadata work, which turns significant on jobs with very many files. The sixth is synthetic full construction, where the target assembles a new full from restore points it already holds.

Why adding storage speed often changes nothing

Replacing a repository with something faster helps exactly one phase. If target write is thirty percent of the night, halving it saves fifteen percent of the window. That is real, but it is not a rescue, and it is often smaller than the variance between two ordinary Tuesdays.

There is a second reason the upgrade underdelivers. Backup software rarely drives a target at its ceiling with the concurrency it has been given. Task counts, proxy numbers, gateway placement and stream settings cap throughput independently of the media underneath, so a repository capable of more idles between tasks. Concurrency settings deserve examination before hardware does, since they cost nothing to change. The third reason is that the read side is often the ceiling, and when it is, no target shortens the night.

Measuring each phase instead of the total

Most backup platforms already record per-task statistics: duration, processing rate, bottleneck attribution across source, proxy, network and target, and data read against data transferred. That sits in the job session detail, a click from the summary everyone looks at, and it is ignored because the summary is green.

The useful practice is to export those statistics weekly for the three or four jobs that carry the business and keep them where months can be compared. A window need not be diagnosed on the night it fails if twelve weeks of phase durations are available. The question changes from why was it slow to which phase grew, and the second has an answer.

PhaseWhat a long duration usually meansWhat to check
Source readProduction storage is the constraint, or something competes for itOther activity in the same hours, proxy placement, datastore queues
Change block trackingTracking was reset, so an incremental read everythingData read against data transferred, recent migrations, host events
Snapshot commitThe snapshot lived too long and the delta grew largePer machine duration, snapshot age at removal, datastore space
Target writeA genuine target limit, or too few tasks to reach itTask slot counts, gateway placement, network path, bottleneck figures
Catalog and metadataFile indexing or restore point bookkeeping is heavy hereFile counts per machine, indexing settings, database location
Synthetic fullThe target is assembling a full on its own scheduleWhich days run synthetic operations, chain type, target read speed

One number deserves attention. When read volume approaches the full size of the protected workload on a job that should be incremental, change tracking has been reset. Nothing else produces that signature.

The decisions available when the window cannot hold

Suppose the phases have been measured and the night does not fit. Four moves exist, and each trades something different.

Staggering is the cheapest. Jobs that all start at 22:00 contend for the same proxies, datastores and target. Spreading start times so heavy jobs do not overlap often recovers an hour with no purchase at all, provided the heavy jobs are known.

Splitting divides a large job into smaller ones. This improves parallelism and shortens snapshot lifetime per machine, reducing commit cost. It also multiplies the chains to manage and the things that fail on their own.

Changing chain type moves work rather than removing it. Forward incremental with periodic synthetic fulls shifts effort onto the target, away from production. Forever forward incremental removes the periodic full but adds a merge at the end of every run. Each option produces a different restore profile, so the choice belongs with whoever owns the recovery time expectations, not only with the operations calendar.

Accepting a longer window is the fourth move, legitimate provided the risk is named. A job finishing at 8 a.m. means the recovery point is older by the length of the overrun, and backup activity now overlaps production load. Copy jobs to a second location start later too, so copy job lag becomes part of the same decision.

Where ARTESCA fits

ARTESCA is object storage software used as a backup target, deployed on infrastructure the customer runs. In window terms it occupies the target write phase and part of the synthetic full phase, since synthetic operations against an object repository work on data the target already holds. It supports immutability through S3 Object Lock and presents an S3-compatible API, so backup software addresses it as it would any object repository.

What that means here is narrow. A capable target removes target write from the suspect list. It does not shorten source reads, restore change tracking, or commit snapshots faster. A team sizing a repository should know which phases it actually has before deciding how much target capability the window needs.

ARTESCA runs from roughly 50 TB to several petabytes and is commonly deployed alongside Veeam through resellers, so sizing and window conversations often involve the same people. That helps only if the phase data is present.

What to measure on a schedule and what to write down

A monthly ten-minute routine covers most of this. Pull session statistics for the jobs that carry the business, record total duration, per-phase duration, data read, data transferred and the reported bottleneck, and put them in the same sheet each time. Trends show in three months and are obvious in six.

Alongside that, keep a written record of the window itself: when each job is expected to finish, the time after which a late finish becomes a problem, and who gets told when it passes. Many teams hold this in one person's head, where it cannot be handed over or defended in a change review.

The last thing to record is the decision already taken. If the team accepted a longer window rather than split a job, that carries a recovery consequence, and it belongs beside the recovery point objectives it affects. Otherwise it gets rediscovered, as a surprise, on the morning it matters.

Try ARTESCA free

Immutable object storage that scales from 20TB to petabytes. Deploy a working cluster in under an hour.

Start a free test drive