A backup job that used to finish by 4 a.m. now finishes at 6:30. Nobody changed anything. The next month it runs into the start of business, and the virtualization team is asked why the hosts are busy at 9 a.m. Backup windows rarely break in a single night. They erode a few minutes at a time, until one added workload or one longer chain pushes the finish past the point anyone can ignore.
The reflex is to treat this as a throughput problem and price faster disks for the repository. Sometimes that is correct. More often the job is not spending its night writing. It is waiting on a snapshot to commit, walking a change tracking map that was invalidated, or reading cold blocks from a datastore that also serves users.
Total job duration hides all of it. Nine reported hours tell nobody which were work and which were waiting. Until the night is broken into phases, any fix is a guess.
Where the overnight hours actually go
A backup job is a sequence of distinct operations, and each stops scaling for a different reason. The first is the source read. A proxy reads blocks from production storage, and that read competes with whatever else is running. A datastore comfortable during the day becomes the constraint at 2 a.m. if a database maintenance plan, a storage rebalance or a replication task was scheduled into the same hours without anyone comparing calendars.
The second is change block tracking. Incremental backups depend on the hypervisor or agent reporting which blocks changed since the last run. When that tracking is reset, and it is reset by storage migrations, certain snapshot operations, host crashes and some upgrades, the next job has no map and reads everything. The job type still says incremental. The duration looks like a full. This is the most common cause of a window that blows out one night and returns to normal the next.
The third is snapshot lifecycle. Creating a snapshot is fast. Committing it is not, since writes that accumulated during the backup must be consolidated into the base disk, and production notices because that work lands on the datastore users hit. The fourth is target write, the phase everyone assumes is the problem. The fifth is catalog and metadata work, which turns significant on jobs with very many files. The sixth is synthetic full construction, where the target assembles a new full from restore points it already holds.
Why adding storage speed often changes nothing
Replacing a repository with something faster helps exactly one phase. If target write is thirty percent of the night, halving it saves fifteen percent of the window. That is real, but it is not a rescue, and it is often smaller than the variance between two ordinary Tuesdays.
There is a second reason the upgrade underdelivers. Backup software rarely drives a target at its ceiling with the concurrency it has been given. Task counts, proxy numbers, gateway placement and stream settings cap throughput independently of the media underneath, so a repository capable of more idles between tasks. Concurrency settings deserve examination before hardware does, since they cost nothing to change. The third reason is that the read side is often the ceiling, and when it is, no target shortens the night.
Measuring each phase instead of the total
Most backup platforms already record per-task statistics: duration, processing rate, bottleneck attribution across source, proxy, network and target, and data read against data transferred. That sits in the job session detail, a click from the summary everyone looks at, and it is ignored because the summary is green.
The useful practice is to export those statistics weekly for the three or four jobs that carry the business and keep them where months can be compared. A window need not be diagnosed on the night it fails if twelve weeks of phase durations are available. The question changes from why was it slow to which phase grew, and the second has an answer.
| Phase | What a long duration usually means | What to check |
|---|---|---|
| Source read | Production storage is the constraint, or something competes for it | Other activity in the same hours, proxy placement, datastore queues |
| Change block tracking | Tracking was reset, so an incremental read everything | Data read against data transferred, recent migrations, host events |
| Snapshot commit | The snapshot lived too long and the delta grew large | Per machine duration, snapshot age at removal, datastore space |
| Target write | A genuine target limit, or too few tasks to reach it | Task slot counts, gateway placement, network path, bottleneck figures |
| Catalog and metadata | File indexing or restore point bookkeeping is heavy here | File counts per machine, indexing settings, database location |
| Synthetic full | The target is assembling a full on its own schedule | Which days run synthetic operations, chain type, target read speed |
One number deserves attention. When read volume approaches the full size of the protected workload on a job that should be incremental, change tracking has been reset. Nothing else produces that signature.
The decisions available when the window cannot hold
Suppose the phases have been measured and the night does not fit. Four moves exist, and each trades something different.
Staggering is the cheapest. Jobs that all start at 22:00 contend for the same proxies, datastores and target. Spreading start times so heavy jobs do not overlap often recovers an hour with no purchase at all, provided the heavy jobs are known.
Splitting divides a large job into smaller ones. This improves parallelism and shortens snapshot lifetime per machine, reducing commit cost. It also multiplies the chains to manage and the things that fail on their own.
Changing chain type moves work rather than removing it. Forward incremental with periodic synthetic fulls shifts effort onto the target, away from production. Forever forward incremental removes the periodic full but adds a merge at the end of every run. Each option produces a different restore profile, so the choice belongs with whoever owns the recovery time expectations, not only with the operations calendar.
Accepting a longer window is the fourth move, legitimate provided the risk is named. A job finishing at 8 a.m. means the recovery point is older by the length of the overrun, and backup activity now overlaps production load. Copy jobs to a second location start later too, so copy job lag becomes part of the same decision.
Where ARTESCA fits
ARTESCA is object storage software used as a backup target, deployed on infrastructure the customer runs. In window terms it occupies the target write phase and part of the synthetic full phase, since synthetic operations against an object repository work on data the target already holds. It supports immutability through S3 Object Lock and presents an S3-compatible API, so backup software addresses it as it would any object repository.
What that means here is narrow. A capable target removes target write from the suspect list. It does not shorten source reads, restore change tracking, or commit snapshots faster. A team sizing a repository should know which phases it actually has before deciding how much target capability the window needs.
ARTESCA runs from roughly 50 TB to several petabytes and is commonly deployed alongside Veeam through resellers, so sizing and window conversations often involve the same people. That helps only if the phase data is present.
What to measure on a schedule and what to write down
A monthly ten-minute routine covers most of this. Pull session statistics for the jobs that carry the business, record total duration, per-phase duration, data read, data transferred and the reported bottleneck, and put them in the same sheet each time. Trends show in three months and are obvious in six.
Alongside that, keep a written record of the window itself: when each job is expected to finish, the time after which a late finish becomes a problem, and who gets told when it passes. Many teams hold this in one person's head, where it cannot be handed over or defended in a change review.
The last thing to record is the decision already taken. If the team accepted a longer window rather than split a job, that carries a recovery consequence, and it belongs beside the recovery point objectives it affects. Otherwise it gets rediscovered, as a surprise, on the morning it matters.
