A backup job that used to run in two hours now runs in five, and the summary screen reports a bottleneck without saying what to do about it. The block being written to the bucket at that moment has already passed through a snapshot, a transport mode, a proxy, a compression stage, an encryption stage, a gateway and an HTTPS request. Each of those stages has its own ceiling and its own way of failing.
Most slow job investigations start at the storage, because that is the component with a vendor attached to it. The path has six or seven distinct hops and the storage is the last one. Without a model of the others, troubleshooting becomes a sequence of guesses that each take a night to test.
The model does not need to be detailed. It needs to be complete, so that every hop is named and every hop has a question attached to it. A slow job then stops being a mystery and becomes an ordered list of things to eliminate.
From the datastore to the proxy
Everything begins with a snapshot. The hypervisor freezes the virtual disk so that writes made during the backup land in a delta file rather than in the base disk. Creating the snapshot is fast. Removing it is not, since the delta must be consolidated back onto the same datastore that serves production. A long backup produces a large delta and a longer commit, which is why job duration and production impact are not independent.
The proxy then reads the frozen disk, and how it reads depends on the transport mode. Direct storage access has the proxy read from the storage fabric itself, with no hypervisor host in the path for the data. Virtual appliance mode attaches the disk to a proxy running as a virtual machine on a host with datastore access, so the read follows the host storage path. Network mode moves data over the management network, usually the slowest option and often the only one available.
Transport mode is the largest architectural decision in this path, and it is frequently inherited rather than chosen. A proxy that lost its storage visibility falls back to a slower mode and keeps working, which is why a job that doubled in duration with no configuration change is worth checking here first. The mode actually used is recorded per task in the session detail.
What the proxy does before anything leaves it
The proxy is where the data stops being a copy of blocks and becomes a backup file, through three transformations that each consume processor cycles on the proxy itself. Compression reduces the volume that will cross the network and land in the bucket. Its effectiveness depends on the workload, since machines full of already compressed media behave nothing like file servers full of documents, so compression by workload is a better planning input than one estate wide assumption. Deduplication removes repeated blocks within the scope the software applies it to, and that scope matters more than the ratio.
Encryption, where enabled, is applied here too, before the data leaves the proxy. That ordering matters: data encrypted at the proxy arrives at the storage already opaque, so anything the storage might do with the content has nothing to work with. The choice is legitimate either way, but it should be made knowingly rather than discovered later.
The practical question here is whether the proxy has enough processor and memory for all three transformations across every concurrent task assigned to it. An oversubscribed proxy does not fail. It queues, and the symptom looks like a slow target.
From the proxy to the bucket
Object storage is reached over HTTPS, and the gateway role speaks that protocol, either on a dedicated server or on a component the software places automatically. Every byte crosses it, so gateway server placement decides which network segments carry the traffic and where the cost of session handling lands. Placed carelessly, it produces a path through a firewall nobody intended to put in the data path.
Above the gateway sits the request itself. Backup data is written as objects, in many parallel requests rather than one large stream, so throughput depends on how many requests are in flight at once. A single stream test against a bucket therefore rarely reflects what a job achieves. The TLS session also has to be established and trusted, which is where certificate problems appear as a job that fails immediately rather than one that runs slowly.
Finally, each object lands, and with immutability enabled it acquires a retention period at that moment. The clock starts at the write, not at the end of the job.
| Hop | What limits throughput there | What failure looks like |
|---|---|---|
| Snapshot and datastore read | Production storage queues, competing activity in the same hours | Long per machine duration, snapshot commit warnings, production complaints |
| Transport mode | Which path the proxy has to the disk | Silent fallback to a slower mode after a fabric or permission change |
| Proxy processing | Processor and memory against concurrent tasks | Tasks queue, rate falls as task count rises, target blamed |
| Gateway and network path | Link speed, firewall inspection, session handling cost | Throughput capped well below link speed, retries under load |
| HTTPS request to the bucket | Requests in flight, object size, latency on the path | Immediate failure on certificate or credential problems |
The second row is the one most often missed, because nothing about it announces itself. The job still succeeds.
Where the ceiling actually sits
Backup software reports a bottleneck for each task, attributing time to source, proxy, network and target. That attribution is useful and frequently misread. It names where time was spent, not where the fixable constraint is, and on a healthy job one component is always highest because something has to be.
The number worth pairing with it is data read against data transferred. When a job that should be incremental reads close to the full size of the protected machines, change tracking was reset and the path is doing full read work regardless of what every other figure says. That comparison resolves many investigations before any hop is examined, and it is the first check when a backup window moves without explanation.
Restores traverse the same hops in reverse, but not symmetrically. Reassembling a disk image from many small objects stresses request handling rather than sequential throughput, so a path tuned for backup can still disappoint on recovery. Finding the slowest step in a restore is a separate exercise.
Where ARTESCA fits
ARTESCA is object storage software used as a backup target, deployed on infrastructure the customer runs. In this path it occupies the last hop only. It receives HTTPS requests from the gateway, stores the resulting objects, and applies S3 Object Lock retention where the repository is configured for immutability, through an S3-compatible API.
Because it sits on infrastructure the customer runs, the network between gateway and bucket is a local path under the same team's control rather than an internet route. Latency and available bandwidth on that segment become measurable facts rather than conditions to be inferred.
What it does not affect is any hop before the gateway. Snapshot behavior, transport mode, proxy sizing and change tracking are properties of the virtualization platform and the backup software. A capable target removes the last hop from the suspect list, which is valuable mainly because it is the hop people suspect first.
What to measure and write down
The record worth keeping is short. For the three or four jobs that carry the business, capture the transport mode actually used, the reported bottleneck, data read, data transferred and total duration once a month in the same sheet. Twelve rows a year make any future regression visible without a forensic exercise.
Alongside the numbers belongs a written description of the path itself: which proxies serve which clusters, which transport mode each is expected to use, where the gateway runs, and which segments the traffic crosses. Most teams hold this in one person's memory, where it cannot be checked during an incident.
Finally, the path deserves re-verification after any change to storage presentation, host configuration or network policy, since those alter transport mode and routing without touching the backup configuration at all. Checking one job session afterward is cheaper than discovering the fallback three months later.
