A repository sized for hypervisor backups often absorbs agent backups badly. Virtual machine jobs arrive as a small number of large, scheduled, predictable streams. Agent jobs arrive as dozens or hundreds of small ones, on schedules the backup team does not fully control, from machines that may be offline for a month and then all reappear on the same Monday morning. The target bucket is the same. The load is not.
The obvious answer is to add the agents to the object storage repository that already exists and leave the tuning alone. That holds until the agent count grows, or a group of traveling laptops comes back at once, and the hypervisor jobs start finishing late for reasons nobody changed. The repository did not get slower. It got busier in a shape the original sizing never described.
Three properties of agent load decide whether sharing a target works: how many concurrent streams it produces, how little its schedule can be predicted, and the fact that retention is owned per machine rather than per job.
How agent load differs from hypervisor load
A hypervisor job protects many machines through a small number of proxies, and the backup server decides how many tasks run and when. Parallelism is a setting. An agent deployment inverts that. Each machine runs its own agent and negotiates its own path to the repository, so twenty agents finishing local processing in the same minute become twenty simultaneous writers, with no job definition above them capping the total.
The data per stream is smaller and far less uniform. A laptop incremental may be a few hundred megabytes of changed blocks, while a physical database server pushes more than anything in the virtual estate. Schedules vary just as widely, since agent jobs usually run at a time of day with catch up behavior if the machine was off. The number of active streams on a given night is therefore not a constant the team chose. It is an outcome of who happened to be logged in.
Concurrency and the shape of the task queue
Concurrency on a shared repository is finite, whether the limit is expressed as repository task slots, gateway task slots, or the practical ceiling of the network path. When agents and hypervisor jobs draw from the same pool, the agents can occupy it in a way that is hard to see afterward, because each individual agent task is short and unremarkable.
Two failure modes follow. The first is starvation, where hypervisor tasks queue behind agent tasks and the window stretches without any single job looking slow. The second is thrash, where more concurrent streams are accepted than the path can sustain, so every stream degrades and total throughput falls below what fewer streams would have achieved. Both look like a storage problem in the summary view, and neither is.
The defense is to make the limit explicit rather than emergent. Where the software allows a task limit on the repository, on the gateway and on the job, the agent population should be bounded somewhere, so its peak cannot displace jobs with tighter recovery expectations. Which controls exist varies by version and by how the repository was defined, so concurrency settings in the deployment at hand are worth reading rather than assumed.
Capacity planning when the machine count moves
Sizing a hypervisor repository starts from a known set of machines and a known change rate. Sizing for agents starts from a headcount, which moves for reasons unrelated to IT. A department hires six people and six endpoints appear. A project ends and ten machines stop checking in without being removed from the console, so their restore points age in place unwatched.
Front end data per machine is also hard to bound. Endpoints accumulate local copies of material that already lives on a file server, and physical servers kept outside virtualization are frequently the ones with the largest local volumes, which is often why they were never virtualized. A model built on an average per machine figure will be wrong in both directions at once.
The workable approach is to size the agent share separately, express it as a band rather than a number, and attach that band to an assumption about machine count growth. That keeps agent growth visible next to the repository sizing work rather than buried inside it.
| Signal | What it usually means | What to check |
|---|---|---|
| Hypervisor jobs finish late with no change in their own data volume | Agent tasks are consuming shared concurrency | Task slot usage on repository and gateway by hour, agent start times |
| Total throughput falls as stream count rises | The path is oversubscribed, not the media | Concurrent task count against measured throughput, network utilization |
| Capacity grows faster than the protected machine count | Per machine front end data is larger than modeled | Data written per agent, local volumes on physical servers |
| Restore points exist for machines nobody recognizes | Decommissioned machines still under retention | Last successful session per agent, console membership against asset records |
| A burst of agent activity at an unusual hour | Missed schedules catching up together | Catch up behavior in the agent policy, machines offline for long periods |
The fourth row surprises teams. On an immutable repository, a machine that left the company six months ago still holds capacity until its locks expire.
Immutability when retention is per machine
Immutability on an object repository prevents deletion of objects until a retention period expires, and that behavior does not care whether the objects came from a hypervisor job or a laptop. What changes with agents is the bookkeeping, since each machine effectively carries its own chain and its own expiry schedule.
Capacity release therefore becomes per machine and staggered rather than aligned to a job cycle, so the repository does not empty in steps that correlate with anything on the backup calendar. Removing a machine from protection also does not remove its data, which remains until the locks expire. Teams setting immutability and retention settings should decide deliberately how long a departed endpoint stays retrievable, because otherwise the default decides.
This is the strongest argument for giving the agents a repository of their own. Endpoints and hypervisor workloads rarely deserve identical retention, and one set of settings forces one of them to be wrong. Splitting also lets capacity alerts sit at thresholds appropriate to each population.
Where ARTESCA fits
ARTESCA is object storage software used as a backup target, deployed on infrastructure the customer runs. It presents an S3-compatible API and supports immutability through S3 Object Lock, so backup software addresses it as it would any object repository, and agent workloads reach it through the same repository definition and gateway path as any other job.
Because ARTESCA runs from roughly 50 TB upward, a separate repository for agents does not necessarily mean separate hardware. Bucket separation with distinct retention settings and distinct repository definitions is usually enough to give the two populations different rules, and it keeps concurrency accounting legible, since each repository definition carries its own task limits.
What the storage layer does not solve is the source side. Agent scheduling, catch up behavior and the machine inventory live in the backup software and in the asset records. A target with room and a clear immutability model removes one variable. It does not identify which laptops stopped reporting.
What to check on a schedule
A monthly review covers most of this. Record the number of agents that completed successfully, the number that have not reported in thirty days, peak concurrent tasks against the configured limit, and capacity consumed by the agent share against the hypervisor share. Four numbers in one place make growth and contention visible early.
Alongside that, keep a written statement of the retention decision for endpoints and for physical servers, including how long data from a removed machine is expected to remain and who may lift a hold. That prevents the common situation where nobody can explain why a former employee's laptop backups still occupy space.
The last item is a rehearsed restore from an agent backup, not only a virtual machine restore. Bare metal recovery of a physical server and file level recovery from an endpoint exercise different code paths, and a team that has only ever restored virtual machines does not yet know how long either takes.
