Integrations

Veeam SOBR: Performance tier vs. capacity tier

Copy mode, move mode and the operational restore window decide what stays local and where a restore reads from.

7 min read
Dark data center showing two rows of racks joined by glowing cyan data paths flowing from one row to the other

A scale-out backup repository presents one target to the backup jobs and two very different sets of storage behind it. The performance tier holds block storage extents. The capacity tier holds object storage. Most teams configure both in an afternoon, then discover months later that they cannot explain why a given restore point is where it is, or why the performance tier filled up while the capacity tier looked half empty.

The confusion usually comes from a model of tiering borrowed from somewhere else: hot data here, cold data there, moved on age. A scale-out repository does move data on age, but age is only one condition, and the others decide what is actually sitting on the fast storage tonight.

How data moves between the two tiers

Two mechanisms operate on the capacity tier, independently of each other. Copy mode writes a copy of backup data to object storage shortly after it lands on the performance tier, leaving the original in place. The result is two copies, one local and one on object storage, and the local copy goes only when normal retention removes it.

Move mode relocates data off the performance tier once it becomes eligible, leaving behind the metadata needed to know the restore point exists and where its data now lives. The restore point stays visible in the console. Its data does not stay on the fast storage.

The two can run together, which is where most deployments end up. Copy mode gives a second copy early, satisfying the requirement for a copy on different storage. Move mode reclaims performance tier space later. Whether both are enabled changes how much data sits where, so it is worth confirming what is set rather than what was intended.

Eligibility for move mode is where the model usually breaks. Age alone is not sufficient. Data belonging to a chain still being written to is generally not moved, however old the individual restore points are, since moving part of an active chain would break what the next incremental depends on. That one rule explains most cases of a performance tier that will not drain.

The operational restore window is a policy, not a prediction

The operational restore window is the number of days of recent restore points intended to stay on the performance tier. It expresses a judgment: restores requested within this period should be fast, and restores older than it can afford to read from object storage.

Setting it well means knowing when restores are actually requested. In most environments the distribution weights heavily toward the last few days, with a thin tail for cases where corruption went unnoticed for weeks. The window should cover the dense part, not the tail, since covering the tail means keeping months of data on the expensive tier for rare events.

The window interacts with chain structure. A forever forward incremental chain stays active indefinitely, so a large amount of data can remain ineligible for move mode even with a short window. Periodic fulls close the previous chain and make it eligible. This is one reason the chain type is a recovery and capacity decision at once, not a scheduling preference. Eligibility rules have changed across releases, so the behavior in the deployment at hand is worth observing over a few weeks rather than assuming it.

SituationWhere the restore reads fromWhat to check if it is slower than expected
Inside the operational restore windowPerformance tier, if the data is still thereWhether move mode already relocated it, and why
Older than the window, move mode onCapacity tier, across the networkPath throughput and latency to the object endpoint
Older than the window, copy mode onlyPerformance tier, until retention removes itPerformance tier free space, since nothing was reclaimed
Restore from a chain still being writtenPerformance tier, since active chains are not relocatedWhether the chain ever closes, and what closes it
Instant recovery of an older machineCapacity tier, read on demand while runningWhether the machine is usable, not merely booted
Performance tier extent in maintenanceWhichever tier still holds the dataExtent state and whether an evacuation completed

Where a restore reads from, and why it matters

The console presents restore points uniformly. It does not make the tier obvious at the moment someone is choosing a recovery point under pressure, and the two tiers can differ substantially in how quickly they deliver data. A restore from a few days ago and one from six weeks ago may be the same operation with very different durations.

This is sharpest for instant recovery, where a machine runs directly from the backup while data is read on demand. Running from the performance tier and running from object storage across a network are different experiences, and the difference shows as application responsiveness rather than failure. Anyone relying on a machine being usable after it boots should have tested it from the tier a real incident will use.

The practical consequence is that a recovery time commitment should name a period, not only a target. A stated time for restore points inside the operational restore window, and a longer figure outside it, is honest. A single number that silently assumes the performance tier is not.

How sizing the two tiers against each other goes wrong

The most common error is sizing the performance tier for the operational restore window alone. That figure omits the active chain, which stays regardless of the window, and whatever copy mode leaves in place until retention removes it. A performance tier sized to the window runs out of space, and the symptom is a full extent rather than an obvious configuration fault.

The second error is treating the capacity tier as where old data lives. With copy mode enabled it holds a copy of everything from the moment it is written, so its size tracks total retention rather than the portion outside the window. Where immutability is applied, locked objects cannot be removed early, which makes this the tier with the least flexibility when space runs short.

The third is sizing both tiers once and leaving them. Retention changes, new workloads and changes in how many restore points are kept move the split between the tiers without anyone adjusting a setting. Reviewing the ratio as part of normal growth and churn planning catches this before an extent fills.

Where ARTESCA fits

ARTESCA is object storage software used as a backup target, deployed on infrastructure the customer runs. In a scale-out backup repository it occupies the capacity tier, addressed through an S3-compatible API, with immutability through S3 Object Lock where the data requires it.

Because it sits on customer infrastructure rather than at a provider, the reads a capacity tier restore performs travel a local network path with no per-request or egress charge attached. That changes the economics of restoring older restore points, and the path can be measured and improved by the team that owns the rest of the network.

Capacity runs from roughly 50 TB to several petabytes. In a copy mode deployment the capacity tier holds the full retention rather than only the older portion, so sizing should start from total retention rather than the data expected to age out.

What to check on a schedule and what to write down

Once a quarter, confirm three things against the configuration rather than memory: which modes are enabled on the capacity tier, what the operational restore window is set to, and how much free space each performance tier extent has. All three drift as jobs are added, and none announce a change.

Alongside that, record the tier assumption behind each recovery commitment. For each workload with a stated recovery time, note whether the figure was measured against the performance tier or the capacity tier. That is the most useful line in the document, since it is the assumption that fails quietly when data moves.

Finally, test one restore from each tier at least twice a year, using a workload that resembles something real rather than the smallest available machine. The performance tier restore confirms the normal case. The capacity tier restore confirms the case that applies on the day the useful recovery point turns out to be older than anyone hoped.

Try ARTESCA free

Immutable object storage that scales from 20TB to petabytes. Deploy a working cluster in under an hour.

Start a free test drive