Integrations

Veeam backup chains: What changes on object storage?

Blocks become objects, metadata becomes objects, and nothing is modified in place. What that changes for restores and growth.

7 min read
A chain of glowing data blocks dissolving into many separate objects across a dark server hall

A backup chain is the same logical thing wherever it lives: a full restore point, a run of incrementals that depend on it, and the metadata that says which blocks belong to which point. Moving that chain from a file system repository to object storage does not change the logic at all. It changes almost everything about how the chain is stored, and most of the operational surprises come from that gap.

On a file system the chain is a small number of large files. The software can open a full backup file, seek to an offset, and write over a region of it. Merging an old incremental into the full is that operation performed many times, and deleting a restore point removes a file so the space comes back at once.

Object storage offers none of those primitives. There is no seek, no partial write, and no way to modify an object that already exists. Every one of those operations is re-expressed as writes of new objects and deletes of old ones, and the operational consequences follow from that translation.

The same chain, stored a different way

In an object repository the blocks that would have lived inside a full backup file are stored as individual objects. Each one is written once, named by the repository layout rather than by a path a person would recognize, and never touched again. A restore point is not a file. It is a list of object keys plus the information needed to reassemble them in order.

That list has to live somewhere, so metadata becomes its own population of objects, written alongside the data and updated far more often. A repository is therefore two very different workloads sharing a bucket: many write once data objects, and fewer metadata objects rewritten whenever the chain changes.

The same split governs direct to object configurations, where backup jobs write to the bucket without a performance tier in front of them. Whether the first write lands on object storage or arrives later from a performance extent is a separate design decision, but the storage layout of the chain is the same either way.

Metadata becomes its own set of objects

Metadata objects are small, read constantly and rewritten frequently. Every job run adds to them, every retention pass edits them, and every restore reads them before a single block of data is fetched. Losing a data object costs one block; losing the metadata for a chain costs the ability to interpret any of it.

Because there is no in place update, an edit to metadata is a new object version written over the old key. In a versioned bucket, and Object Lock requires versioning, the previous version does not disappear. It becomes a noncurrent version that continues to consume capacity until something expires it. A repository with a busy chain accumulates metadata versions faster than most people expect.

Synthetic work without in place modification

A synthetic full on a file system is built by reading blocks from the existing chain and writing them into a new full file. On object storage that would be wasteful, since the blocks are already present as objects. The synthetic operation becomes a metadata operation: a new restore point references existing block objects, and only genuinely new blocks are uploaded.

Merging an old incremental into the full follows the same pattern. Nothing is rewritten. The metadata is updated so that the merged point no longer exists as a separate entity, and the blocks that are now referenced by nothing become candidates for deletion. Cleanup is therefore not truncation of a file, it is a set of delete requests against objects whose reference count has reached zero.

This is why capacity behaves unintuitively. Deleting a restore point frees only the blocks that no surviving point references, which is usually a fraction of the point's nominal size. The tradeoff between full and incremental restore points looks different once storage is shared this way, because the cost of keeping an extra point is the unique blocks it holds rather than its apparent size.

Operation on a file system repositoryEquivalent on object storageWhat changes operationally
Synthetic full written as a new full fileNew metadata referencing existing block objects, plus new blocks onlyFar less data written; work moves to metadata and reference tracking
Merge of the oldest incremental into the fullMetadata rewrite, then deletes for blocks nothing referencesNo large file rewrite; space returns only once deletes complete
Deleting a restore pointDeleting the objects no surviving point referencesFreed capacity rarely matches the point's nominal size
Reading a restore point during a restoreEnumerating metadata, then fetching many individual objectsThroughput depends on request concurrency, not sequential read speed
Compacting or defragmenting a full backup fileNo equivalent, since objects are written onceNothing to schedule, but unreferenced objects need their own cleanup

What a restore has to fetch and how the repository grows

A restore from a file system repository is largely sequential reading with some seeking. A restore from an object repository is a metadata lookup followed by a very large number of individual object fetches, each with its own request overhead. The total bytes are similar; the number of operations is not, and that is what sets the elapsed time.

Block size sets the exchange rate. Smaller blocks mean better change detection and more objects per terabyte, which means more requests for the same restore. Concurrency on the mover or gateway then decides how many of those requests are in flight at once. When a restore from object storage disappoints, the limiting factor is usually request parallelism rather than link speed, which is the same pattern that shows up when looking for the slowest step in a restore path.

Growth follows the same logic in reverse. Because nothing is rewritten, the repository grows by the unique blocks each run contributes, plus metadata, plus whatever noncurrent versions have not yet been expired. It shrinks only when deletes are issued and accepted. A capacity graph that rises smoothly and never falls usually means deletes are failing or noncurrent versions are being kept, not that the environment grew.

Where ARTESCA fits

ARTESCA is object storage software used as a backup target, presented over an S3 compatible API, with immutability available through S3 Object Lock. The chain layout described here is the backup software's, not the storage's: the repository sees objects, versions, retention dates and delete requests, and enforces the rules the S3 API defines for them.

Two properties matter for chains. Versioning is required for Object Lock, so overwritten metadata and deleted blocks leave noncurrent versions that consume capacity until expired. And because the bucket has no knowledge of chain structure, it accepts any delete the backup software is entitled to issue, so reference tracking stays entirely the software's responsibility.

The system runs on infrastructure the customer operates, so object counts, request rates and the network path between mover and bucket are all local and directly measurable. When a restore is slower than expected, the request pattern can be observed at both ends rather than inferred.

What to check after the first month on object storage

The first useful measurement is a restore of a mid chain point, not the newest one, timed end to end and compared against the same restore from the previous repository. A newest point restore touches the fewest objects and flatters the configuration. A point several incrementals deep exercises the metadata lookup and the fan out of fetches a real recovery involves.

The second is a capacity reading taken after a full retention cycle, with object and version counts alongside the byte total. If bytes stay flat while the point count falls, deletes are being issued but noncurrent versions survive. If both stay flat, the deletes themselves are failing, often against immutability.

Worth recording once and keeping: the block size in use, the object count per terabyte that results, the concurrency configured on the movers, and whether the repository sits behind a performance tier or receives writes directly. Those four values explain most of what the repository does later, and they are also the inputs to any sensible discussion about how a performance tier and a capacity tier divide the work.

Try ARTESCA free

Immutable object storage that scales from 20TB to petabytes. Deploy a working cluster in under an hour.

Start a free test drive