What is deduplication?
Deduplication is a way of reducing stored data by keeping only one copy of each repeated chunk and replacing the repeats with references to it. Backup data deduplicates well because each night's backup of the same servers is mostly identical to the night before.
Why deduplication matters for backup storage
Retention is what makes backups useful against ransomware and against problems discovered weeks late, and retention is also what fills repositories. Thirty daily restore points of a 10 TB environment represent 300 TB of logical data. Deduplication is the reason that history fits on a fraction of the capacity, and for most mid-sized organizations the reason long retention is affordable at all.
How deduplication works
- Chunking. Incoming data is split into chunks, typically a few kilobytes to a few hundred kilobytes each.
- Fingerprinting. Each chunk is run through a hash function, producing a short fingerprint that identifies its content.
- Lookup. The fingerprint is checked against an index of chunks already stored.
- Store or reference. A new chunk is written; a known chunk is recorded only as a reference.
- Rehydration. On restore, references are followed and the chunks reassembled into the original data.
Chunks can be fixed-size or set by the content itself. With fixed-size chunks, inserting a few bytes near the start of a file shifts every later boundary and makes the rest of the file look new; content-defined boundaries move with the data, so only nearby chunks change.
Where and when deduplication happens
| Variant | What it means in a backup environment |
|---|---|
| Source-side | The backup agent or proxy deduplicates, so only new chunks cross the network |
| Target-side | The backup storage deduplicates, so the full stream crosses the network first |
| Inline | Duplicates are removed before data is written |
| Post-process | Data lands first and is reduced later, which needs temporary space |
| Per job or global | Chunks are compared within one job, or across all jobs and repositories |
What deduplication means for backup teams
Ratios depend mostly on retention and change rate rather than on the product. Thirty daily backups of a 10 TB source that changes 1% a day hold about 10 + 29 × 0.1 = 12.9 TB of unique data, a ratio near 23:1. Kept for a single day, the same data reduces to roughly 1:1. A quoted ratio says little until the retention behind it is known.
Deduplication also concentrates risk. One stored chunk can be referenced by hundreds of restore points, so losing or corrupting it damages all of them, and an attacker who deletes the deduplicated store removes the entire history at once. Restores are usually slower than backups, because rehydration reads chunks scattered across the store; that shows up in restore throughput when many systems have to come back at the same time.
Order matters as well. Encryption makes identical data look different, so deduplication runs before backup encryption, and files that ransomware has already encrypted barely deduplicate at all. A sudden fall in the ratio, or a jump in new unique data in one night's backup, is often one of the first signs a backup team sees that production has been encrypted.
How deduplication relates to ARTESCA
When a backup application deduplicates data before writing it to ARTESCA over S3, ARTESCA stores the resulting objects as written, and the reduction ratio comes from the application's chunking and scope. Backup software with published ARTESCA compatibility pages includes Veeam, Commvault, Cohesity, Rubrik, HYCU, Veritas NetBackup and Zerto.
An object version locked with S3 Object Lock stays in place until its own retain-until date, in compliance mode against every user including the root account, whatever happens to the restore points that reference it. Its space is reclaimed only after the lock lapses and the application or a per-bucket S3 Lifecycle rule removes it, which ties consumed capacity to retention settings as much as to the deduplication ratio.
Related terms
- Backup target: the storage system a backup application writes restore points to.
- Retention policy design: deciding how long restore points are kept and at what granularity.
- Backup encryption: encrypting backup data so a stolen copy cannot be read.
- Restore throughput: the rate at which data comes back from backup during recovery.
