Home  ›  Glossary  ›  Deduplication

What is deduplication?

Deduplication is a way of reducing stored data by keeping only one copy of each repeated chunk and replacing the repeats with references to it. Backup data deduplicates well because each night's backup of the same servers is mostly identical to the night before.

Why deduplication matters for backup storage

Retention is what makes backups useful against ransomware and against problems discovered weeks late, and retention is also what fills repositories. Thirty daily restore points of a 10 TB environment represent 300 TB of logical data. Deduplication is the reason that history fits on a fraction of the capacity, and for most mid-sized organizations the reason long retention is affordable at all.

How deduplication works

  1. Chunking. Incoming data is split into chunks, typically a few kilobytes to a few hundred kilobytes each.
  2. Fingerprinting. Each chunk is run through a hash function, producing a short fingerprint that identifies its content.
  3. Lookup. The fingerprint is checked against an index of chunks already stored.
  4. Store or reference. A new chunk is written; a known chunk is recorded only as a reference.
  5. Rehydration. On restore, references are followed and the chunks reassembled into the original data.

Chunks can be fixed-size or set by the content itself. With fixed-size chunks, inserting a few bytes near the start of a file shifts every later boundary and makes the rest of the file look new; content-defined boundaries move with the data, so only nearby chunks change.

Where and when deduplication happens

VariantWhat it means in a backup environment
Source-sideThe backup agent or proxy deduplicates, so only new chunks cross the network
Target-sideThe backup storage deduplicates, so the full stream crosses the network first
InlineDuplicates are removed before data is written
Post-processData lands first and is reduced later, which needs temporary space
Per job or globalChunks are compared within one job, or across all jobs and repositories

What deduplication means for backup teams

Ratios depend mostly on retention and change rate rather than on the product. Thirty daily backups of a 10 TB source that changes 1% a day hold about 10 + 29 × 0.1 = 12.9 TB of unique data, a ratio near 23:1. Kept for a single day, the same data reduces to roughly 1:1. A quoted ratio says little until the retention behind it is known.

Deduplication also concentrates risk. One stored chunk can be referenced by hundreds of restore points, so losing or corrupting it damages all of them, and an attacker who deletes the deduplicated store removes the entire history at once. Restores are usually slower than backups, because rehydration reads chunks scattered across the store; that shows up in restore throughput when many systems have to come back at the same time.

Order matters as well. Encryption makes identical data look different, so deduplication runs before backup encryption, and files that ransomware has already encrypted barely deduplicate at all. A sudden fall in the ratio, or a jump in new unique data in one night's backup, is often one of the first signs a backup team sees that production has been encrypted.

How deduplication relates to ARTESCA

When a backup application deduplicates data before writing it to ARTESCA over S3, ARTESCA stores the resulting objects as written, and the reduction ratio comes from the application's chunking and scope. Backup software with published ARTESCA compatibility pages includes Veeam, Commvault, Cohesity, Rubrik, HYCU, Veritas NetBackup and Zerto.

An object version locked with S3 Object Lock stays in place until its own retain-until date, in compliance mode against every user including the root account, whatever happens to the restore points that reference it. Its space is reclaimed only after the lock lapses and the application or a per-bucket S3 Lifecycle rule removes it, which ties consumed capacity to retention settings as much as to the deduplication ratio.

Related terms