What is continuous data protection (CDP)?
Continuous data protection (CDP) is a method that captures every write to protected data as it happens and records it in a time-stamped journal. Data can be recovered to almost any moment the journal covers, instead of only to the last scheduled backup. Veeam's implementation for VMware virtual machines is called Veeam CDP.
From captured write to chosen timestamp
A nightly job can lose up to a day of transactions, which an order system or a busy file server may never recreate. CDP brings the recovery point objective (RPO) down to seconds by following each write through five stages:
- Capture. A component at the hypervisor, host or storage layer duplicates each write as it is issued.
- Transfer. The duplicated writes travel, in order, to a target system, usually at another site.
- Journaling. The target appends each write and its timestamp to a journal and applies it to a replica.
- Trimming. Entries older than the configured history drop out of the journal.
- Recovery. An operator picks a timestamp, and the replica is presented as it stood at that moment.
Because the replica usually waits at a second site, CDP shortens recovery time as well as data loss.
True CDP versus near-CDP intervals
- True CDP journals every individual write and can return to any instant in the history.
- Near-CDP takes frequent snapshots or replication cycles, every few minutes, and can return only to those points. A 5-minute interval gives 288 recovery points a day, against one for a nightly job.
- Replication without a journal carries corruption and ransomware encryption to the replica within seconds and keeps no earlier state to go back to. The journal is what turns replication into point-in-time recovery.
Veeam CDP for vSphere virtual machines
Veeam CDP is the always-on replication feature of Veeam Backup & Replication. It installs an I/O filter, built on VMware's APIs for I/O filtering, on the source and target vSphere clusters, so it sits in the write path of each protected disk and needs no VM snapshots. After an initial sync to an empty replica, every write travels compressed through Veeam CDP proxies to the target host.
Restore points come in two kinds. Short-term points are crash-consistent, can be seconds apart (Veeam describes 15 seconds or more as optimal) and are kept for up to seven days in transaction logs on the target datastore. Long-term points are built on a schedule, are application-consistent when application-aware processing is on, and are kept for a set number of days. Universal CDP extends the method to Windows and Linux machines, physical, virtual or cloud, replicating them to vSphere through an agent.
Everything Veeam CDP keeps lives on a vSphere datastore.
History that runs out before the damage is found
A journal sits on fast storage and absorbs every write to the protected systems, so it usually holds hours to a few days. A server writing a steady 20 MB/s produces about 1.7 TB of journal a day (20 × 86,400 = 1,728,000 MB), which makes weeks of journal history expensive very quickly.
That limit meets ransomware badly. CDP replicates encryption and corruption as faithfully as good writes, and attackers often wait days or weeks before acting, so the last clean moment may have dropped out of the journal by the time anyone looks. The journal and replica also share a management plane with the CDP software. With Veeam CDP, an intruder who reaches vCenter or the Veeam console can delete the source VM, the replica and its restore points in one session.
Bandwidth and target capacity cap how many machines join a CDP policy, so CDP tends to cover a short list of critical systems while backup jobs to separate storage cover the whole estate, those systems included. Copies on storage with enforced retention are what reach past the journal and survive a takeover of the CDP platform.
ARTESCA and continuous data protection
ARTESCA holds no journal or replica, since object storage is not a Veeam CDP target. Its role is the longer history behind CDP. The ARTESCA Zerto compatibility page describes offloading aged recovery points from journal storage to ARTESCA for long-term retention, with Zerto listed under ARTESCA Validated Design status.
Veeam backup jobs for CDP-protected VMs can write to ARTESCA buckets under S3 Object Lock, on storage managed apart from the vCenter that holds the replica. In compliance mode those restore points resist deletion by every account until their retain-until date, which covers corruption found after the journal has trimmed past it.
Related terms
- Point-in-time recovery: restoring data to its state at a chosen moment.
- Instant recovery: running a workload directly from backup storage while it is restored.
- Veeam Backup & Replication: the Veeam application that includes Veeam CDP.
- Ransomware recovery: restoring systems and data after a ransomware attack.
- Data backup: copies of data kept on separate storage so they can be restored.
