What is data backup?
Data backup is a copy of files, systems or databases held apart from the originals, on separate storage, for restoring after loss, corruption or attack. Each copy taken at a point in time is a restore point. NIST defines a backup as "a copy of files and programs made to facilitate recovery if necessary."
From job selection to expiry
- Selection. A job names the virtual machines, servers, databases or applications it protects.
- Consistent capture. A snapshot freezes the source briefly so the copy reflects one moment; an application-consistent capture first asks databases and similar software to flush pending writes.
- Copying. The first run copies everything, and later runs copy only what changed.
- Storing. Data is usually reduced, compressed and encrypted, then written to the backup target.
- Cataloging. The software notes what each restore point contains and where it lives.
- Expiry. Points older than the retention policy allows are deleted to free space.
Backup types differ in what each run copies. A full backup takes everything, an incremental backup takes changes since the previous run, a differential backup takes changes since the last full, and continuous data protection records each write as it happens.
RPO, RTO and how far back restore points reach
Three measures describe whether backups do their job. The recovery point objective (RPO) is how much recent data the organization can afford to lose; with one job a night, up to a day of changes is exposed. The recovery time objective (RTO) is how long systems can stay down, and it depends on restore throughput far more than on backup speed. Retention depth is how far back restore points go, which decides whether an infection noticed three weeks late can still be rolled back past.
Capacity follows that third measure. Every additional month of restore points adds storage, and the data reduction in the backup software decides how much.
The three measures pull against each other inside one budget. A shorter RPO means more frequent jobs and more restore points to hold, deeper retention multiplies capacity, and a tighter RTO calls for storage that reads many systems back at once. Most mid-sized organizations settle them per tier of systems rather than once for the whole estate.
Copies that survive a deliberate attack
Failed disks and files deleted by mistake still happen, but ransomware changed what backups are up against. An attacker who encrypts production also tries to delete or encrypt the backups, because a working copy is what lets an organization refuse to pay.
In mid-sized estates the weak points are rarely the jobs themselves. They are the backup server joined to the production domain, a repository reachable with the same administrator credentials, retention any administrator can shorten, and restores never run at full scale. Each one turns a green job report into a copy that may be gone, or may come back too slowly, on the day it is needed.
Placement matters as much as count. The 3-2-1 rule describes three copies on two different media with one off site, and the 3-2-1-1-0 rule adds one immutable or offline copy and zero errors in restore checks.
ARTESCA and data backup
Backup applications write restore points to ARTESCA over the S3 API, with compatibility pages published for Commvault, Cohesity, HYCU, Rubrik, Veeam, Veritas NetBackup and Zerto. S3 Object Lock fixes the earliest date each restore point can be removed: in compliance mode no account, including root, can delete a point before that date, and in governance mode an identity with the bypass permission still can. Per-bucket S3 Lifecycle rules expire the objects once their date has passed.
ARTESCA accounts are separate from the production domain, which closes the shared-credential weak point described above. Restores that have never been run at scale remain a gap only testing exposes.
Related terms
- Backup target: the storage system a backup application writes restore points to.
- 3-2-1-1-0 backup rule: three copies, two media, one off site, one immutable, zero restore errors.
- Retention policy design: setting how many restore points exist and how far back they go.
- Immutable backup: a backup protected from change or deletion until its retention date.
- Restore throughput: the speed at which data comes back from backup.
