The database server is backed up every night with the rest of the estate. The job is application aware, it reports success, and the restore point is a complete image of the machine. Then someone runs an update without a where clause at eleven in the morning, and the only recovery on offer is the whole server as it stood at two. The backup was not wrong. It was the wrong granularity.
Image level backup protects a server. Database incidents are rarely about a server. They are about one table, one database inside an instance holding a dozen, or a specific minute before a change went wrong. An image restore answers none of those precisely, because its unit is the machine and its resolution is the interval between jobs.
What an image level restore actually returns
An image restore returns the entire virtual machine as it existed at the snapshot: operating system, instance configuration, every database in that instance, agent jobs and logins. All of it moves together. There is no way to take one database back four hours while leaving the others alone, because the restore is not aware of databases as separate things.
Resolution is the second limit. With a nightly job, the available points are twenty four hours apart. An incident at eleven in the morning is ten hours past the nearest one, and those hours are ordinary business activity. Losing them to recover from a single bad statement is usually worse than the statement.
Consistency is the third. A quiesced image gives a database that opens cleanly. A crash consistent image gives one that performs its own recovery on start-up, rolling back transactions in flight. That normally works, but it is a different guarantee, and it is worth knowing which kind the job produces rather than assuming from its name.
Transaction logs and the hours between images
Every serious database engine keeps a log of changes, and that log is what makes recovery to an arbitrary moment possible. The names differ: transaction log backups, archived redo logs, write ahead log archiving, binary logs. The mechanism is the same. Restore a consistent copy, then replay logs forward to a chosen time.
Point in time recovery depends on an unbroken chain from the data copy to the target moment. Any missing log, any gap where archiving was off or the destination was full, and the replay stops there. That is why a log backup failure is more serious than a data backup failure that will simply run again tomorrow.
Log backup frequency, not image frequency, sets the real recovery point for a database. Logs captured every fifteen minutes mean fifteen minutes of exposure regardless of when the nightly image ran. That one setting usually does more for database recovery than anything done on the infrastructure side, and it is often left at the installation default.
What application aware processing does and does not handle
Application aware processing quiesces the database before the snapshot, through the operating system writer framework on Windows or pre and post scripts elsewhere, so the image holds a consistent copy rather than files caught mid-write. Most implementations can also truncate transaction logs after a successful backup, which keeps the log volume from filling.
Truncation is where the trouble starts. Truncating logs as part of the image job removes the record of everything since the previous one. If no separate log backup exists, point in time recovery between images is gone, and it goes quietly, since the job that removed the capability is the one reporting success. Many products offer a log backup feature for this reason, and it has to be enabled deliberately.
Other limits are worth confirming in the deployment at hand, since behavior varies by product and version. Databases on storage attached inside the guest may sit outside the image entirely. Clustered instances and availability groups have their own rules about which node is backed up. And an image job that truncates logs can interfere with the database team's schedule if nobody compared the two.
| Incident | What an image restore alone gives back | What is needed as well |
|---|---|---|
| Bad statement at midday | The whole server as of last night, losing the morning | Log backups and a point in time restore, or an export from a mounted copy |
| One database corrupt, the others fine | Every database in the instance rolled back together | A single database restore from a native backup or item level extraction |
| Server lost to hardware failure | A working server as of the last image | Log replay to close the gap since that image |
| Schema change went wrong | Instance state before the change, if an image sits between | A backup taken by the change process itself, immediately before |
| Log volume filled, instance stopped | A server that boots into the same condition | Log backup schedule, retention and free space monitoring there |
| Database files on guest attached storage | Operating system disks, possibly without the data | Which volumes the image job actually includes |
Restoring one database, and where the database team's own backups fit
Getting a single database back without disturbing the rest takes one of two routes. The first is an item level restore from the image, where the backup product mounts the restore point and presents database objects for publishing or extraction. The second is a native restore from a backup produced by the engine itself, which is the route the database team already knows.
A common and workable arrangement runs both. The database team takes native backups and log backups on its own schedule to a location outside the database server, and the infrastructure team captures that location as ordinary files alongside the image backups. The database team keeps familiar tooling and fine grained control of time, while the infrastructure team provides the offsite copy, the retention and the immutability. It is also the pattern that makes file level restores genuinely useful, since what is recovered are files with predictable names.
Two failure modes recur. Native backups written to a local disk inside the same virtual machine die with the machine. And a dump schedule that overwrites the same file each night gives exactly one recovery point, which is not what anyone assumes on hearing that the database team has its own backups.
Where ARTESCA fits
ARTESCA is object storage software used as a backup target. In a database context it usually holds two kinds of data: image backups of the database servers written by the backup software, and native database and log backups, either written to it directly by tools that speak S3 or captured from a file location.
Because it presents an S3 compatible API, one repository can serve both paths, which keeps the offsite copy and the retention policy consistent regardless of which team produced the data. Immutability is available through S3 Object Lock, applied per object for a defined period, so a native backup and a backup chain written by the backup software are protected by the same mechanism.
What object storage does not change is granularity. Recovering one database to one minute still depends on the log chain and the tooling that replays it. The storage layer decides where copies live and how long they survive, not how precisely they can be rewound.
What to establish before the next database incident
For each database instance, write down three things: the recovery point it can actually meet, which is set by log backup frequency rather than by the nightly job; the procedure for restoring a single database without taking the instance back; and who runs it. Most organizations find these differ from what the service catalog claims.
Confirm truncation behavior once, explicitly: which job truncates logs, whether a separate log backup exists, and where those log backups are written. A short test proves it. Take an image backup, make a change, then attempt recovery to a moment between the two and see whether the tooling can reach it.
Include a database in the restore tests that use representative workloads, not only a file server, and record how long a single database recovery took end to end. The interval between such tests matters as much as the tests, which is the argument that applies to recovery testing generally. A procedure written once and never run is an assumption with formatting.
