Cyber resilience

Backup security alerts: Which changes need attention?

Backup systems emit constant noise. Which configuration changes signal an attempt to destroy backups, and what to check.

7 min read
Dark data center aisle with server racks where one alert indicator glows amid rows of steady status lights

A backup environment generates hundreds of status events a day. Jobs finish, jobs retry, agents reconnect, repositories fill and drain. Buried in that stream are a handful of changes that look administrative and are not: a retention period cut in half, a second storage target added, a service account given a new key. Those changes are what an attacker makes before deleting anything, and they are the ones that rarely get read.

The usual response is to send everything the backup software emits into a ticket queue and assume the important items will stand out. They do not. Failed jobs are loud, frequent and almost always benign, so the team learns to skim, while configuration changes are quiet and rendered in the same typeface as a warning about a full disk. Tuning severity levels does not help, because the software cannot know which changes are legitimate.

A retention reduction made during a capacity squeeze and one made with stolen credentials produce an identical log line. What separates them is context: who made it, when, and what else happened in the same hour. That context comes from the team rather than the tool, so the list of events worth the effort has to be short.

The events that fill the alert feed

Most of what a backup system reports is operational. A job fails because a virtual machine was migrated mid backup. An agent stops checking in because a server was rebooted for patching. A repository crosses a capacity threshold because a file server was added to a job weeks ago and the sizing assumption was never revisited.

These events share a useful property: each correlates with something the team already knows about. Patch windows, migrations and storage growth all have a visible cause elsewhere in the estate, and a failure with an obvious explanation belongs in the operational queue and nowhere else. The events that matter are the ones with no corresponding activity anywhere.

Changes that shorten how long data survives

Retention reduction destroys backups without triggering a deletion alert, because from the software's point of view nothing is being deleted. A policy is edited, and the deletions follow later as cleanup. A retention period cut from ninety days to seven removes nothing at the moment it is saved. It removes months of restore points the next time the retention job runs, through the same code path that clears expired data every night.

In normal operation retention changes are rare, planned, and upward as often as downward. They accompany a compliance conversation or a capacity project, they are made by someone named in a change record, and they usually apply to one job rather than everything at once.

The suspicious version is broad, downward and unaccompanied. A single edit touching multiple jobs, or a change applied at the repository or bucket level rather than the job level, reaches far more data and deserves separate treatment for that reason alone. Where retention is enforced by the storage layer, the equivalent event is a change to a bucket default or a lock mode, which behaves differently again, as the interaction between retention and lock expiry shows.

Changes that move data somewhere new

A new storage target is the second pattern. Adding a repository is normal, since storage gets replaced, capacity gets expanded and copy jobs get introduced. It is also exactly what someone does when they want backups written to a location they control, or an existing job pointed away from immutable storage toward something erasable.

Normal target additions arrive with a purchase, a rack, and a matching change to the job that fills it. The version worth investigating has pieces missing: a target on a network segment the backup estate does not otherwise use, a target added without any job being modified, or a target that appears and then an existing job repointed at it in the same session. Repointing is the tell, since adding storage is additive and moving a job is not.

SignalWhat it usually meansWhat to check
Retention period reducedCapacity pressure or a policy revisionWhether the edit is job level or repository level, and how many restore points it retires
New storage target addedExpansion, refresh or a new copy jobWhether an existing job was repointed at it, and whether the target sits on a familiar network
Repository removedDecommissioning of replaced hardwareWhether the data behind it still exists, and whether the last copy of any chain lived there
Service account credential rotatedScheduled rotation or a password resetWhether rotation was due, which host performed it, and whether jobs continued under the same identity
Job disabled or schedule clearedWorkload retired or a maintenance windowWhether the protected workload still exists, and what else was disabled in the same session
Bucket policy or lock setting editedTuning during a storage projectWhich principal made the call, and whether it came through the backup software or around it

Reading a change in sequence rather than alone

Every event in that table is defensible in isolation. The sequence is not. Credentials change, a target is added, retention drops, jobs are disabled. Each step reads as routine administration. Together they describe someone removing the ability to restore before doing anything visible to users.

This is why timestamp proximity beats severity. Two configuration changes six weeks apart are almost certainly unrelated. The same two changes four minutes apart were made by one person in one sitting, and a sitting that touches credentials, targets and retention is not maintenance. Grouping configuration events by principal and session, even manually once a week, surfaces that pattern.

The other half of the question is where the change came from. A retention change made through the backup console and one made directly against the storage bucket look similar in outcome and very different in meaning, since the second goes around the backup software entirely. Telling them apart depends on being able to trust what the console reports, which is its own problem after an incident touches the backup console.

Where ARTESCA fits

ARTESCA is object storage software used as a backup target, deployed on infrastructure the customer runs. Because it sits beneath the backup software rather than inside it, changes made at the storage layer are recorded separately from changes made in the backup console. A retention setting altered in the backup application and a bucket level change made with S3 credentials are distinct events in distinct places, which is what makes it possible to tell them apart.

Immutability is enforced through S3 Object Lock, so a retention reduction in the backup software does not shorten a lock already applied to an object. The lock travels with the object rather than with the job that wrote it. The set of changes capable of reducing what survives is therefore narrower at the storage layer, and the credentials and bucket policies governing it are a smaller surface to review. That surface is specific rather than general, which is the substance of what Object Lock actually guarantees.

Physical location, network boundaries and operational access remain the customer's, so who holds the storage side credentials, and how they differ from production credentials, is decided locally rather than inherited.

What to check on a schedule

A weekly review of configuration changes is enough for most mid market environments, and it is short if the scope is right. Pull the configuration change log from the backup application and the access log from the storage account, filter to the six event types above, and look at nothing else. Job failures have their own queue and do not belong here.

For each change, write down three things: who made it, what business activity explains it, and whether that explanation was known before the log was read. The third item does the work. A change the team can explain only after seeing it in a log is a change nobody authorized, even when the explanation turns out to be true.

Keep a written baseline of the settings that should not drift: retention per job, the list of storage targets, and which accounts hold storage credentials. Compare live configuration against it quarterly. Drift is often the first sign that permissions are broader than intended, which is where narrowing what the backup service account can reach starts.

Try ARTESCA free

Immutable object storage that scales from 20TB to petabytes. Deploy a working cluster in under an hour.

Start a free test drive