Backup storage monitoring tends to arrive in one of two states. Either nothing is sent anywhere, and problems are found the morning after a job failed, or every event the storage system can emit is forwarded to an operations channel, where it joins thousands of others and stops being read within a month. Both states produce the same outcome, which is that a real fault is noticed late.
The difference between a useful alert and noise is not severity as the storage system labels it. It is whether the person who receives the alert at two in the morning can do something that changes the outcome. A disk that failed in a system with redundancy intact is important and is not urgent. An endpoint that stopped answering while jobs are running is both.
Sorting events along that line is most of the work. What remains is choosing thresholds that fire early enough to act on and rarely enough to stay credible, and making sure a handful of symptoms that look like backup software problems are routed to whoever owns the storage.
An operations center at two in the morning has a narrow set of actions available. It can check whether a host or endpoint answers, restart a service on an approved list, disable a credential, stop or reschedule a job, and escalate to a named person. It cannot size an expansion, judge whether a rebuild is progressing acceptably, or decide that a retention policy should change.
That list is the test for every alert being considered. If the only correct response is to note it and tell somebody in the morning, the alert belongs in a morning report, not on a phone. Sending it to the phone anyway does more harm than leaving it out, since it teaches the recipient that the channel does not require attention.
The reverse error is quieter. Events that do warrant a call are sometimes suppressed because an earlier round of tuning removed a whole category. Reviewing what was suppressed, and why, is worth an hour after any cut in alert volume.
A fixed percentage threshold on a backup repository produces alerts at the wrong times. Backup storage fills in steps, since a synthetic full, a monthly restore point or a new workload arrives as a block rather than a trickle, and it also releases space in steps when retention expires. A repository can sit at a high percentage for months quite safely and can go from comfortable to full in two nights.
The number worth alerting on is the point at which tonight's scheduled writes will not fit. That is derived from recent daily change, the largest single job in the schedule, and space expected to be released by expiry before the window opens. Crossing it is actionable now, since the response is to free space, move a job, or reduce tonight's scope, and those are decisions a night shift can carry out with a short instruction.
The slower question, how many weeks remain at the current rate, is a morning matter. It leads to a procurement conversation rather than an action, and belongs in a weekly view alongside the reasoning about how much headroom a repository should carry. Sending a trajectory warning to an operations channel invites the recipient to do nothing and learn that doing nothing was correct, which is how a channel becomes background noise. Separating the two thresholds is the substance of capacity alerting that stays credible.
Reachability is the most valuable single check, and it is best made from the backup server rather than a monitoring host, because that is the path that matters. An endpoint answering a probe on a management network while the backup server cannot reach it is a failure a well intentioned check reports as healthy.
The check should exercise more than a port. Establishing a session and listing a bucket confirms the name resolves, the certificate validates and credentials work, which is three common failure modes in one test. Running it a few minutes before each backup window, and continuously during it, gives a responder time to act before the jobs start rather than after they fail.
Authentication failures deserve their own route. A handful on a wrong key is an ordinary mistake. A rising count against a valid account, or attempts from an address that is not a known backup component, is a security signal and should reach a person immediately, with enough context to identify the identity and the source. The action available at two in the morning is narrow but real, since disabling a key stops the attempt while leaving the investigation for daylight.
| Signal | Route | What a responder can do at once |
|---|---|---|
| Endpoint unreachable from the backup server | Call | Check the path, restart an approved service, escalate |
| Authentication failures rising, or from a new source | Call | Disable the credential, record the source address |
| Redundancy lost after a second fault in one group | Call | Reduce load, confirm the rebuild is progressing |
| Tonight's scheduled writes will not fit | Call | Free space, reschedule or reduce the largest job |
| Single drive failed, redundancy intact | Morning | Log it, order the part, confirm the rebuild started |
| Capacity trend reaching full in weeks | Morning | Plan expansion, review retention and growth |
Hardware alerts are where volume usually comes from, and the useful line is redundancy rather than component count. A failed drive in an erasure coded or replicated system is a maintenance item. A second failure in the same protection group, a node lost while another is already out, or a rebuild that has stopped progressing are different in kind, because the next fault has consequences. Alerting on redundancy state rather than individual components cuts volume while raising the value of what remains.
Power supplies, fans and network paths follow the same logic. One path down with a second carrying traffic is a morning item. Both down is a call. The condition worth sending is the loss of the spare, not the loss of a part.
Several job side symptoms are storage symptoms wearing a different label. Jobs that succeed but take much longer than usual, a copy job falling behind, a health check failing, or retention tasks that queue without completing all point at the repository rather than at the backup software. They are often the first visible sign of a fault or of maintenance work still settling. A job marked successful is a weaker signal than it appears, which is the practical reason job success and recovery success are tracked separately.
ARTESCA runs on infrastructure the customer operates, so its health, capacity and authentication events are available locally rather than through a provider's interface, and can be forwarded to whatever monitoring the organization already runs rather than requiring a separate console to watch.
Because it is used as a backup target with S3 Object Lock, two categories deserve explicit routing: the endpoint check made from the backup server, which covers name, certificate and credentials together, and the state of retention and expiry processing, since a repository that accepts writes while expiry has stalled will fill without any obvious fault.
A workable starting point is short. Two capacity thresholds, one for tonight and one for the quarter, routed differently. One endpoint check from the backup server, before and during each window. One authentication rule. One redundancy rule per protection domain. Everything else goes to a daily digest somebody reads with a coffee.
Each alert that reaches a phone needs a written response, even a single line, naming what to check and who to escalate to. An alert without one is an interruption rather than a task, and it is the reason night staff stop treating a channel as real.
The review matters more than the initial configuration. Once a quarter, count how many calls were made, how many led to an action, and which faults were found some other way. Alerts that never fired and alerts that fired without producing an action both need changing, and the record of that review keeps the threshold conversation from restarting every year.