Integrations

Backup storage certificates: Avoid connection failures

Which certificate backup software validates, which hosts must trust it, how expiry surfaces, and what recovery needs offline.

7 min read
Dark data center with a teal padlock-shaped light structure forming across a link between two racks

A backup job that has run cleanly for a year fails one night with a connection error, and the first assumption in the room is that the storage is down. The storage is fine. A certificate on the path expired, or was replaced, or was never trusted by the host that started complaining. Certificate problems are a common cause of sudden backup failures and among the most consistently misdiagnosed.

Part of the reason is that backup software reports the failure poorly. The message describes the symptom, an endpoint that could not be reached, rather than the cause. The other part is timing. Certificates expire on their own schedule, unrelated to any change the team made, so the incident begins with nobody having touched anything.

The work that prevents this is not complicated, but it has to be deliberate, since none of it is visible in the backup console. It comes down to knowing which certificate is validated, by which hosts, against which trust store, and holding the material needed to rebuild that trust when the systems that normally provide it are unavailable.

Which certificate the backup software actually validates

The backup server validates the certificate presented by whatever answers the endpoint name in the repository configuration. That is not always the storage. If a load balancer or reverse proxy terminates TLS in front of the nodes, its certificate is the one being checked, and the certificates on the nodes behind it are irrelevant to the client. Teams have renewed the wrong one more than once.

The name matters as much as the dates. The certificate must cover the exact string configured in the repository, and S3 endpoints complicate this through addressing style. With virtual hosted addressing the bucket name becomes part of the hostname, so the certificate needs a wildcard entry covering names one level below the endpoint. With path style addressing the endpoint name alone is enough, so a certificate tested by hand can still fail once the product switches style.

If the repository is configured by IP address rather than name, that address has to appear in the certificate as an IP entry. Internal certificate requests often include names only, producing an endpoint that validates in a browser and fails from the backup server.

Trust stores are per host, not per organization

An internal certificate authority is trusted only by machines that have its root certificate installed. That is a per host condition, and the set of hosts touching backup storage is larger than it first appears: the backup server, every gateway or proxy that moves data on its behalf, every agent that writes to the repository directly, and any management or monitoring host that polls the endpoint.

The store also has to be the right one. On Windows, a certificate installed into a user store does nothing for a service running under a different account, and the machine store is what matters. On Linux, the system bundle and an application's own bundle can differ. An administrator testing from a command line can succeed while the service account fails, which produces a long argument about whether the certificate is installed.

Intermediates are the other half. A chain is valid only if the client can build a path from the presented certificate to a trusted root, and most endpoints are expected to present their intermediates alongside the leaf. An endpoint configured with the leaf alone works from hosts that cached the intermediate and fails from those that did not, which looks like an inconsistent fault rather than a configuration error. It is sharper where backup infrastructure is deliberately kept off the production domain, since those hosts never receive the root through the usual distribution mechanism.

How expiry surfaces, and why it is misread

Expiry rarely announces itself. The job fails, the console reports an error about the repository, and the storage dashboard shows every node healthy, which sends the investigation toward the network. The certificate gets checked late, because it is not part of the storage view.

A second pattern is worse. Some tasks tolerate a broken session and others do not, so the first visible effect can be a failing health check, a retention pass that does not run, or a lagging copy job, while the main backup still completes on a session established earlier. The failure then looks gradual and unrelated to certificates.

SymptomWhat it usually isWhat to check
Job fails at once with a trust error, nodes healthyExpired or untrusted certificate on the pathThe certificate the endpoint presents, and its dates
Works from one host, fails from anotherRoot certificate missing from that host's trust storeThe machine store, not the logged in user's store
Failures began after a network device changeTLS terminated elsewhere, presenting a different certificateWhich device answers the endpoint name today
Connection hangs, then times outRevocation checking cannot reach its distribution pointWhether the host can reach the listed responder address
Error names the host rather than the datesThe name in use is not covered by the certificateEntries against the exact endpoint string, wildcard included
Trust fails only in an isolated recovery networkThe chain is unavailable without the production authorityWhether root and intermediates are stored offline

Renewing without taking the repository down

Renewal is usually less disruptive than it is treated. Most object storage endpoints reload a certificate by restarting a listener, which affects sessions in flight but not stored data, and a node at a time keeps the endpoint answering. The requirement is that the work misses the middle of a job, the same constraint that governs any planned maintenance on the repository.

The sequencing that avoids an outage is to distribute trust before it is needed. When an authority is replaced, the new root can be installed on every client while the old one is still in place, so both are trusted during the overlap and the endpoint certificate can be swapped without a synchronized change. Short lived certificates renew often enough that the process has to be automated, and what needs verifying there is that the automation reloads the endpoint rather than only writing a file.

Clock accuracy belongs in the same conversation. A certificate is rejected as not yet valid on a host whose clock is behind, and an isolated backup network is where time synchronization quietly stops working.

Where ARTESCA fits

ARTESCA presents an S3 compatible endpoint over TLS, and like any such endpoint it can be configured with a certificate from a public authority or from the organization's own. Because it is deployed on infrastructure the customer runs, the certificate, the naming and the trust distribution are all under the customer's control rather than a provider's.

That control is the practical advantage during recovery. The endpoint name, the addressing style and the certificate are all known locally, so the chain needed to reach immutable backup storage can be captured in advance and kept with the recovery material rather than requested from an external service at the worst moment.

What to record and when to check it

A short record kept with the repository documentation prevents most of this. It names the endpoint string exactly as configured, which device terminates TLS, the issuing authority, the expiry date, every host holding the root, and where renewal automation runs. That record turns a two in the morning outage into a five minute fix, so it belongs in the runbook the on call person opens.

Expiry deserves a calendar entry well ahead of the date, not an alert on the day. Sixty days is a reasonable lead for a certificate renewed by hand, since it covers a request, an approval and a change window without urgency. The same check should note the expiry of the root distributed to clients, because a root outliving everyone's attention is how an estate loses trust at once.

The recovery copy is the item most often missing. The public part of the root and any intermediates belong with the offline recovery material, along with the endpoint certificate and a note on whether revocation checking has to be relaxed when the authority cannot be reached. Confirming a rebuilt host reaches storage without the production authority is worth a line in the next drill run with production down.

Try ARTESCA free

Immutable object storage that scales from 20TB to petabytes. Deploy a working cluster in under an hour.

Start a free test drive