A backup I have never read back is a hope. This is how mine get read back: what the checks count, and the one email that tells me the checker itself is still alive.
A backup I have never read back is a hope. This is how mine get read back: what the checks count, and the one email that tells me the checker itself is still alive.
As of September 2026, eleven databases live on one Postgres cluster in my Kubernetes setup. The cluster already had a backup: CloudNativePG ships every change to object storage, plus a nightly base backup, kept for 30 days. It can put the whole cluster back to any minute in that month. I had restored it once in a drill, and every row count matched. But a drill is a one-off, and there were two smaller jobs it cannot do.
The cluster backup is the right tool when the machine is gone. It is the wrong one for two smaller jobs. One is putting back a single app's database without touching the other ten. The other is going back further than 30 days.
So this month I added a second layer next to it: one dump file per database, every night, with its own checks.
At 03:30 UTC a Kubernetes job finds every database on the cluster by itself and dumps each one to a compressed file. All eleven come to about 32 MB a night. Each file lands in Cloudflare R2 in one of two tiers:
The monthly copies sit under a bucket lock. No key the backup system holds can delete or overwrite one, the job's own included. Only the account's admin token can lift the lock. On 27 September, an overwrite and a delete with the job's own key were both refused with 409 ObjectLockedByBucketPolicy. A script bug cannot erase the history.
A green job only says the job exited 0. So the checks never ask the job. They list what is actually in storage.
Freshness, every morning. For each database, status.sh needs three things:
It also reads the newest run log, which must be recent and say failed=0. That catches a database the job could not dump at all. It never got a copy, so it would never show up as stale. Only when all of this holds does the script print BACKUP_FRESH=1.
Two details matter here. If the check itself cannot run, because the network is down or storage refuses the key, it still ends BACKUP_FRESH=0. And running it with a maximum age of zero hours must print BACKUP_FRESH=0, which proves it can fail.
Restore, by hand. restore-check.sh downloads one dump and loads it into a throwaway Postgres container with no network. It lists the schemas, the table count and the largest tables. It prints RESTORE_OK=1 only if the restore reported no errors and tables came back. Then it deletes everything.
On 27 September all eleven databases restored with zero errors. Two of them were also counted against the live database:
The other database compared came back with 71 of 71 tables. One of the eleven uses pgvector, and a plain Postgres image fails that restore with extension "vector" is not available. So the check defaults to Postgres 18 with pgvector, to match the cluster. My secrets manager has its own version of this test: a throwaway copy boots, and all 60 test keys decrypt, from both a daily and a monthly copy.
At 06:15 UTC a small script runs the freshness checks and decides whether to email me.
| Day | Every check passes | Any check fails |
|---|---|---|
| Monday, or the 1st of the month | "Backups OK", with the full report | "BACKUPS FAILING", and which one |
| Any other day | No email | "BACKUPS FAILING", and which one |
An alert that only fires on failure also goes quiet when the alerting breaks.
If a Monday passes without one, silence on the other days stops meaning anything. The 1st gets a report because that is the night the monthly copies are written.
At 06:15 the night's dumps are about three hours old. So the watcher allows 6 hours, not the default 26, and a night that did not run is reported that same morning.
All three outcomes were tested on 27 September. A forced report reached my inbox, not spam, one second later. A forced failure arrived with the failing backup named in the subject. That day was a Sunday with everything green, and the normal run sent nothing.
Since this month I count a backup as done only when a check has read it back and counted it, and I have watched that check fail once.
Backups: never report success from a job's exit code or a green dashboard. Read what is in storage. Freshness: every database needs a copy newer than the last scheduled run, above a size floor, and a run log that says failed=0. Restore: load the copy into a throwaway database with no network. Compare table and row counts with live at least once. A check that cannot run reports failure, never success. Before you trust a check, run it once with a threshold that must fail, and watch it fail.