The obvious way to restore my Kubernetes cluster brings back an empty database that reports healthy. I first thought it would also overwrite the backup. It would not, and the reason changed how I plan for it.
The obvious way to restore my Kubernetes cluster brings back an empty database that reports healthy. I first thought it would also overwrite the backup. It would not, and the reason changed how I plan for it.
In August 2026 I wrote the recovery plan for the small Kubernetes cluster that runs some of my products. It answers a boring question: if the machine dies, what do I type, and in what order? One step looked harmless and turned out to be the most dangerous one in the document.
The cluster's state is backed up as a snapshot, and restoring it brings everything back. That includes the operator that runs Postgres, a small program that watches the database definition and makes reality match it.
I read the live definitions on 29 August. Two details lined up badly:
On new hardware, the disk under that database is empty.
In the obvious order, the operator starts before you can do anything. It finds an empty disk and an instruction to create a database, so it creates one and reports healthy. Then the rest of the plan points every app at it.
My plan said the new database would then archive over the real backup. A reviewer agent reading a draft of this post asked whether the operator allows that. It does not. CloudNativePG checks that the archive folder is empty before a new cluster writes to it, and refuses if it is not, unless you set an annotation that tells it to skip the check. The backup survives.
The empty database does not go away, though. An open report on the project describes a cluster that logged the failed check and still came up healthy, with archiving silently stopped. So that is the case I plan for: an empty database, green checks, apps pointed at it, and no new backups being taken.
The fix is the same either way. Starting k3s without a scheduler is a documented option in the version I run; I checked its help text. Kubernetes creates the pods, but nothing can place them on a node, so nothing runs. The API still works, which is all you need to stop the operator, save the database definition, and delete it. Then the scheduler goes back on and the backup is restored into a new cluster with a new name.
Before anything points at it, I check it. A restored Postgres comes back on timeline 2, while a freshly created one starts on timeline 1. Then the row counts have to match.
The two details above were read from the live cluster on 29 August. Since then, nightly backups and a second, separate tier of per-database dumps have been added. The operator's empty-archive check comes from its documentation, not from a run of mine.
I have not driven a full restore this way on throwaway hardware. Rehearsing it is on the list. Until then this is a plan with its reasoning written down.