Notes

Proxmox backup retention that won’t surprise you

Backups are one of those things everyone agrees are important and almost nobody enjoys thinking about. The part that gets the least attention is retention: how many backups you keep and for how long. Get it wrong one way and your datastore quietly fills up until jobs start failing. Get it wrong the other way and the one backup you need was pruned last Tuesday.

This note covers how I think about retention for Proxmox VE and Proxmox Backup Server (PBS), and the settings I’ve landed on for a home lab.

Start with the question, not the setting

Before touching any configuration, answer two questions for each class of workload:

  • How much data can I afford to lose? That sets how often you back up.
  • How long might it take me to notice a problem? That sets how far back you need to go.

The second question is the one people skip. Ransomware, silent corruption or a bad config change might not be noticed for days or weeks. If you only keep three days of backups and you notice on day five, every copy you have is already bad.

How Proxmox prunes

Proxmox uses a “keep” model that’s easy to reason about once it clicks. You set how many backups to keep in each time bucket:

  • keep-last: the most recent N backups, regardless of age
  • keep-daily: one backup per day for the last N days
  • keep-weekly, keep-monthly, keep-yearly: the same idea at coarser granularity

The buckets combine. A single backup can satisfy more than one rule. Anything that doesn’t match any rule is pruned. You can configure this on the backup job in Proxmox VE or, better, on the PBS datastore or a prune job, so the policy lives next to the data.

What I use

For most lab VMs and containers, backed up nightly:

keep-last: 3
keep-daily: 7
keep-weekly: 4
keep-monthly: 3

That gives me every recent backup for quick rollbacks, a week of dailies, a month of weeklies and a quarter of monthlies. Thanks to PBS deduplication, the storage cost of the older points is far smaller than you’d expect, because most blocks are shared between snapshots.

For throwaway test VMs, I use a much shorter policy, often just keep-last: 2. Not everything deserves three months of history, and treating every guest the same wastes space that the important systems could use.

Pruning isn’t the same as freeing space

This catches a lot of people. On PBS, pruning removes the index of old backups, but the underlying chunks stay on disk until garbage collection runs. If you prune and wonder why the datastore is still full, that’s why.

Schedule garbage collection to run regularly, after your prune jobs. Also be aware that chunks are only eligible for removal after a grace period of roughly a day, so space comes back gradually rather than instantly.

Verify, then actually restore

PBS can run verification jobs that re-read backup chunks and check them against their checksums. Turn them on. They catch bit rot and storage problems long before you’re in a hurry.

Verification doesn’t prove that a restore will boot, though. Every so often I restore a VM to an isolated network, start it, and confirm the application works. It takes twenty minutes and it’s the only real proof that the backups are worth what they cost.

Snapshots are not backups

ZFS or LVM-thin snapshots are wonderful for “undo the last hour” recovery, and I take them frequently with a short retention, typically a day or two of hourlies. But they live on the same storage as the data. If the pool dies, the snapshots go with it. They complement backups; they don’t replace them.

Don’t forget off-site

A backup server in the same rack as the hypervisor protects against a lot, but not against fire, theft or a power event that takes out both. PBS sync jobs make it straightforward to pull a copy to a second datastore somewhere else. Even a smaller retention policy off-site, such as dailies for a week and monthlies for a few months, covers the disaster cases.

A quick checklist

  1. Decide retention per workload class, based on how long problems might go unnoticed.
  2. Configure prune rules on the datastore, not scattered across jobs.
  3. Schedule garbage collection after pruning.
  4. Turn on verification jobs.
  5. Alert on failed jobs and on datastore usage above about 80%.
  6. Restore something on a schedule, and write down how long it took.

Retention isn’t glamorous, but a little thought up front turns backups from a vague comfort into something you can actually rely on.

Questions or corrections? Email me.

← Lessons from DNS caching and NXDOMAINRootless Podman with Quadlet: the one tip that matters →