Files
hermes-skills/devops/proxmox-ve-administration/references/pbs-offsite-backup-jobs-2026-07.md
T

4.1 KiB

PBS Offsite Backup Jobs: Adding Guests & Running on Correct Nodes

When to Use

When expanding offsite PBS backup coverage to include more CTs/VMs beyond the initial HA-only job, considering bandwidth constraints of the remote site.

Key Lessons

1. vzdump Must Run on the Node Hosting the CT

Critical pitfall: vzdump <vmid> --storage noris_offsite run on proxmox1 (10.0.20.10) will silently exit 0 for any CT hosted on a different node — no backup is created, no error is logged. The CT config file (/etc/pve/lxc/<vmid>.conf) only exists on the node where the CT runs.

Finding which node hosts a CT:

ssh root@10.0.20.10 "pvesh get /cluster/resources --type vm --output-format json" | \
  python3 -c "
import sys,json
data=json.load(sys.stdin)
for v in data:
    if v.get('vmid') == TARGET_VMid:
        print(f\"node={v['node']} status={v['status']}\")
"

Running backup on the correct node:

# CT108 on proxmox6 (10.0.20.60)
ssh root@10.0.20.60 "vzdump 108 --storage noris_offsite --mode snapshot --compress zstd"

# CT99999 on proxmox7 (10.0.20.70)
ssh root@10.0.20.70 "vzdump 99999 --storage noris_offsite --mode snapshot --compress zstd"

# CT104 on n5pro (10.0.20.91)
ssh root@10.0.20.91 "vzdump 104 --storage noris_offsite --mode snapshot --compress zstd"

2. Updating the Offsite Backup Job

The offsite job config lives in /etc/pve/jobs.cfg on any PVE node (shared filesystem). Edit with sed:

ssh root@10.0.20.10 "
sed -i '/offsite-ha-daily/,/^$/{s/vmid 106/vmid 106,108,104,99999/}' /etc/pve/jobs.cfg
"

Job config after update:

vzdump: offsite-ha-daily
	compress zstd
	enabled 1
	mode snapshot
	schedule 02:00
	storage noris_offsite
	vmid 106,108,104,99999

3. Bandwidth-Conscious Guest Selection

The remote PBS at noris (213.95.54.60) has limited upload bandwidth. PBS dedup means only changed chunks are transferred after the initial full backup, but the first backup of each guest is a full upload.

Selection criteria for offsite:

Criterion Include Exclude
Size < 20 GB Small enough for initial full
Size > 100 GB Too large for WAN
Change rate low (config, code) Daily delta is tiny
Change rate high (video, media) Daily delta too large
Infrastructure-critical Traefik, Gitea, HA
Replaceable data Frigate recordings

This homelab's offsite selection (2026-07-08):

  • VM106 HA (70 GB) — already in job, critical
  • CT108 Gitea (1.2 GB) — added, critical code
  • CT99999 Traefik (3 GB) — added, critical infra
  • CT104 Paperless (16 GB) — added, documents

Excluded from offsite:

  • CT111 Seafile (558 GB) — too large, needs separate strategy
  • CT120 Frigate (40 GB) — replaceable video data
  • CT115 Immich (20 GB) — photos, needs separate strategy
  • CT121/124/127 — low priority, small

4. Verifying Offsite Backups

After running manual backups, verify on the remote PBS:

ssh root@10.0.20.10 "pvesh get /nodes/proxmox1/storage/noris_offsite/content --output-format json" | \
  python3 -c "
import sys,json,datetime
data=json.load(sys.stdin)
for b in sorted(data, key=lambda x: x.get('ctime',0), reverse=True)[:10]:
    volid=b.get('volid','')
    sz=b.get('size',0)//1024//1024
    dt=datetime.datetime.fromtimestamp(b.get('ctime',0)).strftime('%Y-%m-%d %H:%M')
    print(f'{volid:55s} {sz:>8} MB  {dt}')
"

Note: PBS content listing may lag by a few minutes after upload. If backups don't appear immediately, wait and re-query.

5. Lock Timeout on Concurrent vzdump

If two vzdump jobs run simultaneously on different nodes targeting the same PBS storage, one may fail with:

ERROR: can't acquire lock '/var/run/vzdump.lock' - got timeout

This is a global lock across the PVE cluster. Retry the failed backup after the first one completes. For manual triggers, run sequentially.

Session Reference

  • 2026-07-08: Added CT108 (Gitea), CT99999 (Traefik), CT104 (Paperless) to offsite-ha-daily job. Initial full backups run manually on correct nodes. CT104 failed with lock timeout (concurrent with scheduled job), retried successfully.