Files

16 KiB

RKE2 Upgrade Attempt & Tarball Recovery — 2026-07-12

Goal

Update K8s cluster: OS updates on all VMs, RKE2 upgrade v1.35.2→v1.35.6, Helm chart updates, fix broken components. Scope: K8s only (no PVE hosts).

Phase 1: OS Updates (COMPLETED)

All 7 VMs updated via qm guest exec through Proxmox jumphost chain.

VM→PVE-Node Discovery

Initial assumption that all VMs were on proxmox3 was WRONG. Had to scan all PVE nodes to find actual placement:

118 (CP-01)    → proxmox5    (NOT proxmox3)
130 (CP-02)    → proxmox3
129 (CP-03)    → proxmox2
128 (Worker-01) → proxmox5   (later moved to proxmox)
132 (Worker-02) → proxmox2
131 (Worker-03) → proxmox3
200 (mgmt-runner) → proxmox3

Update Process

  • VM200: Direct SSH (10.0.30.124) — 0 updates remaining
  • VM118, VM129: Updated via qm guest exec on respective PVE nodes
  • VM132: dpkg interrupted (timeout) — fixed with dpkg --configure -a
  • All VMs: 0 remaining updates after completion
  • Cluster stayed healthy throughout (6/6 nodes Ready)

Command Pattern

ssh -i ~/.ssh/id_ed25519_proxmox root@10.0.20.10 \
  "ssh -o ConnectTimeout=5 root@PVE_NODE 'qm guest exec VMID --timeout 600 \
   -- sh -c \"apt-get update && apt-get upgrade -y\"'"

Phase 2: RKE2 Upgrade (PARTIAL — TARBALL DISASTER)

Initial Approach (WRONG)

Attempted manual tarball installation:

  1. Downloaded rke2.linux-amd64.tar.gz (v1.35.6) to VM118
  2. Extracted with tar xf rke2.linux-amd64.tar.gz -C /FATAL ERROR

Tarball Destruction

The RKE2 tarball contains:

bin/      → RKE2 binaries
lib/      → systemd unit files
share/    → RKE2 data files

Extracting to / on Debian 12 (which uses usrmerge) replaced:

  • /bin symlink (→ usr/bin) with a real directory containing only Rke2 binaries
  • /lib symlink (→ usr/lib) with a real directory containing only systemd files
  • Created bogus /share directory

Result: ALL system binaries inaccessible (ls, sh, bash, systemctl, mount). qm guest exec returns exit code 29. VM was still running (kernel + RKE2 processes in memory) but no new processes could start.

Recovery: Ceph RBD Offline Mount

  1. Stopped VM118 (qm stop 118) — safe because 2/3 CP nodes still provide etcd quorum
  2. Mapped RBD image on proxmox5:
    rbd map vm_disks/vm-118-disk-0  # → /dev/rbd1
    sleep 2  # Wait for udev
    
  3. Found partitions: rbd1p1 (root fs 39.9G), rbd1p14 (BIOS boot), rbd1p15 (EFI 124M)
  4. Mounted root partition: mount /dev/rbd1p1 /mnt/vm118-rescue
  5. Verified damage:
    • /mnt/vm118-rescue/bin = real directory with only rke2, rke2-killall.sh, rke2-uninstall.sh
    • /mnt/vm118-rescue/lib = real directory with only systemd/
    • /mnt/vm118-rescue/usr/bin/bash = INTACT (original system binaries fine)
  6. Fixed symlinks:
    # Save RKE2 binaries
    cp /mnt/vm118-rescue/bin/rke2 /mnt/vm118-rescue/usr/local/bin/
    cp /mnt/vm118-rescue/bin/rke2-*.sh /mnt/vm118-rescue/usr/local/bin/
    
    # Save systemd units
    mkdir -p /mnt/vm118-rescue/usr/lib/systemd/system
    cp -a /mnt/vm118-rescue/lib/systemd/system/rke2-*.service /mnt/vm118-rescue/usr/lib/systemd/system/
    cp -a /mnt/vm118-rescue/lib/systemd/system/rke2-*.env /mnt/vm118-rescue/usr/lib/systemd/system/
    
    # Restore symlinks
    rm -rf /mnt/vm118-rescue/bin
    ln -s usr/bin /mnt/vm118-rescue/bin
    rm -rf /mnt/vm118-rescue/lib
    ln -s usr/lib /mnt/vm118-rescue/lib
    rm -rf /mnt/vm118-rescue/share
    
  7. Unmounted, unmapped, started VM:
    umount /mnt/vm118-rescue
    rbd unmap /dev/rbd1
    qm start 118
    
  8. Verified: Guest agent responded after 60s boot. Shell works. BUT RKE2 server stuck "activating" — v1.35.6 binary couldn't find v1.35.6 runtime images (cached images were v1.35.2).

Second Fix: Restore Matching Binary

Downloaded v1.35.2 tarball and extracted CORRECTLY:

wget https://github.com/rancher/rke2/releases/download/v1.35.2%2Brke2r1/rke2.linux-amd64.tar.gz
tar xf rke2.linux-amd64.tar.gz -C /usr/local/  # CORRECT!
rke2 --version  # v1.35.2+rke2r1
systemctl start rke2-server

Node came back as Ready v1.35.2+rke2r1. Cluster fully healthy again.

RBD Pitfalls Encountered

  • Device disappearance: First rbd map created /dev/rbd1 but it vanished before mount. Had to unmap stale mappings and re-map.
  • No kpartx: PVE node didn't have kpartx. Partition devices were created by kernel automatically after sleep 2 post-mapping.
  • partprobe failed: partprobe /dev/rbd1 gave "Could not stat device" because the device had already disappeared. Re-mapping fixed it.

Phase 2b: Proper Upgrade Path (PREREQS IDENTIFIED, BLOCKED ON SSH)

IaC Repo Analysis

The epic-2-k8s/ansible/playbook.yml uses lablabs.rke2 role v1.50.1:

rke2_version: v1.35.2+rke2r1  # ← Change this to upgrade
rke2_token: "{{ lookup('env', 'RKE2_TOKEN') }}"
rke2_ha_mode: true
rke2_api_ip: 10.0.30.50

Play 2 does rolling restart: Workers (serial:1) → Masters (serial:1).

Prerequisites Status

  1. Ansible role installed on VM200 (ansible-galaxy install -r requirements.yml -p /root/.ansible/roles)
  2. Inventory generated manually (tofu state shows VM IDs, IPs known)
  3. RKE2 token extracted from VM130 config (/etc/rancher/rke2/config.yamltoken: line)
  4. SSH access from VM200 to K8s VMsPermission denied (publickey)
    • mgmt-runner's SSH key NOT in debian@VMs/.ssh/authorized_keys
    • Ansible inventory uses ansible_user: debian
    • Need to deploy VM200's public key to all 6 K8s VMs before Ansible can run

Tofu State Details

Tofu state at /opt/tofu-state/rke2-cluster.tfstate has VM IDs:

rke2-cp-01: vm_id=118, node=proxmox (original placement, HA may move)
rke2-cp-02: vm_id=130, node=proxmox2
rke2-cp-03: vm_id=129, node=proxmox3
rke2-worker-01: vm_id=128, node=proxmox
rke2-worker-02: vm_id=132, node=proxmox2
rke2-worker-03: vm_id=131, node=proxmox3

User Directive

User instructed: "Nutze GitOps!" and "Neue nodes mit neuer version erstellen. Cluster join. Verify. Workloads migrieren. Alte nodes löschen"

This means: prefer GitOps/Ansible for upgrades, or blue-green with new VMs if rolling upgrade is too risky. NEVER manual tarball installation.

Phase 3: openclaw-memory Removal (COMPLETED via GitOps)

Context

User: "Openclaw memory kann weg. Nutze ich nicht mehr"

The openclaw-memory namespace had 3 broken pods:

  • memory-api — CreateContainerConfigError (missing secret)
  • qdrant-0 — CreateContainerConfigError (missing secret)
  • ollama — Running but unused

Two ExternalSecrets in SecretSyncedError state (wrong 1Password vault).

Removal Steps (Clean GitOps)

  1. Deleted from Git repo:

    rm clusters/main/apps/memory.yaml          # ArgoCD Application
    rm -rf clusters/main/memory/                # All manifests
    rm .github/workflows/epic6_prepare-memory-secrets.yml
    git commit -m "chore: remove openclaw-memory (unused, deprecated)"
    git push origin main                        # commit ebc4da5
    
  2. Deleted ArgoCD Application (cascade=foreground):

    kubectl --server=https://10.0.30.52:6443 \
      delete application memory -n argocd --cascade=foreground
    
  3. Deleted namespace (pods cleared, namespace went Terminating → Gone):

    kubectl --server=https://10.0.30.52:6443 delete ns openclaw-memory
    
  4. Verified: Namespace NotFound, 0 non-running pods, ArgoCD apps all Healthy except backups (OutOfSync, non-critical).

Key Observations

  • ArgoCD Application --cascade=foreground cleaned up managed resources (deployments, services) but namespace remained — needed manual kubectl delete ns.
  • No operator CRs involved, so no finalizer issues (unlike MariaDB removal).
  • backups ArgoCD app remains OutOfSync (pre-existing, not related).

Phase 2c: SSH Key Deployment & Ansible Launch (COMPLETED)

Problem

Ansible on VM200 could not SSH to K8s VMs — Permission denied (publickey). VM200 had no SSH keypair at all (~/.ssh/ only had authorized_keys and known_hosts).

Fix

  1. Generated SSH key on VM200:

    ssh root@10.0.30.124 "ssh-keygen -t ed25519 -f ~/.ssh/id_ed25519 -N '' -C 'vm200-ansible'"
    
  2. Confirmed actual VM→PVE-Node mapping (all 6 nodes scanned):

    118 (CP-01)      → proxmox5
    130 (CP-02)      → proxmox3
    129 (CP-03)      → proxmox2
    128 (Worker-01)  → proxmox5  (later: proxmox)
    131 (Worker-03)  → proxmox3
    132 (Worker-02)  → proxmox2
    
  3. Distributed public key to all 6 VMs via qm guest exec — each VM got the key appended to /home/debian/.ssh/authorized_keys with correct ownership (debian:debian) and permissions (700/600). All 6 returned OK.

  4. Verified Ansible connectivity:

    cd /root/iac-homelab/epic-2-k8s/ansible && ansible all -m ping
    

    All 6 nodes: SUCCESS => { "ping": "pong" }

Ansible Playbook Execution

With all prerequisites met, launched the upgrade:

cd /root/iac-homelab/epic-2-k8s/ansible
export RKE2_TOKEN='EyMCxVDobAkWjgyPDcwvXRJQHgWHoH6qgDcvFhQUkd'
ansible-playbook playbook.yml -v

The playbook:

  • Play 1: lablabs.rke2 role applies rke2_version: v1.35.6+rke2r1 to all nodes (writes config, downloads/installs new binaries)
  • Play 2: Rolling restart Workers (serial: 1) — each node restarted, waits for RKE2 agent active + node Ready
  • Play 3: Rolling restart Masters (serial: 1) — same pattern for CP

GitOps trail:

  • Commit 6cb1b3c: feat(rke2): upgrade v1.35.2 → v1.35.6+rke2r1
  • Commit ebc4da5: chore: remove openclaw-memory (unused, deprecated)

Phase 2d: dpkg Lock Blocking Ansible (RESOLVED)

Problem

First Ansible playbook run failed on 5 of 6 nodes. Only Worker-01 succeeded (it had been updated cleanly earlier). The other 5 failed at the pre_tasks stage with:

E: dpkg was interrupted, you must manually run 'sudo dpkg --configure -a'

Root cause: Previous OS updates via qm guest exec left dpkg in a half-configured state on 5 nodes. The openssh-server package was mid-upgrade (held a config file conflict prompt that hung without a TTY).

Fix Sequence

Had to fix each node individually via qm guest exec (could not use Ansible ad-hoc commands because Ansible itself triggers apt):

# For each broken node, on its hosting PVE node:
ssh root@PVE_NODE "qm guest exec VMID --timeout 120 -- bash -c \"
  fuser -k /var/lib/dpkg/lock-frontend 2>/dev/null;
  fuser -k /var/lib/dpkg/lock 2>/dev/null;
  fuser -k /var/cache/debconf/config.dat 2>/dev/null;
  sleep 2;
  rm -f /var/lib/dpkg/lock-frontend /var/lib/dpkg/lock;
  DEBIAN_FRONTEND=noninteractive dpkg --configure -a &&
  apt-get install -y ceph-common rbd-nbd &&
  echo FIXED
\""

All 5 nodes fixed (CP-01, CP-02, CP-03, Worker-02, Worker-03). Then re-ran ansible-playbook playbook.yml successfully.

Key Insight

The lablabs.rke2 role's pre_tasks install ceph-common and rbd-nbd via apt. If any node has dpkg in an interrupted state, the ENTIRE playbook fails on that node. Must fix dpkg on ALL nodes before running the upgrade playbook. Use DEBIAN_FRONTEND=noninteractive to avoid config file conflict prompts hanging under qm guest exec (no TTY).

Phase 2e: Upgrade Completion & Transient CP Failure (COMPLETED)

Result

Second Ansible playbook run succeeded. All 6 nodes upgraded to v1.35.6+rke2r1. PLAY RECAP showed failed=1 on CP-03 (VM129) — but investigation via qm guest exec 129 confirmed the service was already active and running v1.35.6+rke2r1. The failure was a timing race: Ansible's systemctl restart rke2-server raced with RKE2's internal reload from the binary swap. The service had already started before Ansible tried to restart it, causing a transient "already activating" error.

Verification

kubectl get nodes -o wide
# All 6 nodes: Ready, v1.35.6+rke2r1, containerd 2.2.5-k3s2
kubectl get pods -A | grep -v Running | grep -v Completed
# Empty — 0 broken pods

Lesson: Transient CP restart failures can be benign

If a CP node shows failed=1 in the PLAY RECAP but is Ready and running the target version, the failure was a race condition between Ansible's systemctl restart and RKE2's internal reload during the binary swap. No remediation needed — verify cluster state with kubectl get nodes rather than blindly re-running the playbook.

Remaining Work

  1. OS Updates — DONE
  2. RKE2 Upgrade v1.35.2→v1.35.6 — COMPLETED (commit 6cb1b3c)
  3. Helm Chart Updates — COMPLETED (commits 36eb5c4, acd3231) See references/helm-chart-upgrades-2026-07.md for full detail. ArgoCD v2→v3, External Secrets, Ceph CSI RBD, CNPG, Velero all upgraded.
  4. openclaw-memory — REMOVED via GitOps (commit ebc4da5)
  5. ArgoCD backups app OutOfSync — ROOT CAUSE FOUND (cosmetic: ESO default fields in live spec not in Git YAML). Fix pending user approval (read-only investigation completed).

Lessons Learned

  1. NEVER extract RKE2 tarball to / — always use /usr/local/
  2. Debian 12 usrmerge: /bin, /lib, /sbin are symlinks, not real directories. Any tarball that creates bin/ or lib/ at root will destroy the system.
  3. Ceph RBD offline mount is viable for VM filesystem repair when guest agent is broken — stop VM, map RBD, mount partition, fix, unmount, start.
  4. Always discover VM→PVE-Node mapping dynamically — tofu state and actual placement diverge due to HA migrations.
  5. Etcd quorum: With 3 CP nodes, stopping 1 is safe (2/3 quorum). Never stop 2 simultaneously.
  6. RKE2 version must match cached runtime images — a newer binary without matching images will loop "activating" forever.
  7. GitOps-first: User expects cluster changes through IaC repo, not manual interventions. Manual SSH/qm-guest-exec should be last resort.
  8. Non-operator app removal via GitOps is straightforward: delete manifests → push → delete ArgoCD App (cascade) → delete namespace. No finalizer complications unlike operator-managed resources.
  9. Ansible SSH prerequisite: mgmt-runner (VM200) needs its SSH public key in debian@VMs/.ssh/authorized_keys before Ansible can reach K8s nodes. Generate keypair on VM200, distribute via qm guest exec to all 6 VMs, verify with ansible all -m ping before running the playbook.
  10. GitOps-first principle reinforced: User explicitly directed to use the IaC repo (rke2_version variable in Ansible playbook) for upgrades rather than manual operations. All cluster changes should go through dominik/iac-homelab on Gitea → commit → push → Ansible/ArgoCD.
  11. dpkg interrupted state blocks Ansible: If OS updates via qm guest exec leave dpkg half-configured (e.g. openssh-server mid-upgrade with a config file conflict prompt hanging without TTY), the Ansible playbook's pre_tasks (apt install ceph-common rbd-nbd) will fail on that node. Fix: fuser -k on dpkg AND debconf locks, then DEBIAN_FRONTEND=noninteractive dpkg --configure -a on each affected node BEFORE running the upgrade playbook.
  12. debconf has its own lock: Beyond /var/lib/dpkg/lock-frontend and /var/lib/dpkg/lock, debconf maintains /var/cache/debconf/config.dat. A stale process holding this lock causes dpkg --configure -a to fail even after clearing dpkg locks. Must fuser -k all three.
  13. Transient CP restart failures can be benign: Ansible's systemctl restart rke2-server can race with RKE2's internal reload during the binary swap. If a CP node shows failed=1 in PLAY RECAP but is Ready and running the target version, the failure was a timing race — no remediation needed. Always verify with kubectl get nodes -o wide rather than blindly re-running the playbook.