# RKE2 Upgrade Attempt & Tarball Recovery — 2026-07-12 ## Goal Update K8s cluster: OS updates on all VMs, RKE2 upgrade v1.35.2→v1.35.6, Helm chart updates, fix broken components. Scope: K8s only (no PVE hosts). ## Phase 1: OS Updates (COMPLETED) All 7 VMs updated via `qm guest exec` through Proxmox jumphost chain. ### VM→PVE-Node Discovery Initial assumption that all VMs were on proxmox3 was WRONG. Had to scan all PVE nodes to find actual placement: ``` 118 (CP-01) → proxmox5 (NOT proxmox3) 130 (CP-02) → proxmox3 129 (CP-03) → proxmox2 128 (Worker-01) → proxmox5 (later moved to proxmox) 132 (Worker-02) → proxmox2 131 (Worker-03) → proxmox3 200 (mgmt-runner) → proxmox3 ``` ### Update Process - VM200: Direct SSH (10.0.30.124) — 0 updates remaining - VM118, VM129: Updated via qm guest exec on respective PVE nodes - VM132: dpkg interrupted (timeout) — fixed with `dpkg --configure -a` - All VMs: 0 remaining updates after completion - Cluster stayed healthy throughout (6/6 nodes Ready) ### Command Pattern ```bash ssh -i ~/.ssh/id_ed25519_proxmox root@10.0.20.10 \ "ssh -o ConnectTimeout=5 root@PVE_NODE 'qm guest exec VMID --timeout 600 \ -- sh -c \"apt-get update && apt-get upgrade -y\"'" ``` ## Phase 2: RKE2 Upgrade (PARTIAL — TARBALL DISASTER) ### Initial Approach (WRONG) Attempted manual tarball installation: 1. Downloaded `rke2.linux-amd64.tar.gz` (v1.35.6) to VM118 2. Extracted with `tar xf rke2.linux-amd64.tar.gz -C /` ← **FATAL ERROR** ### Tarball Destruction The RKE2 tarball contains: ``` bin/ → RKE2 binaries lib/ → systemd unit files share/ → RKE2 data files ``` Extracting to `/` on Debian 12 (which uses usrmerge) replaced: - `/bin` symlink (→ usr/bin) with a real directory containing only Rke2 binaries - `/lib` symlink (→ usr/lib) with a real directory containing only systemd files - Created bogus `/share` directory Result: ALL system binaries inaccessible (`ls`, `sh`, `bash`, `systemctl`, `mount`). `qm guest exec` returns exit code 29. VM was still running (kernel + RKE2 processes in memory) but no new processes could start. ### Recovery: Ceph RBD Offline Mount 1. **Stopped VM118** (`qm stop 118`) — safe because 2/3 CP nodes still provide etcd quorum 2. **Mapped RBD image** on proxmox5: ```bash rbd map vm_disks/vm-118-disk-0 # → /dev/rbd1 sleep 2 # Wait for udev ``` 3. **Found partitions**: `rbd1p1` (root fs 39.9G), `rbd1p14` (BIOS boot), `rbd1p15` (EFI 124M) 4. **Mounted root partition**: `mount /dev/rbd1p1 /mnt/vm118-rescue` 5. **Verified damage**: - `/mnt/vm118-rescue/bin` = real directory with only `rke2`, `rke2-killall.sh`, `rke2-uninstall.sh` - `/mnt/vm118-rescue/lib` = real directory with only `systemd/` - `/mnt/vm118-rescue/usr/bin/bash` = INTACT (original system binaries fine) 6. **Fixed symlinks**: ```bash # Save RKE2 binaries cp /mnt/vm118-rescue/bin/rke2 /mnt/vm118-rescue/usr/local/bin/ cp /mnt/vm118-rescue/bin/rke2-*.sh /mnt/vm118-rescue/usr/local/bin/ # Save systemd units mkdir -p /mnt/vm118-rescue/usr/lib/systemd/system cp -a /mnt/vm118-rescue/lib/systemd/system/rke2-*.service /mnt/vm118-rescue/usr/lib/systemd/system/ cp -a /mnt/vm118-rescue/lib/systemd/system/rke2-*.env /mnt/vm118-rescue/usr/lib/systemd/system/ # Restore symlinks rm -rf /mnt/vm118-rescue/bin ln -s usr/bin /mnt/vm118-rescue/bin rm -rf /mnt/vm118-rescue/lib ln -s usr/lib /mnt/vm118-rescue/lib rm -rf /mnt/vm118-rescue/share ``` 7. **Unmounted, unmapped, started VM**: ```bash umount /mnt/vm118-rescue rbd unmap /dev/rbd1 qm start 118 ``` 8. **Verified**: Guest agent responded after 60s boot. Shell works. BUT RKE2 server stuck "activating" — v1.35.6 binary couldn't find v1.35.6 runtime images (cached images were v1.35.2). ### Second Fix: Restore Matching Binary Downloaded v1.35.2 tarball and extracted CORRECTLY: ```bash wget https://github.com/rancher/rke2/releases/download/v1.35.2%2Brke2r1/rke2.linux-amd64.tar.gz tar xf rke2.linux-amd64.tar.gz -C /usr/local/ # CORRECT! rke2 --version # v1.35.2+rke2r1 systemctl start rke2-server ``` Node came back as Ready v1.35.2+rke2r1. Cluster fully healthy again. ### RBD Pitfalls Encountered - **Device disappearance**: First `rbd map` created `/dev/rbd1` but it vanished before mount. Had to unmap stale mappings and re-map. - **No kpartx**: PVE node didn't have `kpartx`. Partition devices were created by kernel automatically after `sleep 2` post-mapping. - **partprobe failed**: `partprobe /dev/rbd1` gave "Could not stat device" because the device had already disappeared. Re-mapping fixed it. ## Phase 2b: Proper Upgrade Path (PREREQS IDENTIFIED, BLOCKED ON SSH) ### IaC Repo Analysis The `epic-2-k8s/ansible/playbook.yml` uses `lablabs.rke2` role v1.50.1: ```yaml rke2_version: v1.35.2+rke2r1 # ← Change this to upgrade rke2_token: "{{ lookup('env', 'RKE2_TOKEN') }}" rke2_ha_mode: true rke2_api_ip: 10.0.30.50 ``` Play 2 does rolling restart: Workers (serial:1) → Masters (serial:1). ### Prerequisites Status 1. ✅ Ansible role installed on VM200 (`ansible-galaxy install -r requirements.yml -p /root/.ansible/roles`) 2. ✅ Inventory generated manually (tofu state shows VM IDs, IPs known) 3. ✅ RKE2 token extracted from VM130 config (`/etc/rancher/rke2/config.yaml` → `token:` line) 4. ❌ **SSH access from VM200 to K8s VMs** — `Permission denied (publickey)` - mgmt-runner's SSH key NOT in `debian@VMs/.ssh/authorized_keys` - Ansible inventory uses `ansible_user: debian` - Need to deploy VM200's public key to all 6 K8s VMs before Ansible can run ### Tofu State Details Tofu state at `/opt/tofu-state/rke2-cluster.tfstate` has VM IDs: ``` rke2-cp-01: vm_id=118, node=proxmox (original placement, HA may move) rke2-cp-02: vm_id=130, node=proxmox2 rke2-cp-03: vm_id=129, node=proxmox3 rke2-worker-01: vm_id=128, node=proxmox rke2-worker-02: vm_id=132, node=proxmox2 rke2-worker-03: vm_id=131, node=proxmox3 ``` ### User Directive User instructed: "Nutze GitOps!" and "Neue nodes mit neuer version erstellen. Cluster join. Verify. Workloads migrieren. Alte nodes löschen" This means: prefer GitOps/Ansible for upgrades, or blue-green with new VMs if rolling upgrade is too risky. NEVER manual tarball installation. ## Phase 3: openclaw-memory Removal (COMPLETED via GitOps) ### Context User: "Openclaw memory kann weg. Nutze ich nicht mehr" The `openclaw-memory` namespace had 3 broken pods: - `memory-api` — CreateContainerConfigError (missing secret) - `qdrant-0` — CreateContainerConfigError (missing secret) - `ollama` — Running but unused Two ExternalSecrets in `SecretSyncedError` state (wrong 1Password vault). ### Removal Steps (Clean GitOps) 1. **Deleted from Git repo**: ```bash rm clusters/main/apps/memory.yaml # ArgoCD Application rm -rf clusters/main/memory/ # All manifests rm .github/workflows/epic6_prepare-memory-secrets.yml git commit -m "chore: remove openclaw-memory (unused, deprecated)" git push origin main # commit ebc4da5 ``` 2. **Deleted ArgoCD Application** (cascade=foreground): ```bash kubectl --server=https://10.0.30.52:6443 \ delete application memory -n argocd --cascade=foreground ``` 3. **Deleted namespace** (pods cleared, namespace went Terminating → Gone): ```bash kubectl --server=https://10.0.30.52:6443 delete ns openclaw-memory ``` 4. **Verified**: Namespace NotFound, 0 non-running pods, ArgoCD apps all Healthy except `backups` (OutOfSync, non-critical). ### Key Observations - ArgoCD Application `--cascade=foreground` cleaned up managed resources (deployments, services) but namespace remained — needed manual `kubectl delete ns`. - No operator CRs involved, so no finalizer issues (unlike MariaDB removal). - `backups` ArgoCD app remains OutOfSync (pre-existing, not related). ## Phase 2c: SSH Key Deployment & Ansible Launch (COMPLETED) ### Problem Ansible on VM200 could not SSH to K8s VMs — `Permission denied (publickey)`. VM200 had no SSH keypair at all (`~/.ssh/` only had `authorized_keys` and `known_hosts`). ### Fix 1. **Generated SSH key on VM200**: ```bash ssh root@10.0.30.124 "ssh-keygen -t ed25519 -f ~/.ssh/id_ed25519 -N '' -C 'vm200-ansible'" ``` 2. **Confirmed actual VM→PVE-Node mapping** (all 6 nodes scanned): ``` 118 (CP-01) → proxmox5 130 (CP-02) → proxmox3 129 (CP-03) → proxmox2 128 (Worker-01) → proxmox5 (later: proxmox) 131 (Worker-03) → proxmox3 132 (Worker-02) → proxmox2 ``` 3. **Distributed public key to all 6 VMs** via `qm guest exec` — each VM got the key appended to `/home/debian/.ssh/authorized_keys` with correct ownership (debian:debian) and permissions (700/600). All 6 returned OK. 4. **Verified Ansible connectivity**: ```bash cd /root/iac-homelab/epic-2-k8s/ansible && ansible all -m ping ``` All 6 nodes: SUCCESS => { "ping": "pong" } ### Ansible Playbook Execution With all prerequisites met, launched the upgrade: ```bash cd /root/iac-homelab/epic-2-k8s/ansible export RKE2_TOKEN='EyMCxVDobAkWjgyPDcwvXRJQHgWHoH6qgDcvFhQUkd' ansible-playbook playbook.yml -v ``` The playbook: - Play 1: `lablabs.rke2` role applies `rke2_version: v1.35.6+rke2r1` to all nodes (writes config, downloads/installs new binaries) - Play 2: Rolling restart Workers (serial: 1) — each node restarted, waits for RKE2 agent active + node Ready - Play 3: Rolling restart Masters (serial: 1) — same pattern for CP GitOps trail: - Commit `6cb1b3c`: `feat(rke2): upgrade v1.35.2 → v1.35.6+rke2r1` - Commit `ebc4da5`: `chore: remove openclaw-memory (unused, deprecated)` ## Phase 2d: dpkg Lock Blocking Ansible (RESOLVED) ### Problem First Ansible playbook run failed on 5 of 6 nodes. Only Worker-01 succeeded (it had been updated cleanly earlier). The other 5 failed at the pre_tasks stage with: ``` E: dpkg was interrupted, you must manually run 'sudo dpkg --configure -a' ``` Root cause: Previous OS updates via `qm guest exec` left dpkg in a half-configured state on 5 nodes. The `openssh-server` package was mid-upgrade (held a config file conflict prompt that hung without a TTY). ### Fix Sequence Had to fix each node individually via `qm guest exec` (could not use Ansible ad-hoc commands because Ansible itself triggers apt): ```bash # For each broken node, on its hosting PVE node: ssh root@PVE_NODE "qm guest exec VMID --timeout 120 -- bash -c \" fuser -k /var/lib/dpkg/lock-frontend 2>/dev/null; fuser -k /var/lib/dpkg/lock 2>/dev/null; fuser -k /var/cache/debconf/config.dat 2>/dev/null; sleep 2; rm -f /var/lib/dpkg/lock-frontend /var/lib/dpkg/lock; DEBIAN_FRONTEND=noninteractive dpkg --configure -a && apt-get install -y ceph-common rbd-nbd && echo FIXED \"" ``` All 5 nodes fixed (CP-01, CP-02, CP-03, Worker-02, Worker-03). Then re-ran `ansible-playbook playbook.yml` successfully. ### Key Insight The `lablabs.rke2` role's pre_tasks install `ceph-common` and `rbd-nbd` via apt. If any node has dpkg in an interrupted state, the ENTIRE playbook fails on that node. Must fix dpkg on ALL nodes before running the upgrade playbook. Use `DEBIAN_FRONTEND=noninteractive` to avoid config file conflict prompts hanging under `qm guest exec` (no TTY). ## Phase 2e: Upgrade Completion & Transient CP Failure (COMPLETED) ### Result Second Ansible playbook run succeeded. All 6 nodes upgraded to v1.35.6+rke2r1. PLAY RECAP showed `failed=1` on CP-03 (VM129) — but investigation via `qm guest exec 129` confirmed the service was already active and running v1.35.6+rke2r1. The failure was a timing race: Ansible's `systemctl restart rke2-server` raced with RKE2's internal reload from the binary swap. The service had already started before Ansible tried to restart it, causing a transient "already activating" error. ### Verification ```bash kubectl get nodes -o wide # All 6 nodes: Ready, v1.35.6+rke2r1, containerd 2.2.5-k3s2 kubectl get pods -A | grep -v Running | grep -v Completed # Empty — 0 broken pods ``` ### Lesson: Transient CP restart failures can be benign If a CP node shows `failed=1` in the PLAY RECAP but is `Ready` and running the target version, the failure was a race condition between Ansible's `systemctl restart` and RKE2's internal reload during the binary swap. No remediation needed — verify cluster state with `kubectl get nodes` rather than blindly re-running the playbook. ## Remaining Work 1. ✅ OS Updates — DONE 2. ✅ RKE2 Upgrade v1.35.2→v1.35.6 — COMPLETED (commit 6cb1b3c) 3. ✅ Helm Chart Updates — COMPLETED (commits 36eb5c4, acd3231) See `references/helm-chart-upgrades-2026-07.md` for full detail. ArgoCD v2→v3, External Secrets, Ceph CSI RBD, CNPG, Velero all upgraded. 4. ✅ openclaw-memory — REMOVED via GitOps (commit ebc4da5) 5. ✅ ArgoCD `backups` app OutOfSync — ROOT CAUSE FOUND (cosmetic: ESO default fields in live spec not in Git YAML). Fix pending user approval (read-only investigation completed). ## Lessons Learned 1. **NEVER extract RKE2 tarball to `/`** — always use `/usr/local/` 2. **Debian 12 usrmerge**: `/bin`, `/lib`, `/sbin` are symlinks, not real directories. Any tarball that creates `bin/` or `lib/` at root will destroy the system. 3. **Ceph RBD offline mount** is viable for VM filesystem repair when guest agent is broken — stop VM, map RBD, mount partition, fix, unmount, start. 4. **Always discover VM→PVE-Node mapping dynamically** — tofu state and actual placement diverge due to HA migrations. 5. **Etcd quorum**: With 3 CP nodes, stopping 1 is safe (2/3 quorum). Never stop 2 simultaneously. 6. **RKE2 version must match cached runtime images** — a newer binary without matching images will loop "activating" forever. 7. **GitOps-first**: User expects cluster changes through IaC repo, not manual interventions. Manual SSH/qm-guest-exec should be last resort. 8. **Non-operator app removal via GitOps** is straightforward: delete manifests → push → delete ArgoCD App (cascade) → delete namespace. No finalizer complications unlike operator-managed resources. 9. **Ansible SSH prerequisite**: mgmt-runner (VM200) needs its SSH public key in `debian@VMs/.ssh/authorized_keys` before Ansible can reach K8s nodes. Generate keypair on VM200, distribute via `qm guest exec` to all 6 VMs, verify with `ansible all -m ping` before running the playbook. 10. **GitOps-first principle reinforced**: User explicitly directed to use the IaC repo (`rke2_version` variable in Ansible playbook) for upgrades rather than manual operations. All cluster changes should go through `dominik/iac-homelab` on Gitea → commit → push → Ansible/ArgoCD. 11. **dpkg interrupted state blocks Ansible**: If OS updates via `qm guest exec` leave dpkg half-configured (e.g. openssh-server mid-upgrade with a config file conflict prompt hanging without TTY), the Ansible playbook's pre_tasks (`apt install ceph-common rbd-nbd`) will fail on that node. Fix: `fuser -k` on dpkg AND debconf locks, then `DEBIAN_FRONTEND=noninteractive dpkg --configure -a` on each affected node BEFORE running the upgrade playbook. 12. **debconf has its own lock**: Beyond `/var/lib/dpkg/lock-frontend` and `/var/lib/dpkg/lock`, debconf maintains `/var/cache/debconf/config.dat`. A stale process holding this lock causes `dpkg --configure -a` to fail even after clearing dpkg locks. Must `fuser -k` all three. 13. **Transient CP restart failures can be benign**: Ansible's `systemctl restart rke2-server` can race with RKE2's internal reload during the binary swap. If a CP node shows `failed=1` in PLAY RECAP but is `Ready` and running the target version, the failure was a timing race — no remediation needed. Always verify with `kubectl get nodes -o wide` rather than blindly re-running the playbook.