16 KiB
RKE2 Upgrade Attempt & Tarball Recovery — 2026-07-12
Goal
Update K8s cluster: OS updates on all VMs, RKE2 upgrade v1.35.2→v1.35.6, Helm chart updates, fix broken components. Scope: K8s only (no PVE hosts).
Phase 1: OS Updates (COMPLETED)
All 7 VMs updated via qm guest exec through Proxmox jumphost chain.
VM→PVE-Node Discovery
Initial assumption that all VMs were on proxmox3 was WRONG. Had to scan all PVE nodes to find actual placement:
118 (CP-01) → proxmox5 (NOT proxmox3)
130 (CP-02) → proxmox3
129 (CP-03) → proxmox2
128 (Worker-01) → proxmox5 (later moved to proxmox)
132 (Worker-02) → proxmox2
131 (Worker-03) → proxmox3
200 (mgmt-runner) → proxmox3
Update Process
- VM200: Direct SSH (10.0.30.124) — 0 updates remaining
- VM118, VM129: Updated via qm guest exec on respective PVE nodes
- VM132: dpkg interrupted (timeout) — fixed with
dpkg --configure -a - All VMs: 0 remaining updates after completion
- Cluster stayed healthy throughout (6/6 nodes Ready)
Command Pattern
ssh -i ~/.ssh/id_ed25519_proxmox root@10.0.20.10 \
"ssh -o ConnectTimeout=5 root@PVE_NODE 'qm guest exec VMID --timeout 600 \
-- sh -c \"apt-get update && apt-get upgrade -y\"'"
Phase 2: RKE2 Upgrade (PARTIAL — TARBALL DISASTER)
Initial Approach (WRONG)
Attempted manual tarball installation:
- Downloaded
rke2.linux-amd64.tar.gz(v1.35.6) to VM118 - Extracted with
tar xf rke2.linux-amd64.tar.gz -C /← FATAL ERROR
Tarball Destruction
The RKE2 tarball contains:
bin/ → RKE2 binaries
lib/ → systemd unit files
share/ → RKE2 data files
Extracting to / on Debian 12 (which uses usrmerge) replaced:
/binsymlink (→ usr/bin) with a real directory containing only Rke2 binaries/libsymlink (→ usr/lib) with a real directory containing only systemd files- Created bogus
/sharedirectory
Result: ALL system binaries inaccessible (ls, sh, bash, systemctl,
mount). qm guest exec returns exit code 29. VM was still running
(kernel + RKE2 processes in memory) but no new processes could start.
Recovery: Ceph RBD Offline Mount
- Stopped VM118 (
qm stop 118) — safe because 2/3 CP nodes still provide etcd quorum - Mapped RBD image on proxmox5:
rbd map vm_disks/vm-118-disk-0 # → /dev/rbd1 sleep 2 # Wait for udev - Found partitions:
rbd1p1(root fs 39.9G),rbd1p14(BIOS boot),rbd1p15(EFI 124M) - Mounted root partition:
mount /dev/rbd1p1 /mnt/vm118-rescue - Verified damage:
/mnt/vm118-rescue/bin= real directory with onlyrke2,rke2-killall.sh,rke2-uninstall.sh/mnt/vm118-rescue/lib= real directory with onlysystemd//mnt/vm118-rescue/usr/bin/bash= INTACT (original system binaries fine)
- Fixed symlinks:
# Save RKE2 binaries cp /mnt/vm118-rescue/bin/rke2 /mnt/vm118-rescue/usr/local/bin/ cp /mnt/vm118-rescue/bin/rke2-*.sh /mnt/vm118-rescue/usr/local/bin/ # Save systemd units mkdir -p /mnt/vm118-rescue/usr/lib/systemd/system cp -a /mnt/vm118-rescue/lib/systemd/system/rke2-*.service /mnt/vm118-rescue/usr/lib/systemd/system/ cp -a /mnt/vm118-rescue/lib/systemd/system/rke2-*.env /mnt/vm118-rescue/usr/lib/systemd/system/ # Restore symlinks rm -rf /mnt/vm118-rescue/bin ln -s usr/bin /mnt/vm118-rescue/bin rm -rf /mnt/vm118-rescue/lib ln -s usr/lib /mnt/vm118-rescue/lib rm -rf /mnt/vm118-rescue/share - Unmounted, unmapped, started VM:
umount /mnt/vm118-rescue rbd unmap /dev/rbd1 qm start 118 - Verified: Guest agent responded after 60s boot. Shell works. BUT RKE2 server stuck "activating" — v1.35.6 binary couldn't find v1.35.6 runtime images (cached images were v1.35.2).
Second Fix: Restore Matching Binary
Downloaded v1.35.2 tarball and extracted CORRECTLY:
wget https://github.com/rancher/rke2/releases/download/v1.35.2%2Brke2r1/rke2.linux-amd64.tar.gz
tar xf rke2.linux-amd64.tar.gz -C /usr/local/ # CORRECT!
rke2 --version # v1.35.2+rke2r1
systemctl start rke2-server
Node came back as Ready v1.35.2+rke2r1. Cluster fully healthy again.
RBD Pitfalls Encountered
- Device disappearance: First
rbd mapcreated/dev/rbd1but it vanished before mount. Had to unmap stale mappings and re-map. - No kpartx: PVE node didn't have
kpartx. Partition devices were created by kernel automatically aftersleep 2post-mapping. - partprobe failed:
partprobe /dev/rbd1gave "Could not stat device" because the device had already disappeared. Re-mapping fixed it.
Phase 2b: Proper Upgrade Path (PREREQS IDENTIFIED, BLOCKED ON SSH)
IaC Repo Analysis
The epic-2-k8s/ansible/playbook.yml uses lablabs.rke2 role v1.50.1:
rke2_version: v1.35.2+rke2r1 # ← Change this to upgrade
rke2_token: "{{ lookup('env', 'RKE2_TOKEN') }}"
rke2_ha_mode: true
rke2_api_ip: 10.0.30.50
Play 2 does rolling restart: Workers (serial:1) → Masters (serial:1).
Prerequisites Status
- ✅ Ansible role installed on VM200 (
ansible-galaxy install -r requirements.yml -p /root/.ansible/roles) - ✅ Inventory generated manually (tofu state shows VM IDs, IPs known)
- ✅ RKE2 token extracted from VM130 config (
/etc/rancher/rke2/config.yaml→token:line) - ❌ SSH access from VM200 to K8s VMs —
Permission denied (publickey)- mgmt-runner's SSH key NOT in
debian@VMs/.ssh/authorized_keys - Ansible inventory uses
ansible_user: debian - Need to deploy VM200's public key to all 6 K8s VMs before Ansible can run
- mgmt-runner's SSH key NOT in
Tofu State Details
Tofu state at /opt/tofu-state/rke2-cluster.tfstate has VM IDs:
rke2-cp-01: vm_id=118, node=proxmox (original placement, HA may move)
rke2-cp-02: vm_id=130, node=proxmox2
rke2-cp-03: vm_id=129, node=proxmox3
rke2-worker-01: vm_id=128, node=proxmox
rke2-worker-02: vm_id=132, node=proxmox2
rke2-worker-03: vm_id=131, node=proxmox3
User Directive
User instructed: "Nutze GitOps!" and "Neue nodes mit neuer version erstellen. Cluster join. Verify. Workloads migrieren. Alte nodes löschen"
This means: prefer GitOps/Ansible for upgrades, or blue-green with new VMs if rolling upgrade is too risky. NEVER manual tarball installation.
Phase 3: openclaw-memory Removal (COMPLETED via GitOps)
Context
User: "Openclaw memory kann weg. Nutze ich nicht mehr"
The openclaw-memory namespace had 3 broken pods:
memory-api— CreateContainerConfigError (missing secret)qdrant-0— CreateContainerConfigError (missing secret)ollama— Running but unused
Two ExternalSecrets in SecretSyncedError state (wrong 1Password vault).
Removal Steps (Clean GitOps)
-
Deleted from Git repo:
rm clusters/main/apps/memory.yaml # ArgoCD Application rm -rf clusters/main/memory/ # All manifests rm .github/workflows/epic6_prepare-memory-secrets.yml git commit -m "chore: remove openclaw-memory (unused, deprecated)" git push origin main # commit ebc4da5 -
Deleted ArgoCD Application (cascade=foreground):
kubectl --server=https://10.0.30.52:6443 \ delete application memory -n argocd --cascade=foreground -
Deleted namespace (pods cleared, namespace went Terminating → Gone):
kubectl --server=https://10.0.30.52:6443 delete ns openclaw-memory -
Verified: Namespace NotFound, 0 non-running pods, ArgoCD apps all Healthy except
backups(OutOfSync, non-critical).
Key Observations
- ArgoCD Application
--cascade=foregroundcleaned up managed resources (deployments, services) but namespace remained — needed manualkubectl delete ns. - No operator CRs involved, so no finalizer issues (unlike MariaDB removal).
backupsArgoCD app remains OutOfSync (pre-existing, not related).
Phase 2c: SSH Key Deployment & Ansible Launch (COMPLETED)
Problem
Ansible on VM200 could not SSH to K8s VMs — Permission denied (publickey).
VM200 had no SSH keypair at all (~/.ssh/ only had authorized_keys and
known_hosts).
Fix
-
Generated SSH key on VM200:
ssh root@10.0.30.124 "ssh-keygen -t ed25519 -f ~/.ssh/id_ed25519 -N '' -C 'vm200-ansible'" -
Confirmed actual VM→PVE-Node mapping (all 6 nodes scanned):
118 (CP-01) → proxmox5 130 (CP-02) → proxmox3 129 (CP-03) → proxmox2 128 (Worker-01) → proxmox5 (later: proxmox) 131 (Worker-03) → proxmox3 132 (Worker-02) → proxmox2 -
Distributed public key to all 6 VMs via
qm guest exec— each VM got the key appended to/home/debian/.ssh/authorized_keyswith correct ownership (debian:debian) and permissions (700/600). All 6 returned OK. -
Verified Ansible connectivity:
cd /root/iac-homelab/epic-2-k8s/ansible && ansible all -m pingAll 6 nodes: SUCCESS => { "ping": "pong" }
Ansible Playbook Execution
With all prerequisites met, launched the upgrade:
cd /root/iac-homelab/epic-2-k8s/ansible
export RKE2_TOKEN='EyMCxVDobAkWjgyPDcwvXRJQHgWHoH6qgDcvFhQUkd'
ansible-playbook playbook.yml -v
The playbook:
- Play 1:
lablabs.rke2role appliesrke2_version: v1.35.6+rke2r1to all nodes (writes config, downloads/installs new binaries) - Play 2: Rolling restart Workers (serial: 1) — each node restarted, waits for RKE2 agent active + node Ready
- Play 3: Rolling restart Masters (serial: 1) — same pattern for CP
GitOps trail:
- Commit
6cb1b3c:feat(rke2): upgrade v1.35.2 → v1.35.6+rke2r1 - Commit
ebc4da5:chore: remove openclaw-memory (unused, deprecated)
Phase 2d: dpkg Lock Blocking Ansible (RESOLVED)
Problem
First Ansible playbook run failed on 5 of 6 nodes. Only Worker-01 succeeded (it had been updated cleanly earlier). The other 5 failed at the pre_tasks stage with:
E: dpkg was interrupted, you must manually run 'sudo dpkg --configure -a'
Root cause: Previous OS updates via qm guest exec left dpkg in a
half-configured state on 5 nodes. The openssh-server package was
mid-upgrade (held a config file conflict prompt that hung without a TTY).
Fix Sequence
Had to fix each node individually via qm guest exec (could not use
Ansible ad-hoc commands because Ansible itself triggers apt):
# For each broken node, on its hosting PVE node:
ssh root@PVE_NODE "qm guest exec VMID --timeout 120 -- bash -c \"
fuser -k /var/lib/dpkg/lock-frontend 2>/dev/null;
fuser -k /var/lib/dpkg/lock 2>/dev/null;
fuser -k /var/cache/debconf/config.dat 2>/dev/null;
sleep 2;
rm -f /var/lib/dpkg/lock-frontend /var/lib/dpkg/lock;
DEBIAN_FRONTEND=noninteractive dpkg --configure -a &&
apt-get install -y ceph-common rbd-nbd &&
echo FIXED
\""
All 5 nodes fixed (CP-01, CP-02, CP-03, Worker-02, Worker-03). Then
re-ran ansible-playbook playbook.yml successfully.
Key Insight
The lablabs.rke2 role's pre_tasks install ceph-common and rbd-nbd
via apt. If any node has dpkg in an interrupted state, the ENTIRE
playbook fails on that node. Must fix dpkg on ALL nodes before running
the upgrade playbook. Use DEBIAN_FRONTEND=noninteractive to avoid
config file conflict prompts hanging under qm guest exec (no TTY).
Phase 2e: Upgrade Completion & Transient CP Failure (COMPLETED)
Result
Second Ansible playbook run succeeded. All 6 nodes upgraded to
v1.35.6+rke2r1. PLAY RECAP showed failed=1 on CP-03 (VM129) — but
investigation via qm guest exec 129 confirmed the service was already
active and running v1.35.6+rke2r1. The failure was a timing race:
Ansible's systemctl restart rke2-server raced with RKE2's internal
reload from the binary swap. The service had already started before
Ansible tried to restart it, causing a transient "already activating"
error.
Verification
kubectl get nodes -o wide
# All 6 nodes: Ready, v1.35.6+rke2r1, containerd 2.2.5-k3s2
kubectl get pods -A | grep -v Running | grep -v Completed
# Empty — 0 broken pods
Lesson: Transient CP restart failures can be benign
If a CP node shows failed=1 in the PLAY RECAP but is Ready and
running the target version, the failure was a race condition between
Ansible's systemctl restart and RKE2's internal reload during the
binary swap. No remediation needed — verify cluster state with
kubectl get nodes rather than blindly re-running the playbook.
Remaining Work
- ✅ OS Updates — DONE
- ✅ RKE2 Upgrade v1.35.2→v1.35.6 — COMPLETED (commit 6cb1b3c)
- ✅ Helm Chart Updates — COMPLETED (commits 36eb5c4, acd3231)
See
references/helm-chart-upgrades-2026-07.mdfor full detail. ArgoCD v2→v3, External Secrets, Ceph CSI RBD, CNPG, Velero all upgraded. - ✅ openclaw-memory — REMOVED via GitOps (commit ebc4da5)
- ✅ ArgoCD
backupsapp OutOfSync — ROOT CAUSE FOUND (cosmetic: ESO default fields in live spec not in Git YAML). Fix pending user approval (read-only investigation completed).
Lessons Learned
- NEVER extract RKE2 tarball to
/— always use/usr/local/ - Debian 12 usrmerge:
/bin,/lib,/sbinare symlinks, not real directories. Any tarball that createsbin/orlib/at root will destroy the system. - Ceph RBD offline mount is viable for VM filesystem repair when guest agent is broken — stop VM, map RBD, mount partition, fix, unmount, start.
- Always discover VM→PVE-Node mapping dynamically — tofu state and actual placement diverge due to HA migrations.
- Etcd quorum: With 3 CP nodes, stopping 1 is safe (2/3 quorum). Never stop 2 simultaneously.
- RKE2 version must match cached runtime images — a newer binary without matching images will loop "activating" forever.
- GitOps-first: User expects cluster changes through IaC repo, not manual interventions. Manual SSH/qm-guest-exec should be last resort.
- Non-operator app removal via GitOps is straightforward: delete manifests → push → delete ArgoCD App (cascade) → delete namespace. No finalizer complications unlike operator-managed resources.
- Ansible SSH prerequisite: mgmt-runner (VM200) needs its SSH public
key in
debian@VMs/.ssh/authorized_keysbefore Ansible can reach K8s nodes. Generate keypair on VM200, distribute viaqm guest execto all 6 VMs, verify withansible all -m pingbefore running the playbook. - GitOps-first principle reinforced: User explicitly directed to use
the IaC repo (
rke2_versionvariable in Ansible playbook) for upgrades rather than manual operations. All cluster changes should go throughdominik/iac-homelabon Gitea → commit → push → Ansible/ArgoCD. - dpkg interrupted state blocks Ansible: If OS updates via
qm guest execleave dpkg half-configured (e.g. openssh-server mid-upgrade with a config file conflict prompt hanging without TTY), the Ansible playbook's pre_tasks (apt install ceph-common rbd-nbd) will fail on that node. Fix:fuser -kon dpkg AND debconf locks, thenDEBIAN_FRONTEND=noninteractive dpkg --configure -aon each affected node BEFORE running the upgrade playbook. - debconf has its own lock: Beyond
/var/lib/dpkg/lock-frontendand/var/lib/dpkg/lock, debconf maintains/var/cache/debconf/config.dat. A stale process holding this lock causesdpkg --configure -ato fail even after clearing dpkg locks. Mustfuser -kall three. - Transient CP restart failures can be benign: Ansible's
systemctl restart rke2-servercan race with RKE2's internal reload during the binary swap. If a CP node showsfailed=1in PLAY RECAP but isReadyand running the target version, the failure was a timing race — no remediation needed. Always verify withkubectl get nodes -o widerather than blindly re-running the playbook.