5.9 KiB
RKE2 Kubernetes Cluster — Discovery & Cold-Boot Recovery
TL;DR
An existing RKE2 cluster (3 CP + 3 Workers) was found stopped on Proxmox.
Booting it requires only qm start on each VM. The cluster was fully
bootstrapped 111 days prior with ArgoCD, Cilium CNI, Ceph CSI, External
Secrets, CloudNativePG, MariaDB Operator, and Velero.
Cluster Topology
| VMID | Name | Role | Node (dynamic) | IP | Specs |
|---|---|---|---|---|---|
| 118 | rke2-cp-01 | Control Plane | proxmox5 | 10.0.30.51 | 4c / 12GB / 40GB |
| 130 | rke2-cp-02 | Control Plane | proxmox3 | 10.0.30.52 | 4c / 12GB / 40GB |
| 129 | rke2-cp-03 | Control Plane | proxmox2 | 10.0.30.53 | 4c / 12GB / 40GB |
| 128 | rke2-worker-01 | Worker | proxmox5 | 10.0.30.61 | 4c / 12GB / 80GB |
| 132 | rke2-worker-02 | Worker | proxmox2 | 10.0.30.62 | 4c / 12GB / 80GB |
| 131 | rke2-worker-03 | Worker | proxmox3 | 10.0.30.63 | 4c / 12GB / 80GB |
VMs are HA-managed — PVE may place them on different nodes than where they were originally created. Always scan all nodes for VMs:
for n in proxmox1 proxmox2 proxmox3 proxmox4 proxmox5 proxmox6 proxmox7; do
ssh -i ~/.ssh/id_ed25519_proxmox root@10.0.20.10 \
"ssh -o ConnectTimeout=3 root@$n 'qm list 2>/dev/null | grep rke2'"
done
Management VM
VM 200 (mgmt-runner-01) on proxmox3 has kubectl and helm installed.
Kubeconfig at /root/.kube/config (must use KUBECONFIG=/root/.kube/config
explicitly — kubectl doesn't pick it up from default location via qm guest exec).
Cold-Boot Procedure
1. Start CPs First, Then Workers
# Start all CPs
for vmid in 118 130 129; do
node=$(find_node_for_vmid $vmid)
ssh root@$node "qm start $vmid"
done
sleep 15
# Start all workers
for vmid in 128 132 131; do
node=$(find_node_for_vmid $vmid)
ssh root@$node "qm start $vmid"
done
2. Wait for Convergence (5-10 min)
Workers initially show NotReady — the rke2-agent needs to retrieve
serving-kubelet certs from the API server (returns 503 until CPs converge):
Waiting to retrieve agent configuration;
server is not ready: serving-kubelet.crt: 503 Service Unavailable
This is NORMAL during cold boot. Wait 5-10 minutes. Workers transition
to Ready once cert exchange completes.
3. Verify
# Via mgmt-runner
ssh root@10.0.20.10 "ssh root@proxmox3 'qm guest exec 200 -- sh -c \
\"KUBECONFIG=/root/.kube/config kubectl get nodes\"'"
All 6 nodes should show Ready.
What's Already Installed
| Component | Namespace | Status |
|---|---|---|
| ArgoCD | argocd | Running (App-of-Apps) |
| Cilium CNI | kube-system | Running |
| Ceph CSI (RBD) | kube-system | 2 CrashLoopBackOff on cold boot |
| Traefik Ingress | kube-system | Running (IngressClass) |
| CoreDNS | kube-system | Running |
| Metrics Server | kube-system | Running |
| External Secrets | external-secrets | Running |
| CloudNativePG | cnpg-system | Running (PostgreSQL operator) |
| MariaDB Operator | mariadb-operator | Running |
| MariaDB Galera | mariadb | Starting (needs time) |
| PostgreSQL | postgres | 3-replica cluster |
| Velero | velero | Partially running |
| Memory (Qdrant+Ollama) | openclaw-memory | ContainerCreating |
ArgoCD GitOps
- Admin password:
kubectl -n argocd get secret argocd-initial-admin-secret -o jsonpath='{.data.password}' | base64 -d - Ingress:
argocdhost via Traefik (https://10.0.30.70 with Host: argocd) - Git repo:
http://10.0.30.105:3000/dominik/iac-homelab.git - Root app path:
clusters/main/apps(App-of-Apps pattern) - Auto-sync:
prune: true, selfHeal: true - Applications: root, operators, databases, mariadb-operator, mariadb-operator-crds, cloudnativepg-operator, velero-operator, backups, memory
Known Cold-Boot Issues
Ceph CSI CrashLoopBackOff
Some ceph-csi-rbd-nodeplugin pods enter CrashLoopBackOff after cold boot.
They typically recover after a few restart cycles. If they don't:
kubectl -n kube-system delete pod -l app=csi-rbd-nodeplugin
# Or restart the provisioner
kubectl -n kube-system rollout restart deployment ceph-csi-rbd-provisioner
kube-controller-manager CrashLoopBackOff on CP-01
kube-controller-manager-rke2-cp-01 may CrashLoop briefly during leader
election. It resolves once a stable leader is elected among the 3 CPs.
Old Pods Stuck in Terminating
After 111 days offline, many pods show Terminating or Unknown.
These are ghosts from the previous run. They clear automatically as the
new replicas become ready. Force-delete if stubborn:
kubectl -n <namespace> delete pod <pod-name> --force --grace-period=0
Kubeconfig Transfer
To set up kubectl on a new management machine:
# Get kubeconfig from CP-01 (via qm guest exec)
RAW=$(ssh -i ~/.ssh/id_ed25519_proxmox root@10.0.20.10 \
"ssh root@proxmox5 'qm guest exec 118 -- cat /etc/rancher/rke2/rke2.yaml'" \
| python3 -c "import sys,json; print(json.load(sys.stdin)['out-data'])")
# Replace 127.0.0.1 with CP-01 IP
echo "$RAW" | sed 's/127.0.0.1/10.0.30.51/g' > ~/.kube/config
chmod 600 ~/.kube/config
Migration Candidates (PVE → K8s)
Services suitable for K8s migration (ordered by ease):
| Phase | Services | Current Location |
|---|---|---|
| 1 | (Already done) ArgoCD, Operators, DBs | RKE2 cluster |
| 2 | Grafana, Prometheus, Alertmanager, Blackbox | CT 141 (Docker) |
| 3 | Traefik (→ IngressController), Authelia | CT 99999, CT 112 |
| 4 | Gitea, Paperless-ngx, Paperless-GPT, LibreChat | CT 108, 104, 113, 133 |
| 5 | Seafile, Immich | CT 111 (Docker, 11 containers) |
| 6 | Ollama, LiteLLM, Ideogram4 | CT 123, 124, CT 111 |
Services that STAY on PVE (not K8s candidates):
- Home Assistant (USB/Zigbee, HA OS)
- Frigate (GPU passthrough — though could use K8s GPU operator)
- MariaDB Galera VMs (already clustered, K8s adds complexity)
- Hermes Agent (terminal access, persistent state)
- PBS (PVE-integrated)
- Samba shares (kernel-level file serving)
- Dovecot IMAP