Files

5.9 KiB

RKE2 Kubernetes Cluster — Discovery & Cold-Boot Recovery

TL;DR

An existing RKE2 cluster (3 CP + 3 Workers) was found stopped on Proxmox. Booting it requires only qm start on each VM. The cluster was fully bootstrapped 111 days prior with ArgoCD, Cilium CNI, Ceph CSI, External Secrets, CloudNativePG, MariaDB Operator, and Velero.

Cluster Topology

VMID Name Role Node (dynamic) IP Specs
118 rke2-cp-01 Control Plane proxmox5 10.0.30.51 4c / 12GB / 40GB
130 rke2-cp-02 Control Plane proxmox3 10.0.30.52 4c / 12GB / 40GB
129 rke2-cp-03 Control Plane proxmox2 10.0.30.53 4c / 12GB / 40GB
128 rke2-worker-01 Worker proxmox5 10.0.30.61 4c / 12GB / 80GB
132 rke2-worker-02 Worker proxmox2 10.0.30.62 4c / 12GB / 80GB
131 rke2-worker-03 Worker proxmox3 10.0.30.63 4c / 12GB / 80GB

VMs are HA-managed — PVE may place them on different nodes than where they were originally created. Always scan all nodes for VMs:

for n in proxmox1 proxmox2 proxmox3 proxmox4 proxmox5 proxmox6 proxmox7; do
  ssh -i ~/.ssh/id_ed25519_proxmox root@10.0.20.10 \
    "ssh -o ConnectTimeout=3 root@$n 'qm list 2>/dev/null | grep rke2'"
done

Management VM

VM 200 (mgmt-runner-01) on proxmox3 has kubectl and helm installed. Kubeconfig at /root/.kube/config (must use KUBECONFIG=/root/.kube/config explicitly — kubectl doesn't pick it up from default location via qm guest exec).

Cold-Boot Procedure

1. Start CPs First, Then Workers

# Start all CPs
for vmid in 118 130 129; do
  node=$(find_node_for_vmid $vmid)
  ssh root@$node "qm start $vmid"
done
sleep 15
# Start all workers
for vmid in 128 132 131; do
  node=$(find_node_for_vmid $vmid)
  ssh root@$node "qm start $vmid"
done

2. Wait for Convergence (5-10 min)

Workers initially show NotReady — the rke2-agent needs to retrieve serving-kubelet certs from the API server (returns 503 until CPs converge):

Waiting to retrieve agent configuration;
server is not ready: serving-kubelet.crt: 503 Service Unavailable

This is NORMAL during cold boot. Wait 5-10 minutes. Workers transition to Ready once cert exchange completes.

3. Verify

# Via mgmt-runner
ssh root@10.0.20.10 "ssh root@proxmox3 'qm guest exec 200 -- sh -c \
  \"KUBECONFIG=/root/.kube/config kubectl get nodes\"'"

All 6 nodes should show Ready.

What's Already Installed

Component Namespace Status
ArgoCD argocd Running (App-of-Apps)
Cilium CNI kube-system Running
Ceph CSI (RBD) kube-system 2 CrashLoopBackOff on cold boot
Traefik Ingress kube-system Running (IngressClass)
CoreDNS kube-system Running
Metrics Server kube-system Running
External Secrets external-secrets Running
CloudNativePG cnpg-system Running (PostgreSQL operator)
MariaDB Operator mariadb-operator Running
MariaDB Galera mariadb Starting (needs time)
PostgreSQL postgres 3-replica cluster
Velero velero Partially running
Memory (Qdrant+Ollama) openclaw-memory ContainerCreating

ArgoCD GitOps

  • Admin password: kubectl -n argocd get secret argocd-initial-admin-secret -o jsonpath='{.data.password}' | base64 -d
  • Ingress: argocd host via Traefik (https://10.0.30.70 with Host: argocd)
  • Git repo: http://10.0.30.105:3000/dominik/iac-homelab.git
  • Root app path: clusters/main/apps (App-of-Apps pattern)
  • Auto-sync: prune: true, selfHeal: true
  • Applications: root, operators, databases, mariadb-operator, mariadb-operator-crds, cloudnativepg-operator, velero-operator, backups, memory

Known Cold-Boot Issues

Ceph CSI CrashLoopBackOff

Some ceph-csi-rbd-nodeplugin pods enter CrashLoopBackOff after cold boot. They typically recover after a few restart cycles. If they don't:

kubectl -n kube-system delete pod -l app=csi-rbd-nodeplugin
# Or restart the provisioner
kubectl -n kube-system rollout restart deployment ceph-csi-rbd-provisioner

kube-controller-manager CrashLoopBackOff on CP-01

kube-controller-manager-rke2-cp-01 may CrashLoop briefly during leader election. It resolves once a stable leader is elected among the 3 CPs.

Old Pods Stuck in Terminating

After 111 days offline, many pods show Terminating or Unknown. These are ghosts from the previous run. They clear automatically as the new replicas become ready. Force-delete if stubborn:

kubectl -n <namespace> delete pod <pod-name> --force --grace-period=0

Kubeconfig Transfer

To set up kubectl on a new management machine:

# Get kubeconfig from CP-01 (via qm guest exec)
RAW=$(ssh -i ~/.ssh/id_ed25519_proxmox root@10.0.20.10 \
  "ssh root@proxmox5 'qm guest exec 118 -- cat /etc/rancher/rke2/rke2.yaml'" \
  | python3 -c "import sys,json; print(json.load(sys.stdin)['out-data'])")

# Replace 127.0.0.1 with CP-01 IP
echo "$RAW" | sed 's/127.0.0.1/10.0.30.51/g' > ~/.kube/config
chmod 600 ~/.kube/config

Migration Candidates (PVE → K8s)

Services suitable for K8s migration (ordered by ease):

Phase Services Current Location
1 (Already done) ArgoCD, Operators, DBs RKE2 cluster
2 Grafana, Prometheus, Alertmanager, Blackbox CT 141 (Docker)
3 Traefik (→ IngressController), Authelia CT 99999, CT 112
4 Gitea, Paperless-ngx, Paperless-GPT, LibreChat CT 108, 104, 113, 133
5 Seafile, Immich CT 111 (Docker, 11 containers)
6 Ollama, LiteLLM, Ideogram4 CT 123, 124, CT 111

Services that STAY on PVE (not K8s candidates):

  • Home Assistant (USB/Zigbee, HA OS)
  • Frigate (GPU passthrough — though could use K8s GPU operator)
  • MariaDB Galera VMs (already clustered, K8s adds complexity)
  • Hermes Agent (terminal access, persistent state)
  • PBS (PVE-integrated)
  • Samba shares (kernel-level file serving)
  • Dovecot IMAP