Files
hermes-skills/devops/rke2-cluster-administration/references/cold-boot-recovery-2026-07.md
T

163 lines
5.9 KiB
Markdown

# RKE2 Cluster Cold-Boot Recovery — 2026-07-12
## Situation
RKE2 cluster (6 VMs: 3 CP + 3 Worker) was offline for 111 days. All VMs
were stopped across the PVE cluster. User requested cluster startup
and K8s evaluation for Docker workload migration.
## Startup Sequence
### 1. VM Discovery
VMs had been HA-migrated to various PVE nodes. Had to scan all nodes
to find where each VM landed:
```
118 (CP-01) → proxmox5
130 (CP-02) → proxmox3
129 (CP-03) → proxmox2
128 (Worker-01) → proxmox5
132 (Worker-02) → proxmox2
131 (Worker-03) → proxmox3
200 (mgmt-runner) → proxmox3
```
### 2. Start Order
1. Started all 6 RKE2 VMs via `qm start` (HA placed them on various nodes)
2. CP nodes began etcd convergence within 30s
3. Workers initially showed `NotReady` — waiting for
`serving-kubelet.crt` retrieval (503 from API server)
4. Worker-03 became Ready after ~2 min
5. Worker-02 became Ready after ~3 min (slower due to proxy reconnection)
6. All 6 nodes Ready within 5 min
### 3. Convergence Timeline
| T+0 | VMs started |
| T+30s | CPs booting, etcd forming quorum |
| T+1min | Workers connecting via wss://CP:9345/v1-rke2/connect |
| T+2min | First workers Ready (cert retrieval successful) |
| T+3min | All nodes Ready, Cilium agents starting |
| T+5min | ArgoCD, External Secrets, operators beginning reconciliation |
| T+10min | Most pods Running, old replicas Terminating |
| T+15min | Databases recovering (Galera, PostgreSQL) |
### 4. Pod States After Cold Boot
Most pods showed mixed states initially:
- `Running` — new replicas scheduled by deployment controller
- `Terminating` — old replicas from before shutdown (111 days old)
- `Unknown` — node was down when status was last reported
- `CrashLoopBackOff` — ceph-csi nodeplugin (2 of 6) — Ceph connectivity
- `CreateContainerConfigError` — memory-api, qdrant — missing secrets
(ExternalSecret hadn't synced yet)
All resolved naturally except:
- 3 ExternalSecrets with SecretSyncedError (1Password item issues)
- 2 ceph-csi-rbd-nodeplugin CrashLoopBackOff
## Cilium LB Pool Migration
### Problem
Cilium LB pool was set to `10.0.30.70-10.0.30.89`, overlapping with
Galera/MaxScale VIPs:
- 10.0.30.70 = Keepalived VIP for MaxScale (Galera)
- 10.0.30.71-73 = Galera VM IPs
K8s services had claimed:
- Traefik → 10.0.30.70 (conflicted with MaxScale VIP!)
- PostgreSQL RW → 10.0.30.71 (conflicted with Galera VM!)
- MariaDB MaxScale (K8s) → 10.0.30.72
### Fix Applied
1. Deleted live `CiliumLoadBalancerIPPool` CRD
2. Created new pool with `10.0.30.200-10.0.30.250`
3. Services auto-reassigned:
- Traefik → 10.0.30.202
- MariaDB MaxScale (K8s) → 10.0.30.201
- PostgreSQL RW → 10.0.30.200
4. Updated manifest on all 3 CP nodes:
`/var/lib/rancher/rke2/server/manifests/cilium-l2-lb.yaml`
5. Fixed IaC repo: `epic-2-k8s/ansible/playbook.yml` — committed + pushed
### Verification
- `10.0.30.70` responds to ping (Keepalived VIP for MaxScale — correct!)
- No K8s LoadBalancer service holds IPs in .70-.89 range
- Cilium LB pool shows 48 IPs available in .200-.250 range
## K8s MariaDB Removal
### Decision
User clarified: Galera lives as **native VMs** (VM300-302 + MaxScale
VM310), not in K8s. The K8s MariaDB Galera (3 pods) + MaxScale (2 pods)
in namespace `mariadb` was redundant and should be removed.
### Removal Steps
1. Removed YAML files from `iac-homelab` repo:
- `clusters/main/databases/mariadb/` (5 files: cluster.yaml,
external-secrets.yaml, maxscale.yaml, namespace.yaml, README.md)
- `clusters/main/operators/mariadb-operator.yaml`
- `clusters/main/operators/mariadb-operator-crds.yaml`
2. Committed + pushed to Gitea
3. ArgoCD detected change (`databases``OutOfSync`) but did NOT
auto-prune the operator-managed CRs
4. Manual cleanup required:
- `kubectl delete mariadb mariadb-galera -n mariadb` (CR deletion
triggers operator finalizer processing)
- `kubectl delete maxscale mariadb-maxscale -n mariadb` (already
gone — operator had processed it)
- `kubectl delete namespace mariadb` (clean)
- `kubectl delete namespace mariadb-operator` (stuck Terminating
briefly, cleared after ~15s)
- `kubectl delete crd ...` (9 mariadb CRDs — ArgoCD didn't prune
these, had to delete manually)
5. Forced ArgoCD hard-refresh: `kubectl annotate application root -n
argocd argocd.argoproj.io/refresh-options=hard-refresh --overwrite`
### Result
- Namespaces `mariadb` + `mariadb-operator` deleted
- 9 MariaDB CRDs deleted
- ArgoCD apps `mariadb-operator` + `mariadb-operator-crds` pruned
- `databases` app re-synced (only postgres remains)
- Native Galera VMs (300-302) + MaxScale VM310 untouched
- LB IP 10.0.30.201 freed (was K8s MaxScale, now available in pool)
### Key Lesson
ArgoCD `prune: true` handles standard K8s resources but **does not
reliably prune operator CRs and CRDs** — finalizers require the
operator to be alive, and CRD deletion isn't tracked by ArgoCD's
resource tree. Manual `kubectl delete` of CRs → namespace → CRDs is
required when removing an operator-managed application.
ArgoCD was already installed and configured:
- 8 Applications synced (operators, databases, backups, memory, root)
- Ingress via Traefik (host: `argocd`, routed through Traefik LB IP)
- App-of-Apps pattern pointing to `dominik/iac-homelab` on Gitea
- Auto-sync with `prune: true, selfHeal: true`
Admin password retrieved via:
```bash
kubectl -n argocd get secret argocd-initial-admin-secret \
-o jsonpath='{.data.password}' | base64 -d
```
## IaC Repository
`dominik/iac-homelab` on Gitea (10.0.30.105:3000) contains:
- 7 Epics (management VM, K8s cluster, networking, storage, GitOps, memory, MariaDB)
- OpenTofu for VM provisioning + Helm releases
- Ansible for RKE2 bootstrap
- GitHub Actions CI/CD (self-hosted runner on mgmt-runner VM)
- ArgoCD Application manifests in `clusters/main/apps/`
Separate repo: `dominik/observability-iac` — Loki, Prometheus/Thanos, Grafana