163 lines
5.9 KiB
Markdown
163 lines
5.9 KiB
Markdown
# RKE2 Cluster Cold-Boot Recovery — 2026-07-12
|
|
|
|
## Situation
|
|
|
|
RKE2 cluster (6 VMs: 3 CP + 3 Worker) was offline for 111 days. All VMs
|
|
were stopped across the PVE cluster. User requested cluster startup
|
|
and K8s evaluation for Docker workload migration.
|
|
|
|
## Startup Sequence
|
|
|
|
### 1. VM Discovery
|
|
|
|
VMs had been HA-migrated to various PVE nodes. Had to scan all nodes
|
|
to find where each VM landed:
|
|
|
|
```
|
|
118 (CP-01) → proxmox5
|
|
130 (CP-02) → proxmox3
|
|
129 (CP-03) → proxmox2
|
|
128 (Worker-01) → proxmox5
|
|
132 (Worker-02) → proxmox2
|
|
131 (Worker-03) → proxmox3
|
|
200 (mgmt-runner) → proxmox3
|
|
```
|
|
|
|
### 2. Start Order
|
|
|
|
1. Started all 6 RKE2 VMs via `qm start` (HA placed them on various nodes)
|
|
2. CP nodes began etcd convergence within 30s
|
|
3. Workers initially showed `NotReady` — waiting for
|
|
`serving-kubelet.crt` retrieval (503 from API server)
|
|
4. Worker-03 became Ready after ~2 min
|
|
5. Worker-02 became Ready after ~3 min (slower due to proxy reconnection)
|
|
6. All 6 nodes Ready within 5 min
|
|
|
|
### 3. Convergence Timeline
|
|
|
|
| T+0 | VMs started |
|
|
| T+30s | CPs booting, etcd forming quorum |
|
|
| T+1min | Workers connecting via wss://CP:9345/v1-rke2/connect |
|
|
| T+2min | First workers Ready (cert retrieval successful) |
|
|
| T+3min | All nodes Ready, Cilium agents starting |
|
|
| T+5min | ArgoCD, External Secrets, operators beginning reconciliation |
|
|
| T+10min | Most pods Running, old replicas Terminating |
|
|
| T+15min | Databases recovering (Galera, PostgreSQL) |
|
|
|
|
### 4. Pod States After Cold Boot
|
|
|
|
Most pods showed mixed states initially:
|
|
- `Running` — new replicas scheduled by deployment controller
|
|
- `Terminating` — old replicas from before shutdown (111 days old)
|
|
- `Unknown` — node was down when status was last reported
|
|
- `CrashLoopBackOff` — ceph-csi nodeplugin (2 of 6) — Ceph connectivity
|
|
- `CreateContainerConfigError` — memory-api, qdrant — missing secrets
|
|
(ExternalSecret hadn't synced yet)
|
|
|
|
All resolved naturally except:
|
|
- 3 ExternalSecrets with SecretSyncedError (1Password item issues)
|
|
- 2 ceph-csi-rbd-nodeplugin CrashLoopBackOff
|
|
|
|
## Cilium LB Pool Migration
|
|
|
|
### Problem
|
|
|
|
Cilium LB pool was set to `10.0.30.70-10.0.30.89`, overlapping with
|
|
Galera/MaxScale VIPs:
|
|
- 10.0.30.70 = Keepalived VIP for MaxScale (Galera)
|
|
- 10.0.30.71-73 = Galera VM IPs
|
|
|
|
K8s services had claimed:
|
|
- Traefik → 10.0.30.70 (conflicted with MaxScale VIP!)
|
|
- PostgreSQL RW → 10.0.30.71 (conflicted with Galera VM!)
|
|
- MariaDB MaxScale (K8s) → 10.0.30.72
|
|
|
|
### Fix Applied
|
|
|
|
1. Deleted live `CiliumLoadBalancerIPPool` CRD
|
|
2. Created new pool with `10.0.30.200-10.0.30.250`
|
|
3. Services auto-reassigned:
|
|
- Traefik → 10.0.30.202
|
|
- MariaDB MaxScale (K8s) → 10.0.30.201
|
|
- PostgreSQL RW → 10.0.30.200
|
|
4. Updated manifest on all 3 CP nodes:
|
|
`/var/lib/rancher/rke2/server/manifests/cilium-l2-lb.yaml`
|
|
5. Fixed IaC repo: `epic-2-k8s/ansible/playbook.yml` — committed + pushed
|
|
|
|
### Verification
|
|
|
|
- `10.0.30.70` responds to ping (Keepalived VIP for MaxScale — correct!)
|
|
- No K8s LoadBalancer service holds IPs in .70-.89 range
|
|
- Cilium LB pool shows 48 IPs available in .200-.250 range
|
|
|
|
## K8s MariaDB Removal
|
|
|
|
### Decision
|
|
|
|
User clarified: Galera lives as **native VMs** (VM300-302 + MaxScale
|
|
VM310), not in K8s. The K8s MariaDB Galera (3 pods) + MaxScale (2 pods)
|
|
in namespace `mariadb` was redundant and should be removed.
|
|
|
|
### Removal Steps
|
|
|
|
1. Removed YAML files from `iac-homelab` repo:
|
|
- `clusters/main/databases/mariadb/` (5 files: cluster.yaml,
|
|
external-secrets.yaml, maxscale.yaml, namespace.yaml, README.md)
|
|
- `clusters/main/operators/mariadb-operator.yaml`
|
|
- `clusters/main/operators/mariadb-operator-crds.yaml`
|
|
2. Committed + pushed to Gitea
|
|
3. ArgoCD detected change (`databases` → `OutOfSync`) but did NOT
|
|
auto-prune the operator-managed CRs
|
|
4. Manual cleanup required:
|
|
- `kubectl delete mariadb mariadb-galera -n mariadb` (CR deletion
|
|
triggers operator finalizer processing)
|
|
- `kubectl delete maxscale mariadb-maxscale -n mariadb` (already
|
|
gone — operator had processed it)
|
|
- `kubectl delete namespace mariadb` (clean)
|
|
- `kubectl delete namespace mariadb-operator` (stuck Terminating
|
|
briefly, cleared after ~15s)
|
|
- `kubectl delete crd ...` (9 mariadb CRDs — ArgoCD didn't prune
|
|
these, had to delete manually)
|
|
5. Forced ArgoCD hard-refresh: `kubectl annotate application root -n
|
|
argocd argocd.argoproj.io/refresh-options=hard-refresh --overwrite`
|
|
|
|
### Result
|
|
|
|
- Namespaces `mariadb` + `mariadb-operator` deleted
|
|
- 9 MariaDB CRDs deleted
|
|
- ArgoCD apps `mariadb-operator` + `mariadb-operator-crds` pruned
|
|
- `databases` app re-synced (only postgres remains)
|
|
- Native Galera VMs (300-302) + MaxScale VM310 untouched
|
|
- LB IP 10.0.30.201 freed (was K8s MaxScale, now available in pool)
|
|
|
|
### Key Lesson
|
|
|
|
ArgoCD `prune: true` handles standard K8s resources but **does not
|
|
reliably prune operator CRs and CRDs** — finalizers require the
|
|
operator to be alive, and CRD deletion isn't tracked by ArgoCD's
|
|
resource tree. Manual `kubectl delete` of CRs → namespace → CRDs is
|
|
required when removing an operator-managed application.
|
|
|
|
ArgoCD was already installed and configured:
|
|
- 8 Applications synced (operators, databases, backups, memory, root)
|
|
- Ingress via Traefik (host: `argocd`, routed through Traefik LB IP)
|
|
- App-of-Apps pattern pointing to `dominik/iac-homelab` on Gitea
|
|
- Auto-sync with `prune: true, selfHeal: true`
|
|
|
|
Admin password retrieved via:
|
|
```bash
|
|
kubectl -n argocd get secret argocd-initial-admin-secret \
|
|
-o jsonpath='{.data.password}' | base64 -d
|
|
```
|
|
|
|
## IaC Repository
|
|
|
|
`dominik/iac-homelab` on Gitea (10.0.30.105:3000) contains:
|
|
- 7 Epics (management VM, K8s cluster, networking, storage, GitOps, memory, MariaDB)
|
|
- OpenTofu for VM provisioning + Helm releases
|
|
- Ansible for RKE2 bootstrap
|
|
- GitHub Actions CI/CD (self-hosted runner on mgmt-runner VM)
|
|
- ArgoCD Application manifests in `clusters/main/apps/`
|
|
|
|
Separate repo: `dominik/observability-iac` — Loki, Prometheus/Thanos, Grafana
|