Initial commit: Hermes Agent Skills collection
This commit is contained in:
@@ -0,0 +1,162 @@
|
||||
# RKE2 Cluster Cold-Boot Recovery — 2026-07-12
|
||||
|
||||
## Situation
|
||||
|
||||
RKE2 cluster (6 VMs: 3 CP + 3 Worker) was offline for 111 days. All VMs
|
||||
were stopped across the PVE cluster. User requested cluster startup
|
||||
and K8s evaluation for Docker workload migration.
|
||||
|
||||
## Startup Sequence
|
||||
|
||||
### 1. VM Discovery
|
||||
|
||||
VMs had been HA-migrated to various PVE nodes. Had to scan all nodes
|
||||
to find where each VM landed:
|
||||
|
||||
```
|
||||
118 (CP-01) → proxmox5
|
||||
130 (CP-02) → proxmox3
|
||||
129 (CP-03) → proxmox2
|
||||
128 (Worker-01) → proxmox5
|
||||
132 (Worker-02) → proxmox2
|
||||
131 (Worker-03) → proxmox3
|
||||
200 (mgmt-runner) → proxmox3
|
||||
```
|
||||
|
||||
### 2. Start Order
|
||||
|
||||
1. Started all 6 RKE2 VMs via `qm start` (HA placed them on various nodes)
|
||||
2. CP nodes began etcd convergence within 30s
|
||||
3. Workers initially showed `NotReady` — waiting for
|
||||
`serving-kubelet.crt` retrieval (503 from API server)
|
||||
4. Worker-03 became Ready after ~2 min
|
||||
5. Worker-02 became Ready after ~3 min (slower due to proxy reconnection)
|
||||
6. All 6 nodes Ready within 5 min
|
||||
|
||||
### 3. Convergence Timeline
|
||||
|
||||
| T+0 | VMs started |
|
||||
| T+30s | CPs booting, etcd forming quorum |
|
||||
| T+1min | Workers connecting via wss://CP:9345/v1-rke2/connect |
|
||||
| T+2min | First workers Ready (cert retrieval successful) |
|
||||
| T+3min | All nodes Ready, Cilium agents starting |
|
||||
| T+5min | ArgoCD, External Secrets, operators beginning reconciliation |
|
||||
| T+10min | Most pods Running, old replicas Terminating |
|
||||
| T+15min | Databases recovering (Galera, PostgreSQL) |
|
||||
|
||||
### 4. Pod States After Cold Boot
|
||||
|
||||
Most pods showed mixed states initially:
|
||||
- `Running` — new replicas scheduled by deployment controller
|
||||
- `Terminating` — old replicas from before shutdown (111 days old)
|
||||
- `Unknown` — node was down when status was last reported
|
||||
- `CrashLoopBackOff` — ceph-csi nodeplugin (2 of 6) — Ceph connectivity
|
||||
- `CreateContainerConfigError` — memory-api, qdrant — missing secrets
|
||||
(ExternalSecret hadn't synced yet)
|
||||
|
||||
All resolved naturally except:
|
||||
- 3 ExternalSecrets with SecretSyncedError (1Password item issues)
|
||||
- 2 ceph-csi-rbd-nodeplugin CrashLoopBackOff
|
||||
|
||||
## Cilium LB Pool Migration
|
||||
|
||||
### Problem
|
||||
|
||||
Cilium LB pool was set to `10.0.30.70-10.0.30.89`, overlapping with
|
||||
Galera/MaxScale VIPs:
|
||||
- 10.0.30.70 = Keepalived VIP for MaxScale (Galera)
|
||||
- 10.0.30.71-73 = Galera VM IPs
|
||||
|
||||
K8s services had claimed:
|
||||
- Traefik → 10.0.30.70 (conflicted with MaxScale VIP!)
|
||||
- PostgreSQL RW → 10.0.30.71 (conflicted with Galera VM!)
|
||||
- MariaDB MaxScale (K8s) → 10.0.30.72
|
||||
|
||||
### Fix Applied
|
||||
|
||||
1. Deleted live `CiliumLoadBalancerIPPool` CRD
|
||||
2. Created new pool with `10.0.30.200-10.0.30.250`
|
||||
3. Services auto-reassigned:
|
||||
- Traefik → 10.0.30.202
|
||||
- MariaDB MaxScale (K8s) → 10.0.30.201
|
||||
- PostgreSQL RW → 10.0.30.200
|
||||
4. Updated manifest on all 3 CP nodes:
|
||||
`/var/lib/rancher/rke2/server/manifests/cilium-l2-lb.yaml`
|
||||
5. Fixed IaC repo: `epic-2-k8s/ansible/playbook.yml` — committed + pushed
|
||||
|
||||
### Verification
|
||||
|
||||
- `10.0.30.70` responds to ping (Keepalived VIP for MaxScale — correct!)
|
||||
- No K8s LoadBalancer service holds IPs in .70-.89 range
|
||||
- Cilium LB pool shows 48 IPs available in .200-.250 range
|
||||
|
||||
## K8s MariaDB Removal
|
||||
|
||||
### Decision
|
||||
|
||||
User clarified: Galera lives as **native VMs** (VM300-302 + MaxScale
|
||||
VM310), not in K8s. The K8s MariaDB Galera (3 pods) + MaxScale (2 pods)
|
||||
in namespace `mariadb` was redundant and should be removed.
|
||||
|
||||
### Removal Steps
|
||||
|
||||
1. Removed YAML files from `iac-homelab` repo:
|
||||
- `clusters/main/databases/mariadb/` (5 files: cluster.yaml,
|
||||
external-secrets.yaml, maxscale.yaml, namespace.yaml, README.md)
|
||||
- `clusters/main/operators/mariadb-operator.yaml`
|
||||
- `clusters/main/operators/mariadb-operator-crds.yaml`
|
||||
2. Committed + pushed to Gitea
|
||||
3. ArgoCD detected change (`databases` → `OutOfSync`) but did NOT
|
||||
auto-prune the operator-managed CRs
|
||||
4. Manual cleanup required:
|
||||
- `kubectl delete mariadb mariadb-galera -n mariadb` (CR deletion
|
||||
triggers operator finalizer processing)
|
||||
- `kubectl delete maxscale mariadb-maxscale -n mariadb` (already
|
||||
gone — operator had processed it)
|
||||
- `kubectl delete namespace mariadb` (clean)
|
||||
- `kubectl delete namespace mariadb-operator` (stuck Terminating
|
||||
briefly, cleared after ~15s)
|
||||
- `kubectl delete crd ...` (9 mariadb CRDs — ArgoCD didn't prune
|
||||
these, had to delete manually)
|
||||
5. Forced ArgoCD hard-refresh: `kubectl annotate application root -n
|
||||
argocd argocd.argoproj.io/refresh-options=hard-refresh --overwrite`
|
||||
|
||||
### Result
|
||||
|
||||
- Namespaces `mariadb` + `mariadb-operator` deleted
|
||||
- 9 MariaDB CRDs deleted
|
||||
- ArgoCD apps `mariadb-operator` + `mariadb-operator-crds` pruned
|
||||
- `databases` app re-synced (only postgres remains)
|
||||
- Native Galera VMs (300-302) + MaxScale VM310 untouched
|
||||
- LB IP 10.0.30.201 freed (was K8s MaxScale, now available in pool)
|
||||
|
||||
### Key Lesson
|
||||
|
||||
ArgoCD `prune: true` handles standard K8s resources but **does not
|
||||
reliably prune operator CRs and CRDs** — finalizers require the
|
||||
operator to be alive, and CRD deletion isn't tracked by ArgoCD's
|
||||
resource tree. Manual `kubectl delete` of CRs → namespace → CRDs is
|
||||
required when removing an operator-managed application.
|
||||
|
||||
ArgoCD was already installed and configured:
|
||||
- 8 Applications synced (operators, databases, backups, memory, root)
|
||||
- Ingress via Traefik (host: `argocd`, routed through Traefik LB IP)
|
||||
- App-of-Apps pattern pointing to `dominik/iac-homelab` on Gitea
|
||||
- Auto-sync with `prune: true, selfHeal: true`
|
||||
|
||||
Admin password retrieved via:
|
||||
```bash
|
||||
kubectl -n argocd get secret argocd-initial-admin-secret \
|
||||
-o jsonpath='{.data.password}' | base64 -d
|
||||
```
|
||||
|
||||
## IaC Repository
|
||||
|
||||
`dominik/iac-homelab` on Gitea (10.0.30.105:3000) contains:
|
||||
- 7 Epics (management VM, K8s cluster, networking, storage, GitOps, memory, MariaDB)
|
||||
- OpenTofu for VM provisioning + Helm releases
|
||||
- Ansible for RKE2 bootstrap
|
||||
- GitHub Actions CI/CD (self-hosted runner on mgmt-runner VM)
|
||||
- ArgoCD Application manifests in `clusters/main/apps/`
|
||||
|
||||
Separate repo: `dominik/observability-iac` — Loki, Prometheus/Thanos, Grafana
|
||||
Reference in New Issue
Block a user