# Cluster Analysis — 2026-07-12 ## Session Context Comprehensive Proxmox + K8s cluster analysis requested by user after completing RKE2 v1.35.6 upgrade and Helm chart upgrades. Goal: identify performance optimizations, evaluate n5pro for GPU workloads, assess overall health. ## Proxmox Cluster State (8 nodes) | Node | CPU | Cores | RAM | Disks | OSDs | VMs/CTs | |------|-----|--------|------|-------|------|---------| | proxmox1 | i7-8550U | 8T | 15GB | 3.6TB HDD + 954GB NVMe | 1 (HDD) | VM106 (HA) | | proxmox2 | i3-9100T | 4C | 15GB | 233GB SSD×2 + 239GB NVMe | 2 (SSD) | CP-03, Worker-02 | | proxmox3 | i3-9100T | 4C | 31GB | 233GB SSD + 239GB NVMe | 1 (SSD) | CP-02, Worker-03 | | proxmox4 | i3-9100T | 4C | 15GB | 238GB SSD + 239GB NVMe | 1 (SSD) | VM302, VM310 | | proxmox5 | i3-9100T | 4C | 15GB | 238GB SSD + 239GB NVMe | 1 (SSD) | CP-01, Worker-01 | | proxmox6 | i3-9100T | 4C | 15GB | 931GB HDD + 239GB NVMe | 1 (HDD) | Many CTs | | proxmox7 | i3-9100T | 4C | 15GB | 931GB HDD + 239GB NVMe | 1 (HDD) | CTs, VM230 | | n5pro | Ryzen AI 9 HX 370 | 24C | 91GB | 128GB NVMe only | 0 | CT104, CT116 | ### n5pro Hardware Details - **CPU**: AMD Ryzen AI 9 HX PRO 370 (24C/24T) — most powerful node by far - **GPU**: AMD Radeon 890M (**integrated APU graphics** — NOT discrete PCIe) - No PCIe slots available (Mini PC form factor) - VFIO/GPU passthrough NOT possible - SR-IOV not supported on integrated Radeon - MIG is NVIDIA-only — not applicable - **SATA**: JMB58x AHCI controller present in lspci BUT **no SATA disks connected** - Only a 14.3GB SanDisk USB stick (sda) for installation media - **NVMe**: 119.2GB (nearly full: OS + 2×30GB WAL/DB LVs) - `pve-wal--db--osd-a` (30GB) and `pve-wal--db--osd-b` (30GB) — created but UNUSED - No OSDs deployed on n5pro despite WAL/DB preparation - **RAM**: 91GB — more than all other nodes combined ### CRUSH Map Anomalies Two ghost hosts in CRUSH tree: - `px-tmp20` (weight 0) — empty placeholder, no OSDs - `ubuntu` (weight 5.49) — hosts osd.6 + osd.8 (2×2.8TB HDD) - This is likely an external Ceph contributor node, NOT one of the 8 PVE nodes - Hostname doesn't match any PVE node — verify provenance ## K8s Cluster State ### VM Configurations All 6 K8s VMs have identical config: - 4 vCPU, 12GB RAM, `balloon: 8192`, `cpu: host`, `numa: 0` - `aio=io_uring`, `cache=none`, `discard=on`, `iothread=1` - CPs on `vm_disks` pool (SSD tier), Workers on `hdd_disk` pool (HDD tier) - All have HA configured (no HA groups defined) ### Resource Utilization (very low) | Node | CPU | CPU% | Memory | Mem% | |------|-----|------|--------|------| | CP-01 | 194m | 4% | 3387Mi | 43% | | CP-02 | 250m | 6% | 4398Mi | 48% | | CP-03 | 210m | 5% | 3643Mi | 46% | | Worker-01 | 58m | 1% | 2652Mi | 33% | | Worker-02 | 78m | 1% | 2830Mi | 35% | | Worker-03 | 70m | 1% | 2470Mi | 26% | Top CPU consumers: kube-apiserver (84m × 3), etcd (36-58m × 3), Cilium (27-31m × 6). ### Workloads (light) - ArgoCD (7 apps, 6 Healthy/Synced, `backups` OutOfSync, `hindsight` OutOfSync) - CloudNativePG (postgres-main, 3 instances) - External Secrets Operator - Velero (backups + node-agent) - Ceph CSI RBD (provisioner + nodeplugin) - Hindsight (API + postgres) - Cilium, CoreDNS, Traefik (RKE2-bundled) ## Ceph Critical Issues (July 12) ### Recurring osd.10 Near-Full Problem Same issue as July 4 (see `references/ceph-pool-full-recovery-2026-07.md`) but now worse: | Metric | July 4 | July 12 | |--------|---------|----------| | osd.10 %USE | 95% (full_ratio) | **96.72%** | | hdd_disk pool %USED | 71% | **99.09%** | | Pools backfillfull | 6 | **7** | | PGs backfill_toofull | 0 | **4** | | BlueFS spillover | 0 OSDs | **1 (osd.8)** | | Slow ops | 4 OSDs | **5 OSDs** | | Ceph version mix | uniform | 5 OSDs on 19.2.3, rest on 19.2.4 | **Root cause unchanged**: osd.10 (982GB HDD on proxmox6) is nearly full, blocking backfill for 4 PGs. The `hdd_disk` pool (where Worker VM disks reside) is at 99.09% — essentially no growth capacity. ### Worker VM Disk Placement Problem Worker-01/02/03 disks are on `hdd_disk` pool (HDD tier, 99% full). CP-01/02/03 disks are on `vm_disks` pool (SSD tier, 74.5% full). Workers should be on SSD tier for I/O performance, but SSD pool only has 87GB MAX AVAIL remaining. ### OSD Imbalance (VAR 0.65-2.33, STDDEV 27.69%) | OSD | Size | %USE | VAR | PGs | Issue | |-----|------|------|-----|-----|-------| | osd.10 | 982GB | 96.72% | 2.33 | 283 | Nearly full, blocking backfill | | osd.4 | 233GB | 73.90% | 1.78 | 119 | SSD, slow ops | | osd.3 | 238GB | 71.53% | 1.73 | 114 | SSD, slow ops | | osd.5 | 238GB | 71.33% | 1.72 | 113 | SSD, slow ops | | osd.2 | 233GB | 63.73% | 1.54 | 97 | SSD, slow ops | | osd.7 | 982GB | 63.82% | 1.54 | 215 | HDD | | osd.0 | 188GB | 57.61% | 1.39 | 71 | SSD, slow ops | | osd.1 | 3.6TB | 34.53% | 0.83 | 437 | HDD, slow ops | | osd.6 | 2.8TB | 27.10% | 0.65 | 304 | HDD (ubuntu host) | | osd.8 | 2.8TB | 27.25% | 0.66 | 301 | HDD, BlueFS spillover | ## Performance Optimization Recommendations ### Priority 1: Acute Fixes 1. **Drain osd.10** — `ceph osd reweight 10 0.5` or lower. Move data to osd.6/8 (ubuntu host, 2×2.8TB, only 27% full). 2. **Migrate Worker VM disks** from `hdd_disk` (99% full, HDD) to `vm_disks` (SSD) via PVE storage live migration. Need to free SSD pool space first or add SSD capacity. 3. **Upgrade 5 OSDs** from Ceph 19.2.3 → 19.2.4 (eliminate version mix). 4. **Clean CRUSH map** — remove `px-tmp20`, investigate `ubuntu` host. ### Priority 2: K8s Optimizations 5. **Disable memory ballooning** on K8s VMs — set `balloon: 0`. Ballooning causes non-deterministic latency spikes. 6. **Node-Local DNS Cache** — deploy `node-local-dns` DaemonSet. 7. **Kubelet reserved resources** — verify `systemReserved`/`kubeReserved`. 8. **Pod topology spread** — for multi-replica deployments. ### Priority 3: Infrastructure 9. **Jumbo frames (MTU 9000)** — reduce Ceph network overhead. 10. **CPU pinning** — i3 nodes have 4 cores, 2 VMs each = 100% overcommit. 11. **Utilize n5pro** — 24 cores + 91GB RAM idle. Place worker VM there. 12. **HA group with n5pro** — define HA group so VMs can failover to n5pro. ## n5pro GPU Evaluation **Conclusion: GPU passthrough NOT possible.** - Radeon 890M is integrated APU graphics (shares system RAM) - No discrete PCIe GPU, no PCIe slots for expansion - SR-IOV not supported on consumer integrated GPUs - MIG is NVIDIA-only **Alternative uses for n5pro:** - CPU-pinned worker VM (excellent: 24C, 91GB RAM) - Ceph OSD node (JMB58x SATA + 5 disks → new OSDs, WAL/DB LVs ready) - vLLM CPU inference (possible but slow vs GPU) - General compute workloads (best CPU/RAM in cluster) ## Failover Analysis - All 6 K8s VMs have HA configured ✅ - No HA groups defined ⚠️ - i3 nodes: 4 cores each, 2 K8s VMs per node = 8 vCPUs on 4 physical cores - If proxmox5 fails (CP-01 + Worker-01), no other i3 node can absorb 8 vCPUs - n5pro (24C, 91GB) could host all 6 VMs simultaneously — ideal HA fallback - **Recommendation**: Define HA group `[proxmox2, proxmox3, proxmox5, n5pro]` with n5pro as last-resort fallback ## Data Collection Methodology ### Reliable VM Inventory (avoids SSH quoting issues) ```bash # From PVE coordinator (10.0.20.10): pvesh get /cluster/resources --type vm --output-format json # Returns JSON array with vmid, name, node, status, maxmem, maxcpu, maxdisk # Parse with python3 -c "import sys,json; ..." ``` ### SSH Chain Quoting Pitfall Nested SSH commands with `grep` and `awk` break due to quote escaping through multiple SSH layers. Solutions: 1. Use `pvesh get /cluster/resources --type vm --output-format json` (API) 2. Use heredoc: `ssh root@host 'bash -s' << 'SCRIPT' ... SCRIPT` 3. Use `pvesh` API calls instead of parsing CLI output ### Direct SSH to n5pro n5pro is NOT reachable by hostname from proxmox1 — use IP: ```bash ssh -o StrictHostKeyChecking=no 10.0.20.91 'commands' ``` ### `pvesm list` Fails During Ceph Backfillfull (2026-07-12) When Ceph pool `hdd_disk` is in `backfillfull` state (osd.10 at 96.72%), `pvesm list hdd_disk` returns empty with error: ``` rbd error: rbd: listing images failed: (2) No such file or directory ``` Similarly `pvesh get /nodes/{node}/storage/hdd_disk/content` fails. **Fallback**: Use `rbd ls hdd_disk` directly (works even when PVE storage layer cannot enumerate). Cross-reference with `pvesh get /cluster/resources --type vm` for orphan detection. ### HA Rule Created: k8s-prefer-n5pro (2026-07-12) Successfully created node-affinity rule for K8s VMs: - Rule name: `k8s-prefer-n5pro` - Type: `node-affinity` - Resources: `vm:118,vm:128,vm:129,vm:130,vm:131,vm:132` - Nodes: `n5pro,proxmox5,proxmox3,proxmox2` (ordered preference) - All 6 VMs briefly went through "deleting" → "starting" → "started" transition - Existing rules untouched (ct:104/ct:99999 → n5pro, vm:310/311 + vm:301/302 negative affinity) See `references/ha-manager-and-custom-dashboards-2026-07.md` for full syntax. ## K8s VM Anti-Affinity Problem (2026-07-12) ### Current Distribution (BAD) | Node | CP | Worker | Risk | |------|----|--------|------| | proxmox2 | cp-03 (129) | worker-02 (132) | 2 VMs lost on node failure | | proxmox3 | cp-02 (130) | worker-03 (131) | 2 VMs lost on node failure | | proxmox5 | cp-01 (118) | worker-01 (128) | 2 VMs lost on node failure | Each of the 3 K8s nodes hosts BOTH a CP and a Worker. A single node failure drops 1 CP + 1 Worker simultaneously. Quorum survives (2/3 CPs) but workload capacity drops 33%. ### K8s VM Anti-Colocation: EXECUTED (2026-07-12) ✅ The proposed distribution was successfully implemented via live migration. All 3 worker VMs migrated to separate nodes (no node hosts >1 K8s VM): | Node | CP | Worker | |------|----|--------| | proxmox1 | — | worker-02 (132) ← migrated FROM proxmox2 | | proxmox2 | cp-03 (129) | — | | proxmox3 | cp-02 (130) | — | | proxmox4 | — | worker-03 (131) ← migrated FROM proxmox3 | | proxmox5 | cp-01 (118) | — | | n5pro | — | worker-01 (128) ← migrated FROM proxmox5 | **Migration commands** (ran from source node, not coordinator): ```bash ssh root@10.0.20.20 'qm migrate 132 proxmox1 --online' # proxmox2 → proxmox1 ssh root@10.0.20.30 'qm migrate 131 proxmox4 --online' # proxmox3 → proxmox4 ssh root@10.0.20.50 'qm migrate 128 n5pro --online' # proxmox5 → n5pro ``` HA rule `k8s-prefer-n5pro` was expanded to include ALL 8 nodes before migration (was missing proxmox1, proxmox4 — caused migration failure). See `references/ha-manager-and-custom-dashboards-2026-07.md` for details. Result: Any single node failure now loses at most 1 K8s VM (was 2). Quorum survives any single CP node failure (2/3 CPs remain).