198 lines
7.8 KiB
Markdown
198 lines
7.8 KiB
Markdown
# Cluster Analysis — 2026-07-12
|
||
|
||
## Session Context
|
||
|
||
Comprehensive Proxmox + K8s cluster analysis requested by user after completing
|
||
RKE2 v1.35.6 upgrade and Helm chart upgrades. Goal: identify performance
|
||
optimizations, evaluate n5pro for GPU workloads, assess overall health.
|
||
|
||
## Proxmox Cluster State (8 nodes)
|
||
|
||
| Node | CPU | Cores | RAM | Disks | OSDs | VMs/CTs |
|
||
|------|-----|--------|------|-------|------|---------|
|
||
| proxmox1 | i7-8550U | 8T | 15GB | 3.6TB HDD + 954GB NVMe | 1 (HDD) | VM106 (HA) |
|
||
| proxmox2 | i3-9100T | 4C | 15GB | 233GB SSD×2 + 239GB NVMe | 2 (SSD) | CP-03, Worker-02 |
|
||
| proxmox3 | i3-9100T | 4C | 31GB | 233GB SSD + 239GB NVMe | 1 (SSD) | CP-02, Worker-03 |
|
||
| proxmox4 | i3-9100T | 4C | 15GB | 238GB SSD + 239GB NVMe | 1 (SSD) | VM302, VM310 |
|
||
| proxmox5 | i3-9100T | 4C | 15GB | 238GB SSD + 239GB NVMe | 1 (SSD) | CP-01, Worker-01 |
|
||
| proxmox6 | i3-9100T | 4C | 15GB | 931GB HDD + 239GB NVMe | 1 (HDD) | Many CTs |
|
||
| proxmox7 | i3-9100T | 4C | 15GB | 931GB HDD + 239GB NVMe | 1 (HDD) | CTs, VM230 |
|
||
| n5pro | Ryzen AI 9 HX 370 | 24C | 91GB | 128GB NVMe only | 0 | CT104, CT116 |
|
||
|
||
### n5pro Hardware Details
|
||
|
||
- **CPU**: AMD Ryzen AI 9 HX PRO 370 (24C/24T) — most powerful node by far
|
||
- **GPU**: AMD Radeon 890M (**integrated APU graphics** — NOT discrete PCIe)
|
||
- No PCIe slots available (Mini PC form factor)
|
||
- VFIO/GPU passthrough NOT possible
|
||
- SR-IOV not supported on integrated Radeon
|
||
- MIG is NVIDIA-only — not applicable
|
||
- **SATA**: JMB58x AHCI controller present in lspci BUT **no SATA disks connected**
|
||
- Only a 14.3GB SanDisk USB stick (sda) for installation media
|
||
- **NVMe**: 119.2GB (nearly full: OS + 2×30GB WAL/DB LVs)
|
||
- `pve-wal--db--osd-a` (30GB) and `pve-wal--db--osd-b` (30GB) — created but UNUSED
|
||
- No OSDs deployed on n5pro despite WAL/DB preparation
|
||
- **RAM**: 91GB — more than all other nodes combined
|
||
|
||
### CRUSH Map Anomalies
|
||
|
||
Two ghost hosts in CRUSH tree:
|
||
- `px-tmp20` (weight 0) — empty placeholder, no OSDs
|
||
- `ubuntu` (weight 5.49) — hosts osd.6 + osd.8 (2×2.8TB HDD)
|
||
- This is likely an external Ceph contributor node, NOT one of the 8 PVE nodes
|
||
- Hostname doesn't match any PVE node — verify provenance
|
||
|
||
## K8s Cluster State
|
||
|
||
### VM Configurations
|
||
|
||
All 6 K8s VMs have identical config:
|
||
- 4 vCPU, 12GB RAM, `balloon: 8192`, `cpu: host`, `numa: 0`
|
||
- `aio=io_uring`, `cache=none`, `discard=on`, `iothread=1`
|
||
- CPs on `vm_disks` pool (SSD tier), Workers on `hdd_disk` pool (HDD tier)
|
||
- All have HA configured (no HA groups defined)
|
||
|
||
### Resource Utilization (very low)
|
||
|
||
| Node | CPU | CPU% | Memory | Mem% |
|
||
|------|-----|------|--------|------|
|
||
| CP-01 | 194m | 4% | 3387Mi | 43% |
|
||
| CP-02 | 250m | 6% | 4398Mi | 48% |
|
||
| CP-03 | 210m | 5% | 3643Mi | 46% |
|
||
| Worker-01 | 58m | 1% | 2652Mi | 33% |
|
||
| Worker-02 | 78m | 1% | 2830Mi | 35% |
|
||
| Worker-03 | 70m | 1% | 2470Mi | 26% |
|
||
|
||
Top CPU consumers: kube-apiserver (84m × 3), etcd (36-58m × 3), Cilium (27-31m × 6).
|
||
|
||
### Workloads (light)
|
||
|
||
- ArgoCD (7 apps, 6 Healthy/Synced, `backups` OutOfSync, `hindsight` OutOfSync)
|
||
- CloudNativePG (postgres-main, 3 instances)
|
||
- External Secrets Operator
|
||
- Velero (backups + node-agent)
|
||
- Ceph CSI RBD (provisioner + nodeplugin)
|
||
- Hindsight (API + postgres)
|
||
- Cilium, CoreDNS, Traefik (RKE2-bundled)
|
||
|
||
## Ceph Critical Issues (July 12)
|
||
|
||
### Recurring osd.10 Near-Full Problem
|
||
|
||
Same issue as July 4 (see `references/ceph-pool-full-recovery-2026-07.md`)
|
||
but now worse:
|
||
|
||
| Metric | July 4 | July 12 |
|
||
|--------|---------|----------|
|
||
| osd.10 %USE | 95% (full_ratio) | **96.72%** |
|
||
| hdd_disk pool %USED | 71% | **99.09%** |
|
||
| Pools backfillfull | 6 | **7** |
|
||
| PGs backfill_toofull | 0 | **4** |
|
||
| BlueFS spillover | 0 OSDs | **1 (osd.8)** |
|
||
| Slow ops | 4 OSDs | **5 OSDs** |
|
||
| Ceph version mix | uniform | 5 OSDs on 19.2.3, rest on 19.2.4 |
|
||
|
||
**Root cause unchanged**: osd.10 (982GB HDD on proxmox6) is nearly full,
|
||
blocking backfill for 4 PGs. The `hdd_disk` pool (where Worker VM disks
|
||
reside) is at 99.09% — essentially no growth capacity.
|
||
|
||
### Worker VM Disk Placement Problem
|
||
|
||
Worker-01/02/03 disks are on `hdd_disk` pool (HDD tier, 99% full).
|
||
CP-01/02/03 disks are on `vm_disks` pool (SSD tier, 74.5% full).
|
||
Workers should be on SSD tier for I/O performance, but SSD pool only
|
||
has 87GB MAX AVAIL remaining.
|
||
|
||
### OSD Imbalance (VAR 0.65-2.33, STDDEV 27.69%)
|
||
|
||
| OSD | Size | %USE | VAR | PGs | Issue |
|
||
|-----|------|------|-----|-----|-------|
|
||
| osd.10 | 982GB | 96.72% | 2.33 | 283 | Nearly full, blocking backfill |
|
||
| osd.4 | 233GB | 73.90% | 1.78 | 119 | SSD, slow ops |
|
||
| osd.3 | 238GB | 71.53% | 1.73 | 114 | SSD, slow ops |
|
||
| osd.5 | 238GB | 71.33% | 1.72 | 113 | SSD, slow ops |
|
||
| osd.2 | 233GB | 63.73% | 1.54 | 97 | SSD, slow ops |
|
||
| osd.7 | 982GB | 63.82% | 1.54 | 215 | HDD |
|
||
| osd.0 | 188GB | 57.61% | 1.39 | 71 | SSD, slow ops |
|
||
| osd.1 | 3.6TB | 34.53% | 0.83 | 437 | HDD, slow ops |
|
||
| osd.6 | 2.8TB | 27.10% | 0.65 | 304 | HDD (ubuntu host) |
|
||
| osd.8 | 2.8TB | 27.25% | 0.66 | 301 | HDD, BlueFS spillover |
|
||
|
||
## Performance Optimization Recommendations
|
||
|
||
### Priority 1: Acute Fixes
|
||
|
||
1. **Drain osd.10** — `ceph osd reweight 10 0.5` or lower. Move data to
|
||
osd.6/8 (ubuntu host, 2×2.8TB, only 27% full).
|
||
2. **Migrate Worker VM disks** from `hdd_disk` (99% full, HDD) to
|
||
`vm_disks` (SSD) via PVE storage live migration. Need to free SSD
|
||
pool space first or add SSD capacity.
|
||
3. **Upgrade 5 OSDs** from Ceph 19.2.3 → 19.2.4 (eliminate version mix).
|
||
4. **Clean CRUSH map** — remove `px-tmp20`, investigate `ubuntu` host.
|
||
|
||
### Priority 2: K8s Optimizations
|
||
|
||
5. **Disable memory ballooning** on K8s VMs — set `balloon: 0`.
|
||
Ballooning causes non-deterministic latency spikes.
|
||
6. **Node-Local DNS Cache** — deploy `node-local-dns` DaemonSet.
|
||
7. **Kubelet reserved resources** — verify `systemReserved`/`kubeReserved`.
|
||
8. **Pod topology spread** — for multi-replica deployments.
|
||
|
||
### Priority 3: Infrastructure
|
||
|
||
9. **Jumbo frames (MTU 9000)** — reduce Ceph network overhead.
|
||
10. **CPU pinning** — i3 nodes have 4 cores, 2 VMs each = 100% overcommit.
|
||
11. **Utilize n5pro** — 24 cores + 91GB RAM idle. Place worker VM there.
|
||
12. **HA group with n5pro** — define HA group so VMs can failover to n5pro.
|
||
|
||
## n5pro GPU Evaluation
|
||
|
||
**Conclusion: GPU passthrough NOT possible.**
|
||
|
||
- Radeon 890M is integrated APU graphics (shares system RAM)
|
||
- No discrete PCIe GPU, no PCIe slots for expansion
|
||
- SR-IOV not supported on consumer integrated GPUs
|
||
- MIG is NVIDIA-only
|
||
|
||
**Alternative uses for n5pro:**
|
||
- CPU-pinned worker VM (excellent: 24C, 91GB RAM)
|
||
- Ceph OSD node (JMB58x SATA + 5 disks → new OSDs, WAL/DB LVs ready)
|
||
- vLLM CPU inference (possible but slow vs GPU)
|
||
- General compute workloads (best CPU/RAM in cluster)
|
||
|
||
## Failover Analysis
|
||
|
||
- All 6 K8s VMs have HA configured ✅
|
||
- No HA groups defined ⚠️
|
||
- i3 nodes: 4 cores each, 2 K8s VMs per node = 8 vCPUs on 4 physical cores
|
||
- If proxmox5 fails (CP-01 + Worker-01), no other i3 node can absorb 8 vCPUs
|
||
- n5pro (24C, 91GB) could host all 6 VMs simultaneously — ideal HA fallback
|
||
- **Recommendation**: Define HA group `[proxmox2, proxmox3, proxmox5, n5pro]`
|
||
with n5pro as last-resort fallback
|
||
|
||
## Data Collection Methodology
|
||
|
||
### Reliable VM Inventory (avoids SSH quoting issues)
|
||
|
||
```bash
|
||
# From PVE coordinator (10.0.20.10):
|
||
pvesh get /cluster/resources --type vm --output-format json
|
||
# Returns JSON array with vmid, name, node, status, maxmem, maxcpu, maxdisk
|
||
# Parse with python3 -c "import sys,json; ..."
|
||
```
|
||
|
||
### SSH Chain Quoting Pitfall
|
||
|
||
Nested SSH commands with `grep` and `awk` break due to quote escaping
|
||
through multiple SSH layers. Solutions:
|
||
1. Use `pvesh get /cluster/resources --type vm --output-format json` (API)
|
||
2. Use heredoc: `ssh root@host 'bash -s' << 'SCRIPT' ... SCRIPT`
|
||
3. Use `pvesh` API calls instead of parsing CLI output
|
||
|
||
### Direct SSH to n5pro
|
||
|
||
n5pro is NOT reachable by hostname from proxmox1 — use IP:
|
||
```bash
|
||
ssh -o StrictHostKeyChecking=no 10.0.20.91 'commands'
|
||
```
|