Files
hermes-skills/devops/proxmox-ve-administration/references/cluster-analysis-2026-07-12.md
T

7.8 KiB
Raw Blame History

Cluster Analysis — 2026-07-12

Session Context

Comprehensive Proxmox + K8s cluster analysis requested by user after completing RKE2 v1.35.6 upgrade and Helm chart upgrades. Goal: identify performance optimizations, evaluate n5pro for GPU workloads, assess overall health.

Proxmox Cluster State (8 nodes)

Node CPU Cores RAM Disks OSDs VMs/CTs
proxmox1 i7-8550U 8T 15GB 3.6TB HDD + 954GB NVMe 1 (HDD) VM106 (HA)
proxmox2 i3-9100T 4C 15GB 233GB SSD×2 + 239GB NVMe 2 (SSD) CP-03, Worker-02
proxmox3 i3-9100T 4C 31GB 233GB SSD + 239GB NVMe 1 (SSD) CP-02, Worker-03
proxmox4 i3-9100T 4C 15GB 238GB SSD + 239GB NVMe 1 (SSD) VM302, VM310
proxmox5 i3-9100T 4C 15GB 238GB SSD + 239GB NVMe 1 (SSD) CP-01, Worker-01
proxmox6 i3-9100T 4C 15GB 931GB HDD + 239GB NVMe 1 (HDD) Many CTs
proxmox7 i3-9100T 4C 15GB 931GB HDD + 239GB NVMe 1 (HDD) CTs, VM230
n5pro Ryzen AI 9 HX 370 24C 91GB 128GB NVMe only 0 CT104, CT116

n5pro Hardware Details

  • CPU: AMD Ryzen AI 9 HX PRO 370 (24C/24T) — most powerful node by far
  • GPU: AMD Radeon 890M (integrated APU graphics — NOT discrete PCIe)
    • No PCIe slots available (Mini PC form factor)
    • VFIO/GPU passthrough NOT possible
    • SR-IOV not supported on integrated Radeon
    • MIG is NVIDIA-only — not applicable
  • SATA: JMB58x AHCI controller present in lspci BUT no SATA disks connected
    • Only a 14.3GB SanDisk USB stick (sda) for installation media
  • NVMe: 119.2GB (nearly full: OS + 2×30GB WAL/DB LVs)
    • pve-wal--db--osd-a (30GB) and pve-wal--db--osd-b (30GB) — created but UNUSED
    • No OSDs deployed on n5pro despite WAL/DB preparation
  • RAM: 91GB — more than all other nodes combined

CRUSH Map Anomalies

Two ghost hosts in CRUSH tree:

  • px-tmp20 (weight 0) — empty placeholder, no OSDs
  • ubuntu (weight 5.49) — hosts osd.6 + osd.8 (2×2.8TB HDD)
    • This is likely an external Ceph contributor node, NOT one of the 8 PVE nodes
    • Hostname doesn't match any PVE node — verify provenance

K8s Cluster State

VM Configurations

All 6 K8s VMs have identical config:

  • 4 vCPU, 12GB RAM, balloon: 8192, cpu: host, numa: 0
  • aio=io_uring, cache=none, discard=on, iothread=1
  • CPs on vm_disks pool (SSD tier), Workers on hdd_disk pool (HDD tier)
  • All have HA configured (no HA groups defined)

Resource Utilization (very low)

Node CPU CPU% Memory Mem%
CP-01 194m 4% 3387Mi 43%
CP-02 250m 6% 4398Mi 48%
CP-03 210m 5% 3643Mi 46%
Worker-01 58m 1% 2652Mi 33%
Worker-02 78m 1% 2830Mi 35%
Worker-03 70m 1% 2470Mi 26%

Top CPU consumers: kube-apiserver (84m × 3), etcd (36-58m × 3), Cilium (27-31m × 6).

Workloads (light)

  • ArgoCD (7 apps, 6 Healthy/Synced, backups OutOfSync, hindsight OutOfSync)
  • CloudNativePG (postgres-main, 3 instances)
  • External Secrets Operator
  • Velero (backups + node-agent)
  • Ceph CSI RBD (provisioner + nodeplugin)
  • Hindsight (API + postgres)
  • Cilium, CoreDNS, Traefik (RKE2-bundled)

Ceph Critical Issues (July 12)

Recurring osd.10 Near-Full Problem

Same issue as July 4 (see references/ceph-pool-full-recovery-2026-07.md) but now worse:

Metric July 4 July 12
osd.10 %USE 95% (full_ratio) 96.72%
hdd_disk pool %USED 71% 99.09%
Pools backfillfull 6 7
PGs backfill_toofull 0 4
BlueFS spillover 0 OSDs 1 (osd.8)
Slow ops 4 OSDs 5 OSDs
Ceph version mix uniform 5 OSDs on 19.2.3, rest on 19.2.4

Root cause unchanged: osd.10 (982GB HDD on proxmox6) is nearly full, blocking backfill for 4 PGs. The hdd_disk pool (where Worker VM disks reside) is at 99.09% — essentially no growth capacity.

Worker VM Disk Placement Problem

Worker-01/02/03 disks are on hdd_disk pool (HDD tier, 99% full). CP-01/02/03 disks are on vm_disks pool (SSD tier, 74.5% full). Workers should be on SSD tier for I/O performance, but SSD pool only has 87GB MAX AVAIL remaining.

OSD Imbalance (VAR 0.65-2.33, STDDEV 27.69%)

OSD Size %USE VAR PGs Issue
osd.10 982GB 96.72% 2.33 283 Nearly full, blocking backfill
osd.4 233GB 73.90% 1.78 119 SSD, slow ops
osd.3 238GB 71.53% 1.73 114 SSD, slow ops
osd.5 238GB 71.33% 1.72 113 SSD, slow ops
osd.2 233GB 63.73% 1.54 97 SSD, slow ops
osd.7 982GB 63.82% 1.54 215 HDD
osd.0 188GB 57.61% 1.39 71 SSD, slow ops
osd.1 3.6TB 34.53% 0.83 437 HDD, slow ops
osd.6 2.8TB 27.10% 0.65 304 HDD (ubuntu host)
osd.8 2.8TB 27.25% 0.66 301 HDD, BlueFS spillover

Performance Optimization Recommendations

Priority 1: Acute Fixes

  1. Drain osd.10ceph osd reweight 10 0.5 or lower. Move data to osd.6/8 (ubuntu host, 2×2.8TB, only 27% full).
  2. Migrate Worker VM disks from hdd_disk (99% full, HDD) to vm_disks (SSD) via PVE storage live migration. Need to free SSD pool space first or add SSD capacity.
  3. Upgrade 5 OSDs from Ceph 19.2.3 → 19.2.4 (eliminate version mix).
  4. Clean CRUSH map — remove px-tmp20, investigate ubuntu host.

Priority 2: K8s Optimizations

  1. Disable memory ballooning on K8s VMs — set balloon: 0. Ballooning causes non-deterministic latency spikes.
  2. Node-Local DNS Cache — deploy node-local-dns DaemonSet.
  3. Kubelet reserved resources — verify systemReserved/kubeReserved.
  4. Pod topology spread — for multi-replica deployments.

Priority 3: Infrastructure

  1. Jumbo frames (MTU 9000) — reduce Ceph network overhead.
  2. CPU pinning — i3 nodes have 4 cores, 2 VMs each = 100% overcommit.
  3. Utilize n5pro — 24 cores + 91GB RAM idle. Place worker VM there.
  4. HA group with n5pro — define HA group so VMs can failover to n5pro.

n5pro GPU Evaluation

Conclusion: GPU passthrough NOT possible.

  • Radeon 890M is integrated APU graphics (shares system RAM)
  • No discrete PCIe GPU, no PCIe slots for expansion
  • SR-IOV not supported on consumer integrated GPUs
  • MIG is NVIDIA-only

Alternative uses for n5pro:

  • CPU-pinned worker VM (excellent: 24C, 91GB RAM)
  • Ceph OSD node (JMB58x SATA + 5 disks → new OSDs, WAL/DB LVs ready)
  • vLLM CPU inference (possible but slow vs GPU)
  • General compute workloads (best CPU/RAM in cluster)

Failover Analysis

  • All 6 K8s VMs have HA configured
  • No HA groups defined ⚠️
  • i3 nodes: 4 cores each, 2 K8s VMs per node = 8 vCPUs on 4 physical cores
  • If proxmox5 fails (CP-01 + Worker-01), no other i3 node can absorb 8 vCPUs
  • n5pro (24C, 91GB) could host all 6 VMs simultaneously — ideal HA fallback
  • Recommendation: Define HA group [proxmox2, proxmox3, proxmox5, n5pro] with n5pro as last-resort fallback

Data Collection Methodology

Reliable VM Inventory (avoids SSH quoting issues)

# From PVE coordinator (10.0.20.10):
pvesh get /cluster/resources --type vm --output-format json
# Returns JSON array with vmid, name, node, status, maxmem, maxcpu, maxdisk
# Parse with python3 -c "import sys,json; ..."

SSH Chain Quoting Pitfall

Nested SSH commands with grep and awk break due to quote escaping through multiple SSH layers. Solutions:

  1. Use pvesh get /cluster/resources --type vm --output-format json (API)
  2. Use heredoc: ssh root@host 'bash -s' << 'SCRIPT' ... SCRIPT
  3. Use pvesh API calls instead of parsing CLI output

Direct SSH to n5pro

n5pro is NOT reachable by hostname from proxmox1 — use IP:

ssh -o StrictHostKeyChecking=no 10.0.20.91 'commands'