- New: smart-home/home-assistant-dashboard-conventions (Mushroom cards, view tabs, no Bubble Cards) - Updated: rke2, ceph, galera, proxmox, brainstorming, compound-learning, 1password-cli, smart-home-automation skills - New references: ceph-cluster-administration, docker-volume-forensics, ceph-crush-weight, ceph-ec-mixed-size
321 lines
14 KiB
Markdown
321 lines
14 KiB
Markdown
# HA Manager & Custom Grafana Dashboards
|
||
|
||
## HA Manager (PVE ha-manager)
|
||
|
||
### Adding a CT/VM to HA Manager
|
||
```bash
|
||
ha-manager add ct:127 --max_restart 1 --max_relocate 1
|
||
```
|
||
No `--group` parameter needed (groups migrated to rules in newer PVE).
|
||
Specifying `--group 1` fails with: `invalid configuration ID '1'`.
|
||
Verify: `ha-manager status | grep 127` → `service ct:127 (proxmox6, started)`.
|
||
|
||
### Pitfall: ha-manager groups → rules migration (PVE 9.x)
|
||
`ha-manager groupadd` and `ha-manager groupconfig` fail with "ha groups have been migrated to rules".
|
||
Use `ha-manager rules list` / `ha-manager rules config` instead.
|
||
|
||
**PVE 9.2.3 Rules Syntax** (discovered 2026-07-12, took multiple failed attempts):
|
||
|
||
```bash
|
||
# Create node-affinity rule (replaces HA groups)
|
||
ha-manager rules add node-affinity <rule-name> \
|
||
--nodes "n5pro,proxmox5,proxmox3,proxmox2" \
|
||
--resources "vm:118,vm:128,vm:129,vm:130,vm:131,vm:132"
|
||
|
||
# List all rules (table format)
|
||
ha-manager rules config
|
||
|
||
# Raw config file
|
||
cat /etc/pve/ha/rules.cfg
|
||
```
|
||
|
||
Rule types: `node-affinity` (prefer certain nodes) and `resource-affinity` (keep resources together/apart).
|
||
|
||
**Failed approaches** (do NOT use):
|
||
- `ha-manager groupadd k8s-nodes --nodes "..."` → "cannot create group: ha groups have been migrated to rules"
|
||
- `ha-manager add vm:118 --group k8s-nodes` → "invalid parameter 'group'"
|
||
- `ha-manager rules add location "node==n5pro"` → type must be `node-affinity` or `resource-affinity`
|
||
- `ha-manager rules add node-affinity "n5pro"` without `--nodes` → "missing value for required option 'nodes'"
|
||
|
||
**Removing + re-adding HA services** (needed when migrating from old group config):
|
||
```bash
|
||
ha-manager remove vm:118 # VM keeps running, loses HA protection, shows "deleting"
|
||
ha-manager add vm:118 --state started # Re-add immediately
|
||
ha-manager status | grep "service vm:118" # Verify: "started"
|
||
```
|
||
|
||
⚠️ When `ha-manager remove` runs, the service transitions through "deleting" state. The VM continues running. Re-add immediately to restore HA protection.
|
||
|
||
### CT 127 Added to HA (2026-07-05)
|
||
CT 127 (avahi-reflector, proxmox6) added to ha-manager with max_restart=1,
|
||
max_relocate=1. Was previously NOT HA-managed. Verified: `service ct:127
|
||
(proxmox6, queued)` → `started`.
|
||
|
||
### K8s VMs Node-Affinity Rule (2026-07-12)
|
||
Created rule `k8s-prefer-n5pro` (type: node-affinity) for VMs 118,128-132.
|
||
Initial nodes: n5pro,proxmox5,proxmox3,proxmox2 (n5pro = preferred, others = failover).
|
||
All 6 K8s VMs re-added to HA after brief "deleting" transition.
|
||
Existing rules in cluster: `ha-rule-9a834247-53bf` (ct:104,ct:99999 → n5pro),
|
||
`ha-rule-ea9e852a-8d82` (vm:310,vm:311 negative affinity), `ha-rule-d61699eb-2baf` (vm:301,vm:302 negative affinity).
|
||
|
||
### ⚠️ PITFALL: HA Node-Affinity Blocks Live Migration (2026-07-12)
|
||
When K8s VMs are HA-managed with a `node-affinity` rule restricting them to
|
||
certain nodes, `qm migrate <vmid> <target> --online` FAILS if the target node
|
||
is not in the rule's `--nodes` list:
|
||
|
||
```
|
||
cannot migrate resource 'vm:132' to node 'proxmox1':
|
||
- resource 'vm:132' not allowed on target node 'proxmox1'
|
||
```
|
||
|
||
**Fix:** Expand the rule's node list BEFORE attempting migration. Edit
|
||
`/etc/pve/ha/rules.cfg` directly (it's a shared FUSE filesystem, edits
|
||
propagate instantly):
|
||
|
||
```bash
|
||
# Add all cluster nodes to the rule
|
||
sed -i '/^node-affinity: k8s-prefer-n5pro$/,/^$/{
|
||
s/nodes n5pro,proxmox2,proxmox3,proxmox5/nodes n5pro,proxmox1,proxmox2,proxmox3,proxmox4,proxmox5,proxmox6,proxmox7/
|
||
}' /etc/pve/ha/rules.cfg
|
||
```
|
||
|
||
Then retry `qm migrate` from the SOURCE node (not the coordinator). The
|
||
migration command must run on the node where the VM currently lives:
|
||
|
||
```bash
|
||
# CORRECT: SSH to source node, then migrate
|
||
ssh root@10.0.20.20 'qm migrate 132 proxmox1 --online' # proxmox2 → proxmox1
|
||
ssh root@10.0.20.30 'qm migrate 131 proxmox4 --online' # proxmox3 → proxmox4
|
||
ssh root@10.0.20.50 'qm migrate 128 n5pro --online' # proxmox5 → n5pro
|
||
|
||
# WRONG: Running qm migrate from the coordinator (proxmox1) for a VM on proxmox2
|
||
# → "400 Parameter verification failed. target: target is local node."
|
||
```
|
||
|
||
HA migrations are asynchronous — `qm migrate` returns immediately with
|
||
"Requesting HA migration for VM NNN to node XXX". The HA manager processes
|
||
the migration in the background. Check progress with:
|
||
|
||
```bash
|
||
ha-manager status | grep "vm:NNN"
|
||
# States: migrate → starting → started
|
||
```
|
||
|
||
Poll every 15-30s until all show `started`. Typical duration: 30-90s per VM
|
||
(longer if Ceph is under backfill load).
|
||
|
||
### K8s VM Anti-Colocation Strategy (2026-07-12)
|
||
|
||
**Problem:** Initially each K8s node hosted BOTH a CP AND a Worker:
|
||
|
||
| Node | CP | Worker | Risk |
|
||
|------|-----|--------|------|
|
||
| proxmox2 | cp-03 (129) | worker-02 (132) | ⚠️ Node loss = 2 VMs gone |
|
||
| proxmox3 | cp-02 (130) | worker-03 (131) | ⚠️ Node loss = 2 VMs gone |
|
||
| proxmox5 | cp-01 (118) | worker-01 (128) | ⚠️ Node loss = 2 VMs gone |
|
||
|
||
**Solution:** Spread 6 VMs across 6 different nodes (1 per node):
|
||
|
||
| Node | CP | Worker |
|
||
|------|-----|--------|
|
||
| proxmox1 | — | worker-02 (132) |
|
||
| proxmox2 | cp-03 (129) | — |
|
||
| proxmox3 | cp-02 (130) | — |
|
||
| proxmox4 | — | worker-03 (131) |
|
||
| proxmox5 | cp-01 (118) | — |
|
||
| n5pro | — | worker-01 (128) |
|
||
|
||
Since all disks are on shared Ceph storage, live migration is pure RAM
|
||
transfer — no storage move needed. Quorum survives any single node failure
|
||
(2/3 CPs remain), and only 1 worker is lost per node failure (33% capacity
|
||
reduction, not 66%).
|
||
|
||
**Corosync node IP reference** (from `/etc/pve/corosync.conf`):
|
||
|
||
| Node | Ring0 IP |
|
||
|------|----------|
|
||
| proxmox1 | 10.0.20.10 |
|
||
| proxmox2 | 10.0.20.20 |
|
||
| proxmox3 | 10.0.20.30 |
|
||
| proxmox4 | 10.0.20.40 |
|
||
| proxmox5 | 10.0.20.50 |
|
||
| proxmox6 | 10.0.20.60 |
|
||
| proxmox7 | 10.0.20.70 |
|
||
| n5pro | 10.0.20.91 |
|
||
|
||
### DB Nodes Strict Affinity Rule (2026-07-13, updated for quorum safety)
|
||
|
||
**Problem:** Galera + MaxScale VMs were co-located with K8s CP/worker VMs on 15 GB
|
||
PVE nodes, causing RAM overcommit >147%. An OOM-killer event on proxmox2 killed
|
||
the KVM process of VM300 (mariadb-01), dropping the Galera cluster from 3→2 nodes.
|
||
|
||
**Initial fix (2 hosts — had quorum risk):** Restricted DB VMs to n5pro + proxmox3.
|
||
But this placed 2 of 3 Galera nodes on proxmox3 — if proxmox3 fails, only 1 Galera
|
||
node survives → **no quorum, entire cluster down.**
|
||
|
||
**⚠️ CRITICAL Quorum Rule for Galera Placement:**
|
||
Each Galera node MUST be on a different physical PVE host. With 3 nodes, losing
|
||
any single host must leave ≥2 alive (= quorum). 2 nodes on the same host =
|
||
single point of failure for the entire database cluster.
|
||
|
||
**Final fix (3 hosts, quorum-safe):** Added proxmox6 as third allowed host:
|
||
|
||
```bash
|
||
# /etc/pve/ha/rules.cfg (shared FUSE FS, edits propagate instantly)
|
||
node-affinity: db-nodes
|
||
nodes n5pro,proxmox3,proxmox6
|
||
resources vm:300,vm:301,vm:302,vm:310,vm:311
|
||
strict 1
|
||
```
|
||
|
||
**Resulting placement (3 different hardware hosts):**
|
||
| VM | Node | Quorum impact if host fails |
|
||
|----|------|---------------------------|
|
||
| 300 (mariadb-01) | n5pro | 2/3 survive on proxmox3+proxmox6 ✅ |
|
||
| 301 (mariadb-02) | proxmox3 | 2/3 survive on n5pro+proxmox6 ✅ |
|
||
| 302 (mariadb-03) | proxmox6 | 2/3 survive on n5pro+proxmox3 ✅ |
|
||
|
||
With `strict 1`, the HA manager will NEVER place these VMs on other nodes,
|
||
even during failover. If all 3 allowed nodes are down, the VMs stay down
|
||
rather than starting on an overloaded 15 GB node and risking another OOM.
|
||
|
||
**⚠️ PITFALL: Live Migration Can Kill Recovering Galera Nodes**
|
||
Never live-migrate a Galera VM that is still in `activating` state (SST in
|
||
progress) or was recently OOM-killed. The VM can freeze during migration,
|
||
drop from the cluster, and become unreachable (guest agent down, SSH refused).
|
||
Wait until `wsrep_local_state_comment = Synced` before migrating.
|
||
|
||
**Existing anti-affinity rules** (keep Galera/MaxScale nodes apart):
|
||
- `vm:310,vm:311` negative resource-affinity (MaxScale auseinander)
|
||
- `vm:301,vm:302` negative resource-affinity (Galera auseinander)
|
||
|
||
⚠️ Anti-affinity can block migration: VM 311 couldn't migrate to n5pro because
|
||
VM 310 was already there and they have negative affinity. Had to use proxmox3
|
||
instead. Always check `rules.cfg` before choosing a migration target.
|
||
|
||
### OOM-Killer on PVE: Diagnosis Pattern (2026-07-13)
|
||
|
||
When a VM is unreachable (SSH closed, Guest Agent down) but QEMU shows `running`,
|
||
check the PVE host's kernel log for OOM-killer events:
|
||
|
||
```bash
|
||
# On the PVE host where the VM runs:
|
||
dmesg -T | grep -iE "oom|killed|out of memory" | tail -20
|
||
# Look for: "Out of memory: Killed process <PID> (kvm)" with task_memcg=/qemu.slice/<VMID>.scope
|
||
```
|
||
|
||
**Root cause pattern:** PVE RAM overcommit = sum(VM maxmem) / node physical RAM.
|
||
When this exceeds 100% and a VM demands more memory, the host kernel OOM-killer
|
||
selects the largest memory consumer (usually a KVM process) and kills it.
|
||
|
||
**Why `balloon: 0` on DB VMs is correct:** MySQL/Galera uses InnoDB buffer pool
|
||
(hot pages constantly accessed). If the balloon shrinks, the guest kernel swaps
|
||
hot buffer pool pages → massive latency → Galera flow control throttles the ENTIRE
|
||
cluster. Ballooning is appropriate for idle/web VMs, NOT for database workloads.
|
||
|
||
**Mitigation hierarchy:**
|
||
1. Migrate DB VMs to nodes with RAM headroom (n5pro: 91 GB, proxmox3: 31 GB)
|
||
2. Set strict HA node-affinity rules to prevent failover to small nodes
|
||
3. Deploy RAM-based rebalancer cron job (see below)
|
||
4. Long-term: upgrade RAM on 15 GB nodes or redistribute K8s VMs
|
||
|
||
### RAM-Based Rebalancer (2026-07-13)
|
||
|
||
PVE HA manager does NOT do RAM-based rebalancing — `ha-manager rebalance` counts
|
||
VMs per node, not their resource consumption. A custom script + cron job fills
|
||
this gap:
|
||
|
||
**Script:** `~/.hermes/scripts/pve-ram-rebalancer.sh`
|
||
- Queries `pvesh get /cluster/resources` for VM allocations and node RAM
|
||
- Calculates overcommit = sum(running VM maxmem on node) / node physical RAM
|
||
- If any node > 100%, finds the largest HA-managed VM and migrates it to the
|
||
node with the most free RAM (that won't exceed 100% after adding the VM)
|
||
- Respects HA node-affinity rules (strict rules constrain target selection)
|
||
- Only migrates HA-managed VMs (non-HA VMs can't be safely migrated via ha-manager)
|
||
- **Silent when balanced** (empty stdout = nothing to do)
|
||
- Outputs `MIGRATE <vmid> <from> <to>` lines when action needed
|
||
|
||
**Cron job:** Job ID `a41717349482`, every 10 minutes, delivers to SRE agent.
|
||
The agent reviews proposed migrations, executes them via `qm migrate --online`,
|
||
and reports a summary. Only triggers when the script produces output (changes).
|
||
|
||
**Known limitation:** If ALL nodes with free RAM are constrained by HA rules
|
||
(e.g., db-nodes strict rule limits DB VMs to n5pro+proxmox3, and both are full),
|
||
the script correctly reports "balanced" — it can't violate HA rules. The solution
|
||
is to relax the rule or add more RAM-capable nodes to the rule.
|
||
|
||
### Cluster RAM Inventory (2026-07-13)
|
||
|
||
| Node | Physical RAM | Cores | Role |
|
||
|------|-------------|-------|------|
|
||
| n5pro | 91 GB | 24 | Big node — DB VMs, K8s worker |
|
||
| proxmox3 | 31 GB | 4 | Medium — DB failover, K8s CP |
|
||
| proxmox1 | 15 GB | 8 | Small — K8s worker, HA, embedding |
|
||
| proxmox2 | 15 GB | 4 | Small — K8s CP (was OOM source) |
|
||
| proxmox4 | 15 GB | 4 | Small — K8s worker, openwebui |
|
||
| proxmox5 | 15 GB | 4 | Small — K8s CP |
|
||
| proxmox6 | 15 GB | 4 | Small — various CTs |
|
||
| proxmox7 | 15 GB | 4 | Small — Hermes, monitoring, SMB |
|
||
|
||
⚠️ 6 of 8 nodes have only 15 GB RAM. K8s VMs (12 GB each) + any other VM
|
||
on the same node = overcommit. Distribute carefully.
|
||
|
||
### RKE2 CP Nodes Are Untainted (2026-07-13)
|
||
|
||
RKE2 does NOT set `NoSchedule` taints on control-plane nodes by default.
|
||
All 6 K8s nodes (3 CP + 3 Worker) schedule pods freely:
|
||
|
||
| Node | Role | Pods |
|
||
|------|------|------|
|
||
| rke2-cp-01 | CP | 21 (most!) |
|
||
| rke2-cp-03 | CP | 13 |
|
||
| rke2-cp-02 | CP | 11 |
|
||
| rke2-worker-03 | Worker | 14 |
|
||
| rke2-worker-02 | Worker | 6 |
|
||
| rke2-worker-01 | Worker | 5 |
|
||
|
||
This means CP nodes carry real workload — don't assume they're "just etcd+API".
|
||
When planning RAM allocation, count CP nodes as full workload nodes.
|
||
|
||
## Custom Grafana Dashboard Creation via REST API
|
||
|
||
Custom dashboards can be created directly via the Grafana REST API without
|
||
file provisioning. Faster for one-off dashboards than download-patch-push.
|
||
|
||
```python
|
||
import json, urllib.request, base64
|
||
GRAFANA_URL = "http://10.0.30.141:3000"
|
||
GRAFANA_AUTH = "admin:Grafana2026!"
|
||
DS_UID = "PBFA97CFB590B2093"
|
||
dashboard = {"uid":"my-custom","title":"...","panels":[...]}
|
||
payload = json.dumps({"dashboard":dashboard,"overwrite":True,"folderUid":""}).encode()
|
||
req = urllib.request.Request(f"{GRAFANA_URL}/api/dashboards/db", data=payload, method="POST",
|
||
headers={"Content-Type":"application/json",
|
||
"Authorization":"Basic "+base64.b64encode(GRAFANA_AUTH.encode()).decode()})
|
||
resp = urllib.request.urlopen(req)
|
||
```
|
||
|
||
Panel grid: 24 cols, rows in multiples of 4. Stat h=4, timeseries/table h=8.
|
||
State-timeline for HA status over time.
|
||
|
||
### Available PVE Metrics for Custom Dashboards
|
||
| Metric | Labels | Description |
|
||
|--------|--------|-------------|
|
||
| `pve_up` | `id=node/*` | Node online (0/1) |
|
||
| `pve_guest_info` | `id,name,node,type` | Guest running (0/1) |
|
||
| `pve_cpu_usage_ratio` | `id` | CPU ratio (×100 for %) |
|
||
| `pve_memory_usage_bytes` | `id` | Mem used bytes |
|
||
| `pve_uptime_seconds` | `id` | Uptime seconds |
|
||
| `pve_ha_state` | `id,state` | HA state enum |
|
||
| `pve_network_{receive,transmit}_bytes_total` | `id` | Net counters |
|
||
| `pve_disk_{usage,size}_bytes` | `id,disk` | Disk stats |
|
||
|
||
### Network Dashboard (2026-07-05)
|
||
"Network & Infrastructure Overview" (UID: `network-infra`, 21 panels):
|
||
PVE nodes/guests status+CPU+mem+net, HA state timeline, blackbox ICMP/HTTP,
|
||
SSL cert expiry, node temps. URL: `http://10.0.30.141:3000/d/network-infra`
|
||
|
||
### Pitfall: execute_code blocked in cron-safe mode
|
||
`execute_code` may be blocked if cron approval mode is restrictive.
|
||
Workaround: `write_file` Python script to `/tmp/`, then `python3 /tmp/script.py`.
|