Initial commit: Hermes Agent Skills collection

This commit is contained in:
Debian
2026-07-12 19:02:59 +00:00
commit e9cc106625
789 changed files with 233126 additions and 0 deletions
@@ -0,0 +1,207 @@
# Galera Cluster Recovery for Home Assistant Database
## Architecture
```
HA (VM 106, 10.0.30.10)
└─ recorder: mysql://ha_recorder@10.0.30.70:3306/homeassistant
└─ MaxScale VIP (10.0.30.70, keepalived)
├─ maxscale-01 (10.0.30.81, VM 310, n5pro)
└─ maxscale-02 (10.0.30.82, VM 311, proxmox2)
└─ readwritesplit → Galera cluster
├─ db1 (10.0.30.71, VM 300, proxmox2)
├─ db2 (10.0.30.72, VM 301, proxmox4)
└─ db3 (10.0.30.73, VM 302, proxmox5)
```
- **MariaDB**: 11.4.10-MariaDB-deb12
- **MaxScale**: 24.02.9
- **Galera cluster name**: `mariadb-galera`
- **Cluster address**: `gcomm://10.0.30.71,10.0.30.72,10.0.30.73`
- **DB size**: ~18GB (homeassistant schema)
- **SSH to DB nodes**: `ssh -i ~/.ssh/id_ed25519_proxmox debian@10.0.30.{71,72,73}`
- **SSH to MaxScale**: `ssh -i ~/.ssh/id_ed25519_proxmox debian@10.0.30.81`
## Symptoms of Cluster Failure
- HA logs: `MySQLdb.OperationalError: (1045, "Access denied for user 'ha_recorder'@'10.0.30.10'")` or `(2013, "Lost connection to server during query")`
- HA: Recorder service loaded but History/Logbook domains missing
- All 3 MariaDB nodes: `systemctl status mariadb` shows `Active: failed`
- Port 3306 closed on all DB nodes
- MaxScale: all servers show "Down"
## Recovery Procedure
### Step 1: Verify All Nodes Are Down
```bash
for ip in 10.0.30.{71,72,73}; do
echo "=== $ip ==="
ssh -i ~/.ssh/id_ed25519_proxmox debian@$ip "
sudo systemctl is-active mariadb
sudo journalctl -u mariadb --no-pager -n 5
"
done
```
Look for: `Failed to reach primary view`, `gcomm://...: -110 (Connection timed out)`, `Aborting`
### Step 2: Bootstrap From Node 1
```bash
ssh -i ~/.ssh/id_ed25519_proxmox debian@10.0.30.71 "
sudo systemctl stop mariadb 2>/dev/null
sudo galera_new_cluster
"
```
Wait ~30s, then verify:
```bash
ssh -i ~/.ssh/id_ed25519_proxmox debian@10.0.30.71 "
sudo mariadb -e \"
SHOW GLOBAL STATUS LIKE 'wsrep_ready';
SHOW GLOBAL STATUS LIKE 'wsrep_cluster_size';
SHOW GLOBAL STATUS LIKE 'wsrep_cluster_status';
SHOW GLOBAL STATUS LIKE 'wsrep_local_state_comment';
\"
"
```
Expected: `wsrep_ready=ON`, `wsrep_cluster_size=1`, `wsrep_cluster_status=Primary`, `wsrep_local_state_comment=Synced`
### Step 3: Enable MaxScale Routing to Donor
During SST, the bootstrapped node becomes Donor/Desynced. MaxScale excludes Donor nodes by default.
```bash
ssh -i ~/.ssh/id_ed25519_proxmox debian@10.0.30.81 "
maxctrl alter monitor galera-monitor available_when_donor=true
maxctrl list servers
"
```
Verify db1 shows "Master, Synced, Running, Donor/Desynced" and connections work:
```bash
ssh -i ~/.ssh/id_ed25519_proxmox debian@10.0.30.81 "
mariadb -h 127.0.0.1 -P 3306 -u ha_recorder -p'<password>' --skip-ssl -e \"SELECT 'OK' AS r, @@hostname;\"
"
```
### Step 4: Restart HA (NOW the DB is reachable)
```bash
source ~/.hermes/.env
curl -s -X POST "$HA_URL/api/services/homeassistant/restart" \
-H "Authorization: Bearer $HA_TOKEN" \
-H "Content-Type: application/json" -d '{}'
```
Poll for HA startup (takes 2-3 min with 18GB DB):
```bash
for i in $(seq 1 18); do
sleep 10
STATE=$(curl -s "$HA_URL/api/config" -H "Authorization: Bearer $HA_TOKEN" | python3 -c "import sys,json; print(json.load(sys.stdin).get('state','?'))" 2>/dev/null)
echo "Check $i: $STATE"
[ "$STATE" = "RUNNING" ] && break
done
```
Verify recorder/history/logbook:
```bash
curl -s "$HA_URL/api/services" -H "Authorization: Bearer $HA_TOKEN" | python3 -c "
import sys,json
data=json.load(sys.stdin)
services=set(s['domain'] for s in data)
for d in ['recorder','history','logbook']:
print(f'{d}: {\"✓\" if d in services else \"✗\"}')
"
```
Test history API:
```bash
curl -s -w " HTTP:%{http_code}" "$HA_URL/api/history/period?filter_entity_id=sun.sun" \
-H "Authorization: Bearer $HA_TOKEN" | head -1 | cut -c1-100
```
### Step 5: Join db2 (SEQUENTIAL — one at a time!)
```bash
ssh -i ~/.ssh/id_ed25519_proxmox debian@10.0.30.72 "sudo systemctl start --no-block mariadb"
```
Monitor SST progress (18GB DB takes ~30 min):
```bash
ssh -i ~/.ssh/id_ed25519_proxmox debian@10.0.30.72 "
for i in \$(seq 1 120); do
STATUS=\$(sudo systemctl is-active mariadb 2>/dev/null)
if [ \"\$STATUS\" = 'active' ]; then
WSREP=\$(sudo mariadb -e \"SHOW GLOBAL STATUS LIKE 'wsrep_local_state_comment'\" 2>/dev/null | awk '/wsrep_local_state_comment/{print \$2}')
echo \"Poll \$i: ACTIVE, state=\$WSREP\"
[ \"\$WSREP\" = 'Synced' ] && echo 'DB2 SYNCED!' && break
else
SIZE=\$(sudo du -sh /var/lib/mysql/ 2>/dev/null | awk '{print \$1}')
echo \"Poll \$i: \$STATUS, datadir=\$SIZE\"
fi
sleep 10
done
"
```
Also check from donor side:
```bash
ssh -i ~/.ssh/id_ed25519_proxmox debian@10.0.30.71 "
sudo mariadb -e \"SHOW GLOBAL STATUS LIKE 'wsrep_local_state_comment';\"
ps aux | grep mariabackup | grep -v grep
"
```
### Step 6: Join db3 (ONLY after db2 is Synced)
```bash
ssh -i ~/.ssh/id_ed25519_proxmox debian@10.0.30.73 "sudo systemctl start --no-block mariadb"
```
Same monitoring loop as Step 5.
### Step 7: Restore MaxScale Settings
Once all 3 nodes are Synced:
```bash
ssh -i ~/.ssh/id_ed25519_proxmox debian@10.0.30.81 "
maxctrl alter monitor galera-monitor available_when_donor=false
maxctrl clear server db2 maintenance
maxctrl clear server db3 maintenance
maxctrl list servers
"
```
All 3 should show "Synced, Running".
## Key Pitfalls
1. **Parallel SST deadlock**: Starting 2+ joiners simultaneously → "No donor candidates temporarily available in suitable state". Both stall forever. FIX: start one, wait for Synced, then start next.
2. **MaxScale `available_when_donor`**: Default `false` means MaxScale won't route ANY traffic while the only available node is a donor. HA Recorder gets "Lost connection" / "Access denied". FIX: `maxctrl alter monitor galera-monitor available_when_donor=true` before joining additional nodes.
3. **HA Recorder retry exhaustion**: `db_max_retries: 20` in configuration.yaml. If DB unreachable at boot, Recorder gives up after ~20 retries × 5s = 100s. History/Logbook never load. FIX: ensure DB is reachable via MaxScale BEFORE restarting HA.
4. **`maxctrl set server db1 master` fails on monitored servers**: "monitored server, only maintenance/drain can be altered". Don't try to manually set master role — the galera-monitor does this automatically based on wsrep state.
5. **HA startup time with large DB**: With 18GB DB, HA takes 2-3 minutes to start (schema migration/validation). Don't assume failure if `api/config` returns `NOT_RUNNING` for the first few minutes.
6. **grastate.dat may not exist**: On clean shutdown or fresh install, `grastate.ini`/`grastate.dat` may not be present. This doesn't prevent bootstrap — just run `galera_new_cluster` directly.
## Credentials
All DB credentials should be stored in 1Password (Vault "Hermes"):
- `ha_recorder` MySQL password (used in HA configuration.yaml db_url)
- MaxScale admin user/password
- Mariabackup SST auth credentials
Currently in HA configuration.yaml:
```yaml
recorder:
db_url: mysql://ha_recorder:***@10.0.30.70:3306/homeassistant?charset=utf8mb4
purge_keep_days: 30
db_max_retries: 20
```