Initial commit: Hermes Agent Skills collection
This commit is contained in:
@@ -0,0 +1,66 @@
|
||||
# Galera Cluster Recovery Runbook
|
||||
|
||||
Based on June 27, 2026 recovery of total Galera cluster failure (down since June 24 ~14:46 CEST).
|
||||
|
||||
## Timeline
|
||||
|
||||
1. **June 24 ~14:46** — All 3 Galera nodes crashed simultaneously. Root cause: likely simultaneous network issue preventing inter-node communication → "Failed to reach primary view" → all nodes aborted.
|
||||
2. **June 27 13:00** — Recovery began. db1 bootstrapped, db2/db3 stopped to resolve SST deadlock.
|
||||
3. **June 27 13:07** — db2 SST started (mariabackup, ~18GB)
|
||||
4. **June 27 13:17** — db2 SST completed (~40 min), state=Synced, cluster_size=2
|
||||
5. **June 27 13:17** — db3 SST started
|
||||
6. **June 27 ~13:50** — db3 SST completed, cluster_size=3
|
||||
|
||||
## Key Decisions Made
|
||||
|
||||
### 1. Bootstrap from Node 1 (no grastate.ini)
|
||||
No `grastate.ini` existed on any node, so no `safe_to_bootstrap` flag. Any node can be chosen. Picked db1 (10.0.30.71).
|
||||
|
||||
### 2. Stop db2/db3 to Resolve SST Deadlock
|
||||
Initially tried starting both db2 and db3 after db1 bootstrap. Both entered SST but deadlocked: `"No donor candidates temporarily available in suitable state"`. Solution: stop both, then start them one at a time.
|
||||
|
||||
### 3. MaxScale `available_when_donor=true`
|
||||
Without this, MaxScale refused to route to db1 while it was Donor/Desynced serving SST to db2. HA Recorder got `Lost connection to server during query` errors. Setting this flag fixed routing immediately.
|
||||
|
||||
### 4. Two HA Restarts Needed
|
||||
First restart failed because MaxScale wasn't routing yet (before `available_when_donor=true`). Second restart after the fix succeeded — Recorder connected, history/logbook loaded.
|
||||
|
||||
## Pitfalls Encountered
|
||||
|
||||
### PITFALL: `galera_new_cluster` Appears to Timeout
|
||||
`galera_new_cluster` may appear to hang/timeout in the terminal but actually launches mariadbd in the background successfully. Always verify separately with `systemctl is-active` and `mariadb -e "SHOW STATUS LIKE 'wsrep_%'"`.
|
||||
|
||||
### PITFALL: Parallel SST Causes Deadlock
|
||||
Starting 2+ nodes simultaneously when only 1 donor exists → all joiners compete for the same donor → donor can't serve multiple SST streams → deadlock → all abort.
|
||||
|
||||
### PITFALL: HA Recorder Retries Exhaust Before DB Recovers
|
||||
With `db_max_retries: 20` and 5-second intervals, Recorder gives up after ~100 seconds. If MaxScale isn't routing by then, Recorder shuts down. Requires HA restart after DB is confirmed accessible.
|
||||
|
||||
### PITFALL: MaxScale `set server db1 master` Doesn't Work
|
||||
Monitored servers can only have `maintenance`/`drain` set manually. Role assignment is automatic via the galeramon monitor. Don't waste time trying to force roles.
|
||||
|
||||
### PITFALL: HA Takes Long to Reach RUNNING State
|
||||
With ~18GB DB, HA's schema migration check (`pre_migrate_schema`) takes significant time. State stays `NOT_RUNNING` for 2-3 minutes. Don't assume failure — check docker logs for actual errors.
|
||||
|
||||
## Post-Recovery Cleanup
|
||||
|
||||
```bash
|
||||
# Revert available_when_donor
|
||||
maxctrl alter monitor galera-monitor available_when_donor=false
|
||||
|
||||
# Verify all nodes
|
||||
maxctrl list servers
|
||||
# All should show: "Synced, Running" with no Donor/Desynced
|
||||
|
||||
# Test HA
|
||||
curl -s "$HA_URL/api/services" -H "Authorization: Bearer $HA_TOKEN" | python3 -c "
|
||||
import sys,json
|
||||
data=json.load(sys.stdin)
|
||||
services=set(s['domain'] for s in data)
|
||||
for d in ['recorder','history','logbook']:
|
||||
print(f'{d}: {\"✓\" if d in services else \"✗\"}')
|
||||
"
|
||||
|
||||
# Test history data
|
||||
curl -s "$HA_URL/api/history/period?filter_entity_id=sun.sun" -H "Authorization: Bearer $HA_TOKEN"
|
||||
```
|
||||
Reference in New Issue
Block a user