# Alertmanager → Hermes Webhook Integration Session-verified setup guide for routing Prometheus/Alertmanager alerts through the Hermes webhook platform for automatic RCA delivery to Telegram. ## Architecture ``` Prometheus (rules) → Alertmanager → Hermes Webhook (8644) → Agent RCA → Telegram ``` The old Telegram-bridge-bot is replaced. Every alert triggers an agent run that performs root-cause analysis and posts findings to the user's chat. ## Components - **Prometheus**: scrapes exporters, evaluates alerting rules, fires alerts - **Alertmanager**: deduplicates, groups, routes alerts to receivers - **Hermes Webhook**: receives Alertmanager POST, triggers agent run - **Agent**: analyzes alert payload, investigates Prometheus/APIs, delivers RCA ## Step-by-Step Setup ### 1. Hermes Webhook Subscription ```bash hermes webhook subscribe alertmanager-rca \ --prompt 'Alertmanager Event - Status: {status} Severity: {commonLabels.severity} Alert: {commonLabels.alertname} Instance: {commonLabels.instance} Summary: {commonAnnotations.summary} Description: {commonAnnotations.description} Perform a root cause analysis. Investigate the affected system using available tools (Prometheus queries, SSH to hosts, API calls). Report findings concisely.' \ --description "Alertmanager alerts → automatic RCA" \ --deliver telegram \ --deliver-chat-id "" ``` **Alertmanager payload structure** (top-level fields available in templates): - `status` — "firing" or "resolved" - `commonLabels` — merged labels across all alerts in the group - `commonAnnotations` — merged annotations - `alerts` — array of individual alert objects (NOT templatable with dot notation) - `groupKey`, `receiver`, `externalURL` ### 2. Alertmanager Configuration ```yaml route: receiver: hermes-rca group_by: ['alertname', 'instance'] group_wait: 30s group_interval: 5m repeat_interval: 4h receivers: - name: hermes-rca webhook_configs: - url: http://:8644/webhooks/alertmanager-rca send_resolved: true max_alerts: 0 ``` Push to the monitoring container and reload: ```bash # Config is typically mounted from host filesystem # Adjust path to your setup CONFIG_B64=$(base64 -w0 alertmanager.yml) ssh root@ "echo '$CONFIG_B64' | base64 -d | \ pct exec -- tee /opt/monitoring/alertmanager/alertmanager.yml >/dev/null && \ pct exec -- bash -lc 'docker restart alertmanager'" ``` ### 3. Prometheus Alerting Rules Copy `templates/alerting_rules.yml` to the Prometheus rules directory and reload: ```bash RULES_B64=$(base64 -w0 alerting_rules.yml) ssh root@ "echo '$RULES_B64' | base64 -d | \ pct exec -- tee /opt/monitoring/prometheus/rules/alerting_rules.yml >/dev/null && \ pct exec -- bash -lc 'docker exec prometheus wget -qO- --post-data= http://localhost:9090/-/reload'" ``` ### 4. Verification ```bash # Check rules loaded curl -s http://localhost:9090/api/v1/rules | python3 -m json.tool | head -40 # Check active alerts curl -s http://localhost:9090/api/v1/alerts | python3 -m json.tool # Check alertmanager sees the same alerts curl -s http://localhost:9093/api/v2/alerts | python3 -m json.tool # Test webhook end-to-end hermes webhook test alertmanager-rca --payload '{"status":"firing","commonLabels":{"alertname":"TestAlert","severity":"warning"},"commonAnnotations":{"summary":"Test","description":"Testing webhook"}}' ``` ## Critical Pitfalls ### 0. Webhook Agent Has No Investigation Tools (MOST COMMON FAILURE) **Symptom**: Alert arrives via webhook, agent responds but says "I don't have shell/curl/terminal tools available" and cannot perform RCA. **Root Cause**: The default `hermes-webhook` toolset (`_HERMES_WEBHOOK_SAFE_TOOLS` in `toolsets.py`) only includes `web_search`, `web_extract`, `vision_analyze`, and `clarify` — no terminal, no file, no SSH. This is intentional for security (webhook payloads are untrusted) but makes local investigation impossible. **Fix**: Add `webhook` to `platform_toolsets` in `~/.hermes/config.yaml`: ```yaml platform_toolsets: webhook: - browser - code_execution - delegation - file - memory - session_search - skills - terminal - todo - vision - web ``` Then restart the gateway (from a SEPARATE shell — cannot restart from inside the gateway process). Use a one-shot cronjob if no external shell is available: ``` cronjob(action='create', schedule='1m', prompt='systemctl --user restart hermes-gateway') ``` ### 1. PVE Alert Storm (624 false alerts) **Symptom**: Hundreds of `PVEHAServiceDegraded` or `PVENodeDown` alerts fire immediately after enabling rules. **Root Cause**: The `pve_up` and `pve_ha_state` metrics include ALL guests (including intentionally stopped VMs/CTs) and ALL HA states (15 states per guest). Without filtering: - `pve_up == 0` matches every stopped guest (dozens) - `pve_ha_state != 1` matches every non-started state × every guest (600+) **Fix**: Filter rules strictly: ```yaml # Only alert on actual PVE NODES, not guests expr: pve_up{id=~"node/.*"} == 0 # Only alert on real error states expr: pve_ha_state{state=~"error|fence|freeze|gone"} == 1 ``` ### 2. Stale Alerts in Alertmanager After Rule Changes **Symptom**: Old alerts remain in Alertmanager's API even after Prometheus stops firing them. **Root Cause**: Alertmanager caches alerts and waits for Prometheus to send "resolved" notifications. If the rule that generated them no longer exists, Prometheus never sends the resolution. **Fix**: Restart Alertmanager to flush its cache: ```bash docker restart alertmanager ``` ### 3. Blackbox ICMP False Positives **Symptom**: `HostUnreachableICMP` fires for a host that is actually reachable via TCP. **Root Cause**: Many cloud providers and firewalls block ICMP echo requests while allowing TCP traffic. PBS Remote (noris nbg1) is a typical example. **Fix**: Either whitelist ICMP on the target firewall, or switch the Blackbox probe from ICMP to TCP port check: ```yaml # In blackbox.yml, use tcp module instead of icmp modules: tcp_connect: prober: tcp timeout: 5s tcp: tls: false ``` ### 4. File Transfer to LXC Containers `scp` and `pct push` often fail with permission errors. The reliable method is base64 encoding through SSH pipes: ```bash # Encode locally, decode remotely B64=$(base64 -w0 /local/file.py) ssh root@ "echo '$B64' | base64 -d | \ pct exec -- tee /remote/path/file.py >/dev/null" ``` ### 5. Docker Container Config Reload When config files are bind-mounted from the host, edit the HOST file (not inside the container). Then reload: ```bash # Prometheus (hot reload via HTTP) docker exec prometheus wget -qO- --post-data= http://localhost:9090/-/reload # Alertmanager (hot reload via HTTP) docker exec alertmanager wget -qO- --post-data= http://localhost:9093/-/reload # If hot reload fails, restart the container docker restart prometheus docker restart alertmanager ``` ## RCA Workflow When an alert arrives via webhook, the agent should: 1. **Parse the alert** — extract alertname, severity, instance, summary 2. **Query Prometheus** — check related metrics, history, correlated alerts 3. **Investigate the root cause** — SSH to affected hosts, check logs, APIs 4. **Report findings** — concise summary with root cause and recommended fix Example RCA script pattern (runs inside the monitoring container): ```python import json, urllib.request, urllib.parse, socket PROM = "http://127.0.0.1:9090" def q(expr): url = f"{PROM}/api/v1/query?query={urllib.parse.quote(expr)}" with urllib.request.urlopen(url, timeout=5) as r: d = json.loads(r.read()) return d.get("data", {}).get("result", []) # Check probe success, HTTP status, TCP reachability # Cross-reference with direct curl/socket connections ``` ## Key Metrics Reference | Metric | What it measures | Alert threshold | |--------|-----------------|----------------| | `up{job="node_exporter"}` | Node reachability | `== 0` for 2m | | `pve_up{id=~"node/.*"}` | PVE node status (FILTER!) | `== 0` for 2m | | `pve_ha_state{state=~"error\|fence\|freeze\|gone"}` | HA error states only | `== 1` for 1m | | `ceph_health_status` | Cluster health (0=OK,1=WARN,2=ERR) | `>= 1` | | `mysql_up` | Galera node alive | `== 0` for 1m | | `mysql_global_status_wsrep_cluster_size` | Galera nodes connected | `< 3` for 2m | | `probe_success{job="blackbox_icmp"}` | ICMP reachability | `== 0` for 2m | | `probe_success{job="blackbox_http"}` | HTTP endpoint | `== 0` for 2m | ## Infrastructure Layout (Reference) - Proxmox host: `10.0.20.70`, SSH key: `~/.ssh/id_ed25519_proxmox` - Monitoring container: LXC CT 141 (IP `10.0.30.141`) - Prometheus: `localhost:9090` (in CT141 Docker) - Grafana: `localhost:3000` (in CT141 Docker, admin/Grafana2026!) - Alertmanager: `localhost:9093` (in CT141 Docker) - Hermes Gateway: `10.0.30.230:8644` - Telegram chat ID: `223926918` - PVE exporter: `pve-exporter:9221` (job `pve_api`) - Ceph exporter: `10.0.20.70:9283` (job `ceph`) - Galera nodes: `10.0.30.71/72/73` (job `mysqld_exporter`) - PBS Remote: `213.95.54.60:8007` (job `pbs-remote`, ICMP blocked)