Initial commit: Hermes Agent Skills collection
This commit is contained in:
@@ -0,0 +1,272 @@
|
||||
# Alertmanager → Hermes Webhook Integration
|
||||
|
||||
Session-verified setup guide for routing Prometheus/Alertmanager alerts
|
||||
through the Hermes webhook platform for automatic RCA delivery to Telegram.
|
||||
|
||||
## Architecture
|
||||
|
||||
```
|
||||
Prometheus (rules) → Alertmanager → Hermes Webhook (8644) → Agent RCA → Telegram
|
||||
```
|
||||
|
||||
The old Telegram-bridge-bot is replaced. Every alert triggers an agent run
|
||||
that performs root-cause analysis and posts findings to the user's chat.
|
||||
|
||||
## Components
|
||||
|
||||
- **Prometheus**: scrapes exporters, evaluates alerting rules, fires alerts
|
||||
- **Alertmanager**: deduplicates, groups, routes alerts to receivers
|
||||
- **Hermes Webhook**: receives Alertmanager POST, triggers agent run
|
||||
- **Agent**: analyzes alert payload, investigates Prometheus/APIs, delivers RCA
|
||||
|
||||
## Step-by-Step Setup
|
||||
|
||||
### 1. Hermes Webhook Subscription
|
||||
|
||||
```bash
|
||||
hermes webhook subscribe alertmanager-rca \
|
||||
--prompt 'Alertmanager Event - Status: {status}
|
||||
Severity: {commonLabels.severity}
|
||||
Alert: {commonLabels.alertname}
|
||||
Instance: {commonLabels.instance}
|
||||
Summary: {commonAnnotations.summary}
|
||||
Description: {commonAnnotations.description}
|
||||
|
||||
Perform a root cause analysis. Investigate the affected system using available
|
||||
tools (Prometheus queries, SSH to hosts, API calls). Report findings concisely.' \
|
||||
--description "Alertmanager alerts → automatic RCA" \
|
||||
--deliver telegram \
|
||||
--deliver-chat-id "<your-chat-id>"
|
||||
```
|
||||
|
||||
**Alertmanager payload structure** (top-level fields available in templates):
|
||||
- `status` — "firing" or "resolved"
|
||||
- `commonLabels` — merged labels across all alerts in the group
|
||||
- `commonAnnotations` — merged annotations
|
||||
- `alerts` — array of individual alert objects (NOT templatable with dot notation)
|
||||
- `groupKey`, `receiver`, `externalURL`
|
||||
|
||||
### 2. Alertmanager Configuration
|
||||
|
||||
```yaml
|
||||
route:
|
||||
receiver: hermes-rca
|
||||
group_by: ['alertname', 'instance']
|
||||
group_wait: 30s
|
||||
group_interval: 5m
|
||||
repeat_interval: 4h
|
||||
|
||||
receivers:
|
||||
- name: hermes-rca
|
||||
webhook_configs:
|
||||
- url: http://<hermes-ip>:8644/webhooks/alertmanager-rca
|
||||
send_resolved: true
|
||||
max_alerts: 0
|
||||
```
|
||||
|
||||
Push to the monitoring container and reload:
|
||||
```bash
|
||||
# Config is typically mounted from host filesystem
|
||||
# Adjust path to your setup
|
||||
CONFIG_B64=$(base64 -w0 alertmanager.yml)
|
||||
ssh root@<proxmox-host> "echo '$CONFIG_B64' | base64 -d | \
|
||||
pct exec <ct-id> -- tee /opt/monitoring/alertmanager/alertmanager.yml >/dev/null && \
|
||||
pct exec <ct-id> -- bash -lc 'docker restart alertmanager'"
|
||||
```
|
||||
|
||||
### 3. Prometheus Alerting Rules
|
||||
|
||||
Copy `templates/alerting_rules.yml` to the Prometheus rules directory and reload:
|
||||
```bash
|
||||
RULES_B64=$(base64 -w0 alerting_rules.yml)
|
||||
ssh root@<proxmox-host> "echo '$RULES_B64' | base64 -d | \
|
||||
pct exec <ct-id> -- tee /opt/monitoring/prometheus/rules/alerting_rules.yml >/dev/null && \
|
||||
pct exec <ct-id> -- bash -lc 'docker exec prometheus wget -qO- --post-data= http://localhost:9090/-/reload'"
|
||||
```
|
||||
|
||||
### 4. Verification
|
||||
|
||||
```bash
|
||||
# Check rules loaded
|
||||
curl -s http://localhost:9090/api/v1/rules | python3 -m json.tool | head -40
|
||||
|
||||
# Check active alerts
|
||||
curl -s http://localhost:9090/api/v1/alerts | python3 -m json.tool
|
||||
|
||||
# Check alertmanager sees the same alerts
|
||||
curl -s http://localhost:9093/api/v2/alerts | python3 -m json.tool
|
||||
|
||||
# Test webhook end-to-end
|
||||
hermes webhook test alertmanager-rca --payload '{"status":"firing","commonLabels":{"alertname":"TestAlert","severity":"warning"},"commonAnnotations":{"summary":"Test","description":"Testing webhook"}}'
|
||||
```
|
||||
|
||||
## Critical Pitfalls
|
||||
|
||||
### 0. Webhook Agent Has No Investigation Tools (MOST COMMON FAILURE)
|
||||
|
||||
**Symptom**: Alert arrives via webhook, agent responds but says "I don't have
|
||||
shell/curl/terminal tools available" and cannot perform RCA.
|
||||
|
||||
**Root Cause**: The default `hermes-webhook` toolset
|
||||
(`_HERMES_WEBHOOK_SAFE_TOOLS` in `toolsets.py`) only includes `web_search`,
|
||||
`web_extract`, `vision_analyze`, and `clarify` — no terminal, no file, no SSH.
|
||||
This is intentional for security (webhook payloads are untrusted) but makes
|
||||
local investigation impossible.
|
||||
|
||||
**Fix**: Add `webhook` to `platform_toolsets` in `~/.hermes/config.yaml`:
|
||||
```yaml
|
||||
platform_toolsets:
|
||||
webhook:
|
||||
- browser
|
||||
- code_execution
|
||||
- delegation
|
||||
- file
|
||||
- memory
|
||||
- session_search
|
||||
- skills
|
||||
- terminal
|
||||
- todo
|
||||
- vision
|
||||
- web
|
||||
```
|
||||
Then restart the gateway (from a SEPARATE shell — cannot restart from inside
|
||||
the gateway process). Use a one-shot cronjob if no external shell is available:
|
||||
```
|
||||
cronjob(action='create', schedule='1m', prompt='systemctl --user restart hermes-gateway')
|
||||
```
|
||||
|
||||
### 1. PVE Alert Storm (624 false alerts)
|
||||
|
||||
**Symptom**: Hundreds of `PVEHAServiceDegraded` or `PVENodeDown` alerts fire
|
||||
immediately after enabling rules.
|
||||
|
||||
**Root Cause**: The `pve_up` and `pve_ha_state` metrics include ALL guests
|
||||
(including intentionally stopped VMs/CTs) and ALL HA states (15 states per
|
||||
guest). Without filtering:
|
||||
|
||||
- `pve_up == 0` matches every stopped guest (dozens)
|
||||
- `pve_ha_state != 1` matches every non-started state × every guest (600+)
|
||||
|
||||
**Fix**: Filter rules strictly:
|
||||
```yaml
|
||||
# Only alert on actual PVE NODES, not guests
|
||||
expr: pve_up{id=~"node/.*"} == 0
|
||||
|
||||
# Only alert on real error states
|
||||
expr: pve_ha_state{state=~"error|fence|freeze|gone"} == 1
|
||||
```
|
||||
|
||||
### 2. Stale Alerts in Alertmanager After Rule Changes
|
||||
|
||||
**Symptom**: Old alerts remain in Alertmanager's API even after Prometheus
|
||||
stops firing them.
|
||||
|
||||
**Root Cause**: Alertmanager caches alerts and waits for Prometheus to send
|
||||
"resolved" notifications. If the rule that generated them no longer exists,
|
||||
Prometheus never sends the resolution.
|
||||
|
||||
**Fix**: Restart Alertmanager to flush its cache:
|
||||
```bash
|
||||
docker restart alertmanager
|
||||
```
|
||||
|
||||
### 3. Blackbox ICMP False Positives
|
||||
|
||||
**Symptom**: `HostUnreachableICMP` fires for a host that is actually
|
||||
reachable via TCP.
|
||||
|
||||
**Root Cause**: Many cloud providers and firewalls block ICMP echo requests
|
||||
while allowing TCP traffic. PBS Remote (noris nbg1) is a typical example.
|
||||
|
||||
**Fix**: Either whitelist ICMP on the target firewall, or switch the
|
||||
Blackbox probe from ICMP to TCP port check:
|
||||
```yaml
|
||||
# In blackbox.yml, use tcp module instead of icmp
|
||||
modules:
|
||||
tcp_connect:
|
||||
prober: tcp
|
||||
timeout: 5s
|
||||
tcp:
|
||||
tls: false
|
||||
```
|
||||
|
||||
### 4. File Transfer to LXC Containers
|
||||
|
||||
`scp` and `pct push` often fail with permission errors. The reliable method
|
||||
is base64 encoding through SSH pipes:
|
||||
|
||||
```bash
|
||||
# Encode locally, decode remotely
|
||||
B64=$(base64 -w0 /local/file.py)
|
||||
ssh root@<pve-host> "echo '$B64' | base64 -d | \
|
||||
pct exec <ct-id> -- tee /remote/path/file.py >/dev/null"
|
||||
```
|
||||
|
||||
### 5. Docker Container Config Reload
|
||||
|
||||
When config files are bind-mounted from the host, edit the HOST file (not
|
||||
inside the container). Then reload:
|
||||
|
||||
```bash
|
||||
# Prometheus (hot reload via HTTP)
|
||||
docker exec prometheus wget -qO- --post-data= http://localhost:9090/-/reload
|
||||
|
||||
# Alertmanager (hot reload via HTTP)
|
||||
docker exec alertmanager wget -qO- --post-data= http://localhost:9093/-/reload
|
||||
|
||||
# If hot reload fails, restart the container
|
||||
docker restart prometheus
|
||||
docker restart alertmanager
|
||||
```
|
||||
|
||||
## RCA Workflow
|
||||
|
||||
When an alert arrives via webhook, the agent should:
|
||||
|
||||
1. **Parse the alert** — extract alertname, severity, instance, summary
|
||||
2. **Query Prometheus** — check related metrics, history, correlated alerts
|
||||
3. **Investigate the root cause** — SSH to affected hosts, check logs, APIs
|
||||
4. **Report findings** — concise summary with root cause and recommended fix
|
||||
|
||||
Example RCA script pattern (runs inside the monitoring container):
|
||||
```python
|
||||
import json, urllib.request, urllib.parse, socket
|
||||
|
||||
PROM = "http://127.0.0.1:9090"
|
||||
|
||||
def q(expr):
|
||||
url = f"{PROM}/api/v1/query?query={urllib.parse.quote(expr)}"
|
||||
with urllib.request.urlopen(url, timeout=5) as r:
|
||||
d = json.loads(r.read())
|
||||
return d.get("data", {}).get("result", [])
|
||||
|
||||
# Check probe success, HTTP status, TCP reachability
|
||||
# Cross-reference with direct curl/socket connections
|
||||
```
|
||||
|
||||
## Key Metrics Reference
|
||||
|
||||
| Metric | What it measures | Alert threshold |
|
||||
|--------|-----------------|----------------|
|
||||
| `up{job="node_exporter"}` | Node reachability | `== 0` for 2m |
|
||||
| `pve_up{id=~"node/.*"}` | PVE node status (FILTER!) | `== 0` for 2m |
|
||||
| `pve_ha_state{state=~"error\|fence\|freeze\|gone"}` | HA error states only | `== 1` for 1m |
|
||||
| `ceph_health_status` | Cluster health (0=OK,1=WARN,2=ERR) | `>= 1` |
|
||||
| `mysql_up` | Galera node alive | `== 0` for 1m |
|
||||
| `mysql_global_status_wsrep_cluster_size` | Galera nodes connected | `< 3` for 2m |
|
||||
| `probe_success{job="blackbox_icmp"}` | ICMP reachability | `== 0` for 2m |
|
||||
| `probe_success{job="blackbox_http"}` | HTTP endpoint | `== 0` for 2m |
|
||||
|
||||
## Infrastructure Layout (Reference)
|
||||
|
||||
- Proxmox host: `10.0.20.70`, SSH key: `~/.ssh/id_ed25519_proxmox`
|
||||
- Monitoring container: LXC CT 141 (IP `10.0.30.141`)
|
||||
- Prometheus: `localhost:9090` (in CT141 Docker)
|
||||
- Grafana: `localhost:3000` (in CT141 Docker, admin/Grafana2026!)
|
||||
- Alertmanager: `localhost:9093` (in CT141 Docker)
|
||||
- Hermes Gateway: `10.0.30.230:8644`
|
||||
- Telegram chat ID: `223926918`
|
||||
- PVE exporter: `pve-exporter:9221` (job `pve_api`)
|
||||
- Ceph exporter: `10.0.20.70:9283` (job `ceph`)
|
||||
- Galera nodes: `10.0.30.71/72/73` (job `mysqld_exporter`)
|
||||
- PBS Remote: `213.95.54.60:8007` (job `pbs-remote`, ICMP blocked)
|
||||
Reference in New Issue
Block a user