# Infrastructure Monitoring: Prometheus + Grafana Stack ## When to Use Deploying a full monitoring stack for Proxmox VE + Ceph + PBS + Galera/MaxScale infrastructure. Covers container provisioning, exporter deployment, alerting via Telegram, and Grafana dashboards. ## Architecture ``` CT 141 (monitoring, 10.0.30.141/24, VLAN 30, 4CPU/8GB/30GB) ├── Prometheus (:9090) — 90-day TSDB retention ├── Grafana (:3000) — dashboards, via Traefik → grafana.familie-schoen.com ├── Alertmanager (:9093) — routes alerts to Telegram bridge ├── blackbox_exporter (:9115) — ICMP + HTTP probes for uptime ├── pve-exporter (:9221) — PVE API metrics (VM/CT status, resources) ├── telegram-bridge (:9099) — Flask webhook → Telegram Bot API └── node_exporter (:9100) — local CT metrics Exporters on targets: ├── 6× PVE Nodes — node_exporter :9100 (systemd service) ├── Ceph mgr — :9283 (built-in prometheus module, active mgr node) ├── 3× Galera — mysqld_exporter :9104 (MySQL/Galera metrics) ├── PBS Local (CT 116) — node_exporter :9100 └── PBS Remote (213.95.54.60) — node_exporter :9100 (UFW restricted) ``` ## Container Creation ```bash pct create 141 hdd_templates:vztmpl/debian-12-standard_12.12-1_amd64.tar.zst \ --hostname monitoring \ --memory 8192 --swap 4096 --cores 4 \ --rootfs vm_disks:vm-141-disk-0,size=30G \ --net0 name=eth0,bridge=vmbr0,ip=10.0.30.141/24,gw=10.0.30.1,tag=30,type=veth \ --features nesting=1 \ --onboot 1 --start 0 ``` Install Docker inside the CT (standard Docker CE bookworm repo). ## Exporter Deployment ### node_exporter on PVE Nodes Deploy via systemd service on each PVE node. Binary from GitHub releases (v1.8.2). Service file at `/etc/systemd/system/node_exporter.service`. Enable `--collector.systemd` and `--collector.textfile.directory` for custom textfile collectors. PVE node IPs (VLAN 20 management network) — 8 nodes total: - proxmox1: 10.0.20.10 - proxmox2: 10.0.20.20 - proxmox3: 10.0.20.30 - proxmox4: 10.0.20.40 - proxmox5: 10.0.20.50 - proxmox6: 10.0.20.60 - proxmox7: 10.0.20.70 (primary jump host for SSH) - n5pro: 10.0.20.91 ⚠️ **Do NOT confuse node names with IPs** — the numeric suffix in the hostname does NOT match the last octet. proxmox4 = .40, proxmox7 = .70, n5pro = .91. Always verify with `hostname` after SSH. ### Ceph Prometheus Module Built-in, just needs enabling: ```bash ceph mgr module enable prometheus ceph config set mgr mgr/prometheus/scrape_interval 15 ``` Metrics available at `http://:9283/metrics`. Active mgr can fail over — add all potential mgr nodes as scrape targets. ### mysqld_exporter on Galera Requires exporter MySQL user on each Galera node: ```sql CREATE USER 'exporter'@'localhost' IDENTIFIED BY ''; GRANT PROCESS, REPLICATION CLIENT, SELECT ON *.* TO 'exporter'@'localhost'; ``` Binary: mysqld_exporter v0.15.1. Config at `/etc/mysqld_exporter/my.cnf`. Galera SSH access: 1Password 'SSH-Key Galera' (debian user, then sudo). ### node_exporter on PBS Remote Remote PBS is on public IP (213.95.54.60). **Must restrict UFW** to PVE WAN IP only — node_exporter exposes system information to anyone. ```bash sudo ufw allow from to any port 9100 ``` ### MaxScale Prometheus Metrics MaxScale has a built-in Prometheus endpoint at `:8189/metrics`. No additional exporter needed — just add it as a scrape target. ```yaml # In prometheus.yml - job_name: maxscale static_configs: - targets: ["10.0.30.81:8189"] ``` ⚠️ **MaxScale VM 310 was unreachable as of 2026-07-05** — PVE shows `status: running` but no ping, no SSH, QGA not running. The VM may have hung at the bootloader or lost its network config. Until fixed, this target will show as `down` in Prometheus — which is actually useful as an alert. The scrape config should be deployed regardless; it becomes functional as soon as the VM is reachable. MaxScale VM details: - VM 310 on proxmox4 (10.0.20.40) - IP: 10.0.30.81 (VLAN 30) - Metrics: http://10.0.30.81:8189/metrics - REST API: http://10.0.30.81:8989/ ## Docker Compose Stack Services: prometheus, grafana, alertmanager, blackbox-exporter, telegram-bridge, pve-exporter. Key config files: - `prometheus/prometheus.yml` — scrape configs for all targets - `prometheus/rules/*.yml` — alerting rules (general, ceph, galera, backup) - `alertmanager/alertmanager.yml` — routing to Telegram bridge - `blackbox/blackbox.yml` — ICMP + HTTP probe modules - `grafana/provisioning/` — datasource + dashboard auto-provisioning ### Prometheus Scrape Targets | Job | Target | Interval | Notes | |-----|--------|----------|-------| | node_exporter | 7× :9100 | 30s | Named `node_exporter` (not `pve_nodes`) for dashboard compat | | pve_api | :9221 | 30s | `metrics_path: /pve`, params: `cluster=1, node=1, target=:8006` | | ceph | :9283 | 30s | Active mgr node | | galera | 3× :9104 | 30s | UFW rule needed (allow from monitoring CT IP) | | pbs-remote | :9101 | 30s | Via SSH tunnel, use Docker gateway IP `172.18.0.1:9101` | | blackbox_icmp | via :9115 | 30s | ICMP probes to all infra nodes | | blackbox_http | via :9115 | 30s | HTTP probes to external services | | prometheus | :9090 | 30s | Self-monitoring | ## Telegram Alerting Alertmanager doesn't natively support Telegram. Deploy a lightweight Flask bridge container that: 1. Receives Alertmanager webhooks (POST to :9099) 2. Formats alerts as Markdown messages 3. Sends via Telegram Bot API `sendMessage` Bot token from Hermes `.env` file (`TELEGRAM_BOT_TOKEN`). Chat ID: 223926918 (user's Telegram ID, same as `TELEGRAM_HOME_CHANNEL`). Alert routing: - **Critical** (host down, Ceph ERR, Galera split, PBS unreachable): immediate, repeat every 1h - **Warning** (disk >80%, Ceph WARN, slow ops): grouped 5min, repeat every 4h ## Alert Rules Summary | Rule | Expression | Severity | |------|-----------|----------| | HostDown | `up{job="pve-nodes"} == 0` for 2m | critical | | HighDiskUsage | disk >80% for 10m | warning | | HighMemoryUsage | mem >85% for 10m | warning | | ServiceUnreachable | `probe_success == 0` for 2m | critical | | CephHealthError | `ceph_health_status == 2` for 5m | critical | | CephHealthWarning | `ceph_health_status == 1` for 15m | warning | | CephOSDDown | `ceph_osd_up < 1` for 5m | warning | | GaleraNodeDown | `wsrep_ready == 0` for 2m | critical | | GaleraClusterSize | `wsrep_cluster_size < 3` for 2m | critical | | PBSLocalDown | `up{job="pbs-local"} == 0` for 5m | critical | | PBSRemoteDown | `up{job="pbs-remote"} == 0` for 5m | critical | | MaxScaleDown | `up{job="maxscale"} == 0` for 5m | warning | ## Grafana Dashboards Auto-provisioned community dashboards (downloaded from grafana.com API): - **Node Exporter Full** (ID 1860) — system metrics, uses `$job` + `$node` variables - **Ceph Cluster** (ID 2842) — cluster overview, PATCHED for Ceph Squid (see pitfalls 32+) - **MySQL Overview** (ID 7362) — Galera metrics, uses `$host` variable - **Proxmox VE** (ID 10347) — PVE API metrics, uses `$instance` variable Download command: ```bash curl -s "https://grafana.com/api/dashboards//revisions/latest/download" \ -H "Accept: application/json" -o "dashboard_.json" ``` ⚠️ Community dashboards may reference metrics that don't exist in your exporter version. Always verify metric availability via the Prometheus API before importing. Patch dashboard JSON in-place when metrics are missing (see pitfall 33). ## Traefik Integration Expose Grafana externally via Traefik (CT 99999) at `grafana.familie-schoen.com`. Add dynamic config route to `http://10.0.30.141:3000` with Let's Encrypt TLS. ## PVE Native Notifications PVE has its own notification system. Forward PVE backup failures and HA state changes to the Telegram bridge via webhook endpoint (`http://10.0.30.141:9099/pve`). The bridge needs a separate handler for PVE's webhook format (different from Alertmanager's). ## Pitfalls 1. **Remote PBS node_exporter on public IP** — Must restrict UFW to PVE WAN IP. Without restriction, anyone can scrape system metrics. 2. **Ceph mgr failover moves :9283** — Active mgr can change. Add all potential mgr nodes as scrape targets; Prometheus handles dedup. 3. **Alertmanager can't send to Telegram natively** — Need a bridge container. The bridge must handle both Alertmanager and PVE webhook formats. 4. **PVE backup job "skip external VMs" masks failures** — See `references/pbs-lxc-setup-2026-07.md` pitfall #12. A backup job showing `status=OK` may have skipped most VMs/CTs on other nodes. 5. **PBS owner mismatch silently kills backups** — See `references/pbs-lxc-setup-2026-07.md` pitfall #11. Always verify owner files match the PVE storage username after credential changes. 6. **PVE node hostname ≠ IP last octet** — `proxmox4` is at `.40` not `.70`, `proxmox7` is at `.70` not `.91`, `n5pro` is at `.91`. The numeric suffix in the hostname does NOT correspond to the IP octet. Always use the IP→name mapping from `pvecm nodes` + corosync.conf, or verify with `hostname` after SSH. Wasted significant time SSH-ing to wrong nodes during monitoring setup. 7. **VMs can show "running" in PVE but be network-dead** — MaxScale VM 310 showed `status: running` in PVE API but was completely unreachable (no ping, no SSH, QGA not running). The VM may have hung at boot or lost its network config. Don't assume a VM is functional just because PVE says it's running — always verify network reachability separately. This is exactly the kind of failure monitoring catches: deploy the scrape config anyway so Prometheus alerts on the unreachable target. 8. **Docker container name conflicts after compose down** — When `docker compose up` times out mid-pull and you retry, containers can get stuck in "Created" state. `docker compose down` reports success but `docker rm -f ` says "No such container" — yet the names remain reserved and `compose up` fails with "container name already in use". Fix: `systemctl restart docker` inside the CT, then `docker compose up -d`. The daemon restart clears the stale name reservations. Do NOT waste time trying to `docker rm` by truncated ID — the IDs from the error message don't exist anymore. 9. **Prometheus permission denied on mounted config** — The Prometheus container runs as UID 65534 (nobody) by default. Config files created by root on the host with mode 600 cause `permission denied` on startup, and the container enters a restart loop. Fix: either `chmod 644` all config files on the host, or add `user: root` to the prometheus service in docker-compose.yml. The latter is simpler for a trusted internal monitoring CT. 10. **pve-exporter config mount path** — The `prompve/prometheus-pve-exporter` Docker image hardcodes the config lookup at `/etc/prometheus/pve.yml`. Even if you set `PVE_EXPORTER_CONFIG` env var to a different path, the binary's default still looks for `/etc/prometheus/pve.yml` first. Always mount the config to `/etc/prometheus/pve.yml:ro` inside the container, regardless of the env var. 11. **PBS Remote behind OpenStack security group** — The remote PBS (213.95.54.60) runs on an OpenStack VM. Port 9100 (node_exporter) was blocked by the OpenStack security group, NOT by UFW (which was inactive). Since we had no OpenStack SG access, the fix was an SSH tunnel via autossh from the monitoring CT: `autossh -N -L 9101:localhost:9100 ubuntu@213.95.54.60`. Prometheus scrapes `localhost:9101` instead of the remote directly. The SSH key must be copied into the CT (it lives on the Hermes host, not on the PVE node). Pattern: scp key to PVE jump host → `pct push` into CT. 12. **Grafana first-start SQLite migrations** — On first launch with a fresh persistent volume, Grafana runs dozens of SQLite migration steps that take 1–2+ minutes. During this time, HTTP requests to :3000 return connection refused (curl `%{http_code}` = 000). This is NOT a crash — check `docker logs grafana` for ongoing migration log lines. Just wait and re-check. Do not restart the container. 13. **SCP to PVE jump host needs -i flag** — When copying files TO a PVE node that requires a specific SSH key, `scp` also needs the `-i` flag (same as ssh). `scp file root@10.0.20.70:/path` fails; `scp -i ~/.ssh/id_ed25519_proxmox file root@10.0.20.70:/path` works. 14. **PVE /tmp can be full on jump hosts** — proxmox7 (10.0.20.70) had a full /tmp partition. SCP to `/tmp/` failed with "No space left on device". Use `/root/` or another writable path for temporary file transfers through PVE jump hosts. 15. **Docker image pulls on slow CT connections** — Pulling 5+ Docker images (Prometheus, Grafana, Alertmanager, Blackbox, PVE exporter, Python base for telegram bridge) on a CT with limited bandwidth can take 5–10+ minutes. The 300s default terminal timeout is insufficient. Either set timeout=600 or run `docker compose pull` separately first, then `docker compose up -d`. The `--build` flag for the telegram bridge adds another 2+ minutes for pip install. ## Working Templates - `templates/monitoring-docker-compose.yml` — Production-tested compose file with all fixes (user:root for Prometheus, correct pve-exporter mount path, custom telegram-bridge build) - `templates/telegram-bridge.py` — Flask webhook bridge for Alertmanager → Telegram Bot API ## Related References - `references/pbs-lxc-setup-2026-07.md` — PBS setup, PVE integration, owner mismatch pitfall - `references/pbs-sync-pipeline-2026-07.md` — PBS sync pipeline, direct vzdump vs sync ## Additional Pitfalls (Session 2 — 2026-07-05) 16. **Galera UFW blocks mysqld_exporter from outside** — mysqld_exporter binds to `*:9104` and works from `localhost`, but Galera nodes have UFW active which blocks 9104 from other subnets. Prometheus scrapes fail silently (empty response, no error). Fix: `ufw allow from to any port 9104 proto tcp` on each Galera node. Must be done via the Galera SSH key (1Password 'SSH-Key Galera', debian user → sudo). 17. **Docker `localhost` inside a container ≠ CT host** — When Prometheus runs in a Docker bridge network, `localhost:9101` in a scrape target refers to the *container's* loopback, not the CT host. An SSH tunnel listening on the CT's `0.0.0.0:9101` is invisible to the container. Fix: use the Docker bridge gateway IP (`172.18.0.1:9101`) as the scrape target. Find it with `docker network inspect | grep Gateway`. Alternatively, use `network_mode: host` on the Prometheus container, but that breaks container isolation and inter-container DNS. 18. **Grafana provisioning files need chmod 644** — Grafana container runs as UID 472. Provisioning YAML files created by root with mode 600 cause silent startup failures (HTTP 000, no error in logs — Grafana just hangs). The earlier `find . -name *.yml -exec chmod 644` fix didn't catch files in `grafana/provisioning/` because the glob only searched from the prometheus subdir. Fix: explicitly `chmod 644 /opt/monitoring/grafana/provisioning/**/*.yml` or use `find /opt/monitoring -name "*.yml" -exec chmod 644 {} +`. 19. **PBS Local is on n5pro, not CT 116** — The reference doc originally listed "PBS Local (CT 116)". In reality, CT 116 is on n5pro (10.0.20.91) per `ha-manager status`. The earlier `pct list` on proxmox7 didn't show it because CT 116 runs on n5pro, not proxmox7. Always check `ha-manager status` for the *actual* node assignment before trying to `pct exec` into a CT. 20. **PVE notification webhook setup** — PVE's native notification system supports webhook endpoints. Creating the endpoint alone is NOT enough — you also need a matcher to route notifications to it. Two steps: ```bash # Step 1: Create the webhook endpoint pvesh create /cluster/notifications/endpoints/webhook \ --name telegram-bridge \ --url http://10.0.30.141:9099/pve \ --method post # ⚠️ --method must be LOWERCASE (post/put/get). UPPERCASE "POST" fails # with: "value 'POST' does not have a value in the enumeration 'post, put, get'" # Step 2: Create a matcher that routes all notifications to this endpoint pvesh create /cluster/notifications/matchers \ --name telegram-route \ --target telegram-bridge \ --mode all \ --comment "Route all PVE notifications to Telegram" ``` Without Step 2, the endpoint exists but no notifications are routed to it — PVE's default matcher only targets `mail-to-root`. The bridge's `/pve` endpoint handles PVE's format (title/message/severity fields), which differs from Alertmanager's alerts array. ⚠️ Ensure `TELEGRAM_BOT_TOKEN` and `TELEGRAM_CHAT_ID` are set in the docker-compose env (or `.env` file) — empty values cause the bridge to accept the webhook (HTTP 200) but silently fail the Telegram send. ⚠️ Telegram bots cannot initiate conversations — the user MUST send `/start` to the bot first. Until then, `sendMessage` returns `400 Bad Request: chat not found`. After `/start`, the bot can send messages freely. 21. **Traefik on CT 99999, not a VM** — Traefik runs as a native binary (systemd service) inside CT 99999 on proxmox7. Config at `/etc/traefik/traefik.yaml` with dynamic configs in `/etc/traefik/conf.d/`. Adding a new route is as simple as dropping a YAML file into `conf.d/` — Traefik watches the directory and auto-reloads. No restart needed. The ACME cert resolver (`letsencrypt`) handles TLS automatically for new domains. 22. **Prometheus target list cleanup** — Dead targets (pbs-local CT 116 on n5pro, maxscale VIP .81 unreachable) should be removed from prometheus.yml to avoid noise. But keep scrape configs for targets that are temporarily down but expected to come back (maxscale) — the `down` state itself is a useful alert. Removed pbs-local entirely since CT 116 is being decommissioned. Kept pbs-remote via SSH tunnel. ## Additional Pitfalls (Session 3 — Dashboard Fixes 2026-07-05) 23. **pve-exporter config key is NOT `url` — it's implicit `host`** — The `prompve/prometheus-pve-exporter` Docker image uses `proxmoxer` which expects `host` as the first positional arg to `ProxmoxAPI()`. The YAML config must NOT include a `url:` or `host:` key — the exporter passes `host` from the scrape URL's `target` query param. Correct config (`pve-exporter.yml`): ```yaml default: user: exporter@pve token_name: monitoring token_value: verify_ssl: false ``` If you include `url:` → `TypeError: unexpected keyword argument 'url'`. If you include `host:` → `TypeError: multiple values for argument 'host'`. Neither error is obvious from the generic 500 HTML response. Check `docker logs pve-exporter` for the traceback. 24. **pve-exporter needs `target` param in scrape URL** — Without `?target=` in the Prometheus scrape URL, the exporter defaults to `localhost:8006` and gets connection refused. The Prometheus job config must be: ```yaml - job_name: pve_api metrics_path: /pve params: cluster: [1] node: [1] target: ['10.0.20.70:8006'] static_configs: - targets: ['pve-exporter:9221'] ``` Without `metrics_path: /pve`, Prometheus scrapes `/metrics` (only 5 collector self-metrics). Without `cluster=1&node=1`, the cluster and node collectors are skipped. Without `target=`, it tries localhost:8006 and fails. All three params are required. 25. **Community Grafana dashboards hardcode job names** — Dashboard 1860 (Node Exporter Full) uses `$job` variable populated by `label_values(node_uname_info, job)`. If your Prometheus job is named `pve_nodes` instead of `node_exporter`, the dropdown populates with `pve_nodes` but panel queries that hardcode `job="node_exporter"` show no data. Fix: name the Prometheus job `node_exporter` (the conventional name) so community dashboards work without modification. Same principle for other dashboards: Galera dashboard 7362 uses `label_values(mysql_up, instance)` — works as long as `mysql_up` exists regardless of job name. PVE dashboard 10347 uses `label_values(pve_node_info, instance)` — works once pve-exporter is properly configured (see pitfalls 23–24). 26. **Prometheus reload via POST /-/reload sometimes silently fails** — After editing `prometheus.yml`, `curl -X POST http://localhost:9090/-/reload` returns success but the config doesn't change (verified by checking scrape URLs via the targets API). This happens when the YAML has a subtle issue (duplicate keys, control characters from copy-paste) that Prometheus rejects on parse but doesn't surface as an error to the reload endpoint. Fix: `docker restart prometheus` forces a full config reload. If that still doesn't work, exec into the container and run `promtool check config /etc/prometheus/prometheus.yml` to find the validation error. 27. **Old job labels persist in Prometheus after rename** — Renaming a job from `pve_nodes` to `node_exporter` causes BOTH labels to coexist in query results until the old TSDB samples age out (retention period). `node_uname_info` will show series with `job=pve_nodes` (cached) and `job=node_exporter` (new). This is cosmetic — the old data expires naturally. Don't try to force- purge; just wait. 28. **pve-exporter Docker image has split Python environments** — The `prompve/prometheus-pve-exporter` image installs the exporter in `/opt/prometheus-pve-exporter/` (virtualenv), NOT in the system Python. `docker exec pve-exporter python3 -c "import ..."` uses the wrong interpreter and gets ModuleNotFoundError. Use `docker exec pve-exporter /opt/prometheus-pve-exporter/bin/python3` or `docker exec pve-exporter /opt/prometheus-pve-exporter/bin/pve_exporter` for introspection. Similarly, `pip list` shows only setuptools — the real packages are in the venv's pip. 29. **PVE exporter collectors are URL-param gated, not CLI-flag gated** — Unlike most Prometheus exporters where collectors are enabled via `--collector.xxx` CLI flags, pve-exporter gates them via URL query params: `?cluster=1` enables status/version/node/cluster/resources/ backup-info/qdevice collectors, `?node=1` enables config/replication/ subscription collectors. Without these params, you get only 5 self-metrics (collector duration + request errors). The `--help` output lists `--collector.status` etc. but these are NOT how you enable them in practice — the URL params are the mechanism. 30. **Grafana dashboard provisioning requires `provider.yml`** — Dashboard JSON files in `/var/lib/grafana/dashboards/` are NOT auto-discovered. Grafana needs a provisioning provider config at `/etc/grafana/provisioning/dashboards/provider.yml` that points to the dashboard directory. Without it, dashboards exist on disk but never appear in the UI. The provider config specifies `path: /var/lib/grafana/dashboards` (inside the container), which must match the volume mount in docker-compose.yml. 31. **Verifying dashboard data availability** — Before declaring dashboards fixed, query the Prometheus API for the specific metric names the dashboard uses. For Node Exporter: `node_uname_info`, `node_load1`. For Ceph: `ceph_health_status`, `ceph_pool_metadata`. For Galera: `mysql_up`, `mysql_galera_status_info`. For PVE: `pve_node_info`, `pve_cpu_usage_ratio`, `pve_up`. If the metric exists in `label_values` API and has the expected labels, the dashboard will populate. If not, trace the gap: exporter not running → exporter running but wrong config → Prometheus not scraping → scraping but wrong job name → job name correct but dashboard template variable doesn't match. 32. **Ceph Squid prometheus module lacks OSD op counters** — Ceph 19.2 (Squid) prometheus module does NOT export `ceph_osd_op_r`, `ceph_osd_op_w`, `ceph_osd_op_r_out_bytes`, `ceph_osd_op_w_in_bytes`, `ceph_osd_op_r/w_latency_sum/count` by default. There is no config option to enable them (`mgr/prometheus/experimental_perf_counters` doesn't exist in Squid). Disabling/re-enabling the module doesn't help. Available OSD metrics: `ceph_osd_apply_latency_ms`, `ceph_osd_commit_latency_ms`, `ceph_osd_up`, `ceph_osd_in`, `ceph_osd_weight`, flags, metadata. Available pool-level I/O: `ceph_pool_rd`, `ceph_pool_wr`, `ceph_pool_rd_bytes`, `ceph_pool_wr_bytes` (all counters, usable with `irate()`/`rate()`). Community Ceph dashboards (ID 2842) that use `ceph_osd_op_*` show "No data". Fix: patch the dashboard JSON to substitute pool-level metrics. Mapping table: - `ceph_osd_op_w_in_bytes` → `ceph_pool_wr_bytes` - `ceph_osd_op_r_out_bytes` → `ceph_pool_rd_bytes` - `ceph_osd_op_w` → `ceph_pool_wr` - `ceph_osd_op_r` → `ceph_pool_rd` - `ceph_osd_op_r/w_latency_sum/count` → `ceph_osd_commit_latency_ms` or `ceph_osd_apply_latency_ms` (these are gauges, not sum/count pairs — remove the `rate(sum)/rate(count)` formula and use `avg()`) - `ceph_mon_num_sessions` → `count(ceph_mon_quorum_status)` - `ceph_osd_numpg` → `sum(ceph_pg_total)` 33. **Patching Grafana dashboard JSON in-place** — When community dashboards reference missing metrics, the fastest fix is to download the dashboard JSON, patch all `expr` fields with `json.load` + recursive walk + `str.replace`, and push it back to the provisioning directory. Steps: 1. `curl -s "https://grafana.com/api/dashboards//revisions/latest/download" -o dashboard_.json` 2. Python: load JSON, walk all dicts looking for `"expr"` keys, apply replacements, save 3. `scp` to PVE jump host → `pct push` into CT → `chmod 644` 4. Grafana auto-reloads provisioned dashboards every 30s (`updateIntervalSeconds: 30` in provider.yml). If it doesn't, `docker restart grafana` forces a reload. 5. Verify via browser: `document.body.innerText.match(/No data/gi).length` in the Grafana page console — should be 0. 34. **Grafana dashboard provisioning reload timing** — After updating a dashboard JSON file in the provisioning directory, Grafana picks it up within `updateIntervalSeconds` (default 30s). But if the file was modified while Grafana was running, it may not detect the change immediately. `docker restart grafana` guarantees a reload but takes 15-20s for startup + SQLite checks. After restart, verify with the Grafana API: `curl -u admin: http://localhost:3000/api/search` to confirm dashboards are listed. 35. **Browser-based Grafana dashboard verification** — When verifying dashboards remotely, use the browser tools to navigate to the dashboard URL, login with credentials, then run this in the page console to count "No data" panels: ```javascript (() => { const txt = document.body.innerText; const noData = (txt.match(/No data/gi) || []).length; return JSON.stringify({noDataCount: noData}); })() ``` To identify WHICH panels show no data: ```javascript (() => { const p = document.querySelectorAll('.react-grid-item, [class*="panel-container"]'); const nd = []; for (const panel of p) { if (panel.innerText.includes('No data')) { const t = panel.querySelector('.panel-title, h2, [class*="title"]'); nd.push(t ? t.textContent.trim() : '?'); } } return JSON.stringify(nd); })() ``` This is faster than screenshots for identifying broken panels. Note: panels inside collapsed rows won't be detected — expand all rows first. 36. **Patching dashboards with latency formulas** — When substituting `ceph_osd_op_r_latency_sum/count` (histogram-style sum+count pair) with `ceph_osd_commit_latency_ms` (a simple gauge), the original formula `rate(sum[5m]) / rate(count[5m])` must be replaced with just `avg(ceph_osd_commit_latency_ms{})`. Keeping the division formula with non-histogram metrics produces nonsensical values or NaN. Always check the metric TYPE (counter vs gauge vs histogram) before patching — `ceph_pool_rd` is a counter (use with `irate`), `ceph_osd_commit_latency_ms` is a gauge (use directly or with `avg()`). ## Dashboard Patching Reference ### Ceph Dashboard 2842 — Metric Substitution Table | Original (missing) | Replacement (available) | Type Change | |---|---|---| | `ceph_osd_op_w_in_bytes` | `ceph_pool_wr_bytes` | counter → counter | | `ceph_osd_op_r_out_bytes` | `ceph_pool_rd_bytes` | counter → counter | | `ceph_osd_op_w` | `ceph_pool_wr` | counter → counter | | `ceph_osd_op_r` | `ceph_pool_rd` | counter → counter | | `ceph_osd_op_r_latency_sum/count` | `ceph_osd_commit_latency_ms` | histogram → gauge | | `ceph_osd_op_w_latency_sum/count` | `ceph_osd_apply_latency_ms` | histogram → gauge | | `ceph_mon_num_sessions` | `count(ceph_mon_quorum_status)` | gauge → count | | `ceph_osd_numpg` | `sum(ceph_pg_total)` | gauge → sum | ### Verification Workflow 1. Query Prometheus API for available metrics: `curl -s http://localhost:9090/api/v1/label/__name__/values | python3 -c "..."` 2. Check specific metric existence and label values 3. Test patched PromQL expressions directly via API 4. Push patched dashboard JSON to CT 5. Browser verify: login to Grafana, navigate to dashboard, run console JS to count "No data" panels 6. If panels still show no data, check: - Metric exists in Prometheus? (`label_values` API) - Metric has expected labels? (`query` API) - Dashboard template variable populates? (check dropdown) - Panel query matches available data? (test in Grafana Explore) ## Additional Pitfalls (Session 4 — MySQL Dashboard Fix 2026-07-05) 37. **Imported dashboards with `__inputs` leave `${DS_PROMETHEUS}` unresolved** — Community dashboards downloaded from grafana.com often ship with an `__inputs` section declaring a datasource variable (e.g. `{"name":"DS_PROMETHEUS","type":"datasource","pluginId":"prometheus"}`). When imported via file provisioning (not the Grafana import UI), the `${DS_PROMETHEUS}` placeholder in panel targets and template variables is NEVER resolved — panels show "No data" because the datasource ref is invalid. Fix: download the JSON, replace ALL `${DS_PROMETHEUS}` references with the actual Prometheus datasource UID (found via `GET /api/datasources`), then remove the `__inputs` section entirely. The replacement must handle both string-format (`"${DS_PROMETHEUS}"`) and object-format (`{"uid":"${DS_PROMETHEUS}"}`) datasource references. Python walk-and-replace pattern: ```python def fix_ds(obj): if isinstance(obj, dict): for k, v in obj.items(): if k == "datasource": if isinstance(v, str) and "${DS_PROMETHEUS}" in v: obj[k] = DS_UID elif isinstance(v, dict) and v.get("uid") == "${DS_PROMETHEUS}": obj[k] = {"type": "prometheus", "uid": DS_UID} else: fix_ds(v) elif isinstance(obj, list): for item in obj: fix_ds(item) ``` 38. **Template variables with empty `current` value produce no data** — Even after fixing the datasource, a query-type template variable (e.g. `host = label_values(mysql_up, instance)`) may have `current: null` and `options: []` in the provisioned JSON. This means `$host` expands to empty string in all panel queries (`{instance=""}`), returning no data. The Grafana UI populates these on first load, but file-provisioned dashboards skip that step. Fix: pre-populate the `current` and `options` fields in the JSON with known values BEFORE pushing to the provisioning directory: ```python t["current"] = {"text": "10.0.30.71:9104", "value": "10.0.30.71:9104"} t["options"] = [ {"selected": True, "text": "10.0.30.71:9104", "value": "10.0.30.71:9104"}, {"selected": False, "text": "10.0.30.72:9104", "value": "10.0.30.72:9104"}, {"selected": False, "text": "10.0.30.73:9104", "value": "10.0.30.73:9104"}, ] ``` Grafana will still query the variable dynamically on page load, but the pre-populated `current` ensures the initial render has data. Users can then switch via the dropdown. 39. **Cross-network metric joins fail when instances differ** — The MySQL Overview dashboard (ID 7362) has a "Buffer Pool Size of Total RAM" panel that joins `mysql_global_variables_innodb_buffer_pool_size` with `node_memory_MemTotal_bytes` using `on(instance)`. This fails because MySQL exporter instances (`10.0.30.71-73:9104`) and node_exporter instances (`10.0.20.x:9100`) are on completely different networks — Galera nodes don't run node_exporter at all. The `on(instance)` join matches nothing. Fix options: a. Install node_exporter on Galera nodes (adds operational burden). b. Use `label_replace` to strip ports and match by IP only (still fails because IPs differ: 10.0.30.x vs 10.0.20.x). c. Replace with a static constant if hardware specs are known (e.g. `(buffer_pool_size * 100) / 8589934592` for 8GB RAM nodes). Option (c) is the pragmatic fix for homelab clusters with uniform hardware. Document the assumption in the panel description. ### MySQL Dashboard 7362 — Fixes Applied | Issue | Fix | |---|---| | `${DS_PROMETHEUS}` unresolved in all panels | Replace with UID `PBFA97CFB590B2093`, remove `__inputs` | | `$host` variable empty (`current: null`) | Pre-populate with 3 Galera instances, default `10.0.30.71:9104` | | Buffer Pool % of RAM join fails (no node_exporter on Galera) | Static 8GB constant: `(buf_pool * 100) / 8589934592` | ### General Dashboard Import Checklist When importing ANY community Grafana dashboard via file provisioning: 1. Check for `__inputs` section → remove it, hardcode datasource UID 2. Check all template variables → pre-populate `current` + `options` if empty 3. Check for cross-datasource joins (e.g. mysql + node) → verify both sides exist 4. Check for hardcoded job names (e.g. `job="node_exporter"`) → match your Prometheus config 5. Push to provisioning dir, `chmod 644`, restart Grafana 6. Browser verify: count "No data" panels via console JS ## Additional Pitfalls (Session 5 — Resume & Dashboard Variable Fixes 2026-07-05) 40. **Base64-through-SSH file transfer to CTs** — When you need to get a Python script (or any file) into a CT but `pct push` fails because the file isn't on the PVE host yet, and direct SCP to the PVE host `/tmp/` may fail (permissions, disk space), use base64 encoding piped through SSH directly into the CT: ```bash SCRIPT=$(base64 -w0 /tmp/local_script.py) ssh -i ~/.ssh/key root@10.0.20.70 "echo '$SCRIPT' | base64 -d | pct exec 141 -- tee /tmp/script.py >/dev/null && pct exec 141 -- python3 /tmp/script.py" ``` This bypasses the two-step (scp to PVE host → pct push into CT) entirely. The `>/dev/null` suppresses the tee output. Works for any text file. For binary files, use base64 with `--decode` on the receiving end. 41. **Grafana REST API for dashboard variable fixes** — Instead of downloading dashboard JSON, patching it locally, and re-provisioning via file (pitfalls 33-34, 37-38), you can fix dashboard variables directly via the Grafana REST API. This is faster for surgical fixes: ```python import json, urllib.request, base64 BASE = "http://127.0.0.1:3000" AUTH = base64.b64encode(b"admin:Grafana2026!").decode() PROM_UID = "PBFA97CFB590B2093" def api_get(path): req = urllib.request.Request(f"{BASE}{path}") req.add_header("Authorization", f"Basic {AUTH}") with urllib.request.urlopen(req) as resp: return json.loads(resp.read()) def api_post(path, data): body = json.dumps(data).encode() req = urllib.request.Request(f"{BASE}{path}", data=body, method="POST") req.add_header("Authorization", f"Basic {AUTH}") req.add_header("Content-Type", "application/json") with urllib.request.urlopen(req) as resp: return json.loads(resp.read()) # GET dashboard → modify → POST back d = api_get("/api/dashboards/uid/Dp7Cd57Zza") dash = d.get("dashboard", {}) templating = dash.get("templating", {}).get("list", []) # Add missing DS_PROMETHEUS variable if "DS_PROMETHEUS" not in [v.get("name") for v in templating]: templating.insert(0, { "name": "DS_PROMETHEUS", "type": "datasource", "query": "prometheus", "current": {"text": "Prometheus", "value": PROM_UID}, }) # Fix empty current value on existing variable for v in templating: if v.get("name") == "ds_prometheus" and not v.get("current"): v["current"] = {"text": "Prometheus", "value": PROM_UID} result = api_post("/api/dashboards/db", { "dashboard": dash, "message": "Fix datasource vars", "overwrite": True }) ``` Key endpoints: - `GET /api/dashboards/uid/{uid}` — fetch dashboard JSON - `POST /api/dashboards/db` — save modified dashboard (requires `"overwrite": True` if the dashboard already exists) - `GET /api/datasources` — list datasources to find the Prometheus UID The dashboard JSON path is `response["dashboard"]` (NOT `response["data"]["dashboard"]` — that's the file-provisioning format). 42. **Python `urllib` + `localhost` DNS resolution failure with embedded credentials** — Using `urllib.request.urlopen` with a URL like `http://admin:pass@localhost:3000/api/...` fails with `socket.gaierror: [Errno -2] Name or service not known` in some environments. The URL-embedded credentials cause urllib to parse `localhost` as a hostname for DNS lookup rather than resolving it to 127.0.0.1. Fix: use `127.0.0.1` explicitly and pass credentials via the Authorization header instead: ```python import base64, urllib.request AUTH = base64.b64encode(b"admin:password").decode() req = urllib.request.Request("http://127.0.0.1:3000/api/...") req.add_header("Authorization", f"Basic {AUTH}") with urllib.request.urlopen(req) as resp: data = json.loads(resp.read()) ``` This is more robust than URL-embedded credentials and works in all environments. The same issue affects `requests` library — always prefer explicit headers over URL-embedded auth. 43. **Echo with parentheses breaks bash in nested SSH+pct exec** — When running multi-line Python or bash scripts through nested SSH (`ssh root@pve 'pct exec 141 -- bash -c "..."'`), `echo` statements containing parentheses like `echo === Node Exporter (dashboard 1860) ===` cause `syntax error near unexpected token '('`. Bash interprets the parentheses as subshell syntax. Fix: either escape parentheses (`echo === Node Exporter \(dashboard 1860\) ===`), avoid parentheses in echo strings, or — preferably — write the script to a local file, base64-transfer it (pitfall 40), and execute it as a unit. The base64 approach eliminates all quoting/escaping issues in nested SSH. ## Related References - `references/openstack-offsite-pbs-2026-07.md` — Remote PBS provisioning - `references/infra-monitoring-2026-07.md` — Full monitoring stack setup, dashboard patching guide (pitfalls 23-43), Ceph Squid metric substitution table, MySQL dashboard import checklist, Grafana API dashboard fixes