Initial commit: Hermes Agent Skills collection

This commit is contained in:
Debian
2026-07-12 19:02:59 +00:00
commit e9cc106625
789 changed files with 233126 additions and 0 deletions
@@ -0,0 +1,549 @@
# Ceph Recovery, EC Reweight, PG Merge & Backfillfull
## EC4+1 Reweight Trap — CRUSH Returns 2147483647
### Problem
With exactly 5 HDD OSDs and EC profile `k=4 m=1` (needs exactly 5 OSDs),
rewighting any OSD (e.g. `ceph osd reweight osd.7 0.65`) causes CRUSH
to return `2147483647` ("no OSD available") in the up-set for affected
PGs. These PGs become `active+clean+remapped` and **cannot recover**
there is no alternative OSD to place them on.
```
8.20 active+clean+remapped [6,1,8,10,2147483647] 6
8.78 active+clean+remapped [1,6,8,2147483647,10] 1
```
### Diagnosis
```bash
# Check EC profile
ceph osd erasure-code-profile get ec41-hdd
# k=4 m=1 → needs 5 OSDs
# Count HDD OSDs
ceph osd tree | grep hdd | wc -l
# 5 → exactly the minimum. Reweighting breaks CRUSH.
# Check for 2147483647 in up-sets
ceph pg dump pgs_brief 2>/dev/null | grep 2147483647 | wc -l
```
### Fix
**Reset reweight to 1.0 for all EC-participating OSDs.** With only 5
HDD OSDs and EC4+1, there is no spare capacity for reweighting. Every
OSD must participate at full weight.
```bash
ceph osd reweight osd.7 1.0
ceph osd reweight osd.10 1.0
```
After resetting, recovery resumes (67 MiB/s, 16 obj/s observed).
### General Rule
- **Replicated pools** (size=3): reweighting is safe — CRUSH redirects
to other OSDs in the failure domain.
- **Erasure-coded pools**: reweighting is only safe when there are MORE
OSDs than `k+m`. With exactly `k+m` OSDs, reweighting traps PGs.
- **Long-term fix**: add more HDD OSDs (e.g. on the ubuntu host with
2×2.8 TB spare) to give CRUSH room to redistribute.
## Backfillfull Sticky Flag
### Problem
osd.10 at 80% utilization has a `backfillfull` flag. Even after raising
the ratio:
```bash
ceph osd set-backfillfull-ratio 0.95
ceph osd set-nearfull-ratio 0.95
```
The flag persists and blocks backfill on 12+ PGs:
```
HEALTH_WARN ... 1 backfillfull osd(s); Low space hindering backfill:
12 pgs backfill_toofull; 6 pool(s) backfillfull
```
### Attempted Fixes That Don't Work
```bash
# Does not exist in this Ceph version (Reef/Quincy):
ceph osd clear-backfillfull osd.10 # → EINVAL
# Cannot be modified at runtime:
ceph tell osd.7 injectargs --mon_osd_backfillfull_ratio 0.95 # → EPERM
# Forces backfill on PGs that need it, but doesn't clear toofull:
ceph osd pool force-backfill hdd_disk # instructs PGs to force-backfill
# → PGs still show backfill_toofull
```
### What Happened
The `backfillfull` flag was set when osd.10 crossed the OLD threshold
(default 0.85). Raising the ratio afterward does NOT retroactively
clear the flag — it only prevents future flag-setting. The OSD remains
flagged until its utilization drops below the new threshold naturally
(via recovery moving data away).
### Mitigation
- **Patience**: Once the reweight trap is fixed and recovery proceeds,
data redistributes and osd.10 utilization drops, eventually clearing
the flag.
- **PG merge**: Reducing PG count (128→32) eliminates some remapped PGs
entirely, reducing the number of PGs that need to backfill to the
stuck OSD.
- **Do NOT reweight osd.10** to reduce its load — with EC4+1 and only 5
HDD OSDs, reweighting makes things worse (see above).
## PG Merge (128 → 32)
### Procedure
```bash
# 1. Disable autoscaler for the pool
ceph osd pool set cephfs_data pg_autoscale_mode off
ceph osd pool set cephfs_metadata pg_autoscale_mode off
# 2. Set pgp_num first (must be ≤ pg_num)
ceph osd pool set cephfs_data pgp_num 32
ceph osd pool set cephfs_metadata pgp_num 32
# 3. Set pg_num_target (async merge preparation)
ceph osd pool set cephfs_data pg_num_target 32
ceph osd pool set cephfs_metadata pg_num_target 32
# 4. Force the merge — works when all PGs are clean
ceph osd pool set cephfs_data pg_num 32
ceph osd pool set cephfs_metadata pg_num 32
# 5. Verify
ceph osd pool get cephfs_data pg_num # → 32
ceph osd pool get cephfs_metadata pg_num # → 32
```
### When It Works
- All PGs in the pool must be `active+clean` for the merge to execute.
- If PGs are `remapped` or `backfilling`, the merge waits.
- With `pg_autoscale_mode off`, the merge is manual and deterministic.
### Impact
Reducing `cephfs_data` (0 bytes stored) and `cephfs_metadata` (85 KiB
stored) from 128→32 PGs eliminates 384 PG instances across OSDs, helping
with the "too many PGs per OSD" health warning (262 > 250).
## Recovery Speed Tuning
```bash
# Increase concurrent backfills per OSD (default 1)
ceph config set osd osd_max_backfills 5
# Increase recovery op priority (default 1, lower = higher priority)
ceph config set osd osd_recovery_op_priority 3
# Ensure no recovery sleep (default 0)
ceph config get osd osd_recovery_sleep # → 0.000000 (already optimal)
```
### Monitoring Recovery Progress
```bash
# Overall
ceph status | grep -E 'recovery|pgs:'
# Count remapped/backfill PGs
ceph pg dump pgs_brief 2>/dev/null | tail -n +2 | awk '{print $2}' | grep -c remapped
ceph pg dump pgs_brief 2>/dev/null | tail -n +2 | awk '{print $2}' | grep -c backfill_toofull
# Recovery rate
ceph status | grep -A2 'recovery:'
# → recovery: 67 MiB/s, 16 objects/s
```
### Diagnosing Stalled Recovery (2026-07-07)
When recovery appears stalled (misplaced objects constant for hours),
check these in order BEFORE tuning recovery speed:
1. **EC reweight trap**: `ceph pg dump pgs_brief | grep 2147483647`
— If PGs contain `2147483647` ("no OSD"), CRUSH cannot place them.
Reset all EC-participating OSDs to reweight 1.0. See section above.
2. **Backfillfull sticky flag**: `ceph health detail | grep backfillfull`
— An OSD flagged `backfillfull` blocks all backfill to it, even after
raising the ratio. The flag is set when crossing the OLD threshold and
does NOT clear retroactively. See "Backfillfull Sticky Flag" section.
3. **Stale `rbd rm` process**: A previous `rbd rm` stuck in I/O holds a
watcher that blocks new operations. Check `ps aux | grep 'rbd.*rm'`.
4. **Recovery actually progressing but slow**: Compare `objects misplaced`
count over time — if decreasing, it's just slow. If truly constant,
investigate items 1-3.
## OSD Utilization Assessment
```bash
# Full tree with utilization
ceph osd df tree
# Compact view
ceph osd df | grep -E 'osd\.|TOTAL'
```
### This Cluster (2026-07-07)
| OSD | Host | Size | Used | %Used | VAR | PGs | Issue |
|-----|------|------|------|-------|-----|-----|-------|
| osd.7 | proxmox7 | 982G | 863G | 88% | 1.94 | 272 | Overloaded, EC trap |
| osd.10 | proxmox6 | 982G | 790G | 80% | 1.78 | 269 | Backfillfull flag |
| osd.3-5 | proxmox3-5 | 233-238G | ~148G | 62-64% | 1.36-1.41 | 185-189 | SSD, balanced |
| osd.0-2 | proxmox2 | 188-233G | 94-129G | 50-55% | 1.10-1.22 | 128-158 | SSD, lighter |
| osd.6 | ubuntu | 2.8T | 913G | 33% | 0.72 | 345 | Underutilized |
| osd.8 | ubuntu | 2.8T | 928G | 33% | 0.72 | 354 | Underutilized |
| osd.1 | proxmox1 | 3.6T | 1.5T | 40% | 0.88 | 538 | Balanced |
**Root imbalance**: osd.7/10 (982G each) are half the size of osd.6/8
(2.8T each) but receive nearly equal PG counts from CRUSH. The small
OSDs fill up faster. Long-term: add more large HDDs or remove osd.7/10
and replace with larger drives.
## EC Profile Immutability & Pool Recreation Cost
### Key Rule
**Erasure-code profiles are immutable after pool creation.** You cannot
change `k` or `m` on an existing pool. To switch from EC4+1 (k=4, m=1)
to EC3+1 (k=3, m=1), you must:
1. Create a new EC profile (`ceph osd erasure-code-profile set ec31-hdd k=3 m=1 ...`)
2. Create new pools (`media_ec2` + `media_meta2`)
3. Copy all RBD images to the new pool (`rbd copy --data-pool media_ec2`)
4. Update PVE storage config + all CT/VM disk references
5. Delete old pools
### When to Change EC Profile vs. Add OSDs
| Approach | Pros | Cons |
|----------|------|------|
| **Add HDD OSD** | No downtime, no data migration, gives CRUSH room to rebalance, preserves EC4+1 efficiency (25% overhead) | Physical disk needed, OSD count grows |
| **Switch to EC3+1** | Needs only 4 OSDs (one fewer), same 1-failure tolerance | 33% overhead (vs 25%), full pool migration, downtime risk |
**Recommendation**: When the problem is "EC4+1 with exactly 5 HDD OSDs
can't tolerate reweighting," **adding an OSD is almost always better**
than changing the EC profile. The reweight trap exists because CRUSH
has no spare OSD to redirect to — adding one OSD solves that directly,
preserves the better storage efficiency of EC4+1, and requires zero
data migration.
### Trade-off: EC3+1 vs EC4+1
- EC3+1 (k=3, m=1): 4 OSDs needed, 33% overhead, tolerates 1 failure
- EC4+1 (k=4, m=1): 5 OSDs needed, 25% overhead, tolerates 1 failure
- EC4+2 (k=4, m=2): 6 OSDs needed, 50% overhead, tolerates 2 failures
- Both EC3+1 and EC4+1 tolerate only 1 OSD failure — the extra OSD in
EC4+1 buys storage efficiency, not more resilience.
## Samba Time Machine Sparsebundle on Ceph EC
After migrating a Time Machine sparsebundle (via `rbd copy` to EC storage),
Time Machine may report "Backup fehlgeschlagen" with `RESULT: 70` in
`com.apple.TimeMachine.Results.plist`. The sparsebundle is accessible
and bands are being written (verified via `smbstatus` — active locks on
band files), but TM aborts.
### Possible Causes
1. **I/O latency from Ceph recovery**`BLUESTORE_SLOW_OP_ALERT` on
HDD OSDs during recovery causes high write latency. TM is sensitive
to I/O timeouts.
2. **Sparsebundle corruption from rbd copy**`rbd copy` copies at
the block level; if the source was being written to during copy,
the sparsebundle may be inconsistent.
3. **EC pool latency** — EC writes involve chunk calculation across
multiple OSDs, inherently slower than replicated pools.
### Diagnosis
```bash
# Check TM result code
cat "/mnt/timemachine/backup/<name>.sparsebundle/com.apple.TimeMachine.Results.plist"
# RESULT 70 = sparsebundle error
# Check active SMB connections + locked files
smbstatus
# Check I/O latency
ceph status | grep -E 'slow|io:'
dmesg | grep -i "slow\|error\|timeout"
# Check sparsebundle integrity from Mac:
# hdiutil verify /Volumes/TimeMachine/<name>.sparsebundle
# hdiutil repair /Volumes/TimeMachine/<name>.sparsebundle
```
### Resolution Options
1. **Wait for Ceph recovery to complete** — if I/O latency is the cause,
TM backups should resume once HEALTH_OK and slow ops clear.
2. **Verify + repair sparsebundle** — run `hdiutil verify` / `hdiutil repair`
from the Mac (requires mounting the share).
3. **Fresh start** — delete sparsebundle, begin new TM backup (loses history).
## Concurrent I/O Competition
Running RBD copy + Ceph recovery + Seafile fsck simultaneously on the
same HDD OSDs causes:
- `BLUESTORE_SLOW_OP_ALERT` on 4+ OSDs
- Extreme latency on all operations
- Recovery appears stalled (not actually stalled, just throttled)
### Sequencing Recommendation
1. Complete RBD copies first (or pause them)
2. Let Ceph recovery finish (monitor with `ceph status`)
3. Then run fsck or other I/O-intensive operations
If all three must run concurrently, expect 5-10x slowdown on each.
## Playwright MCP Container Restoration
### Scenario
Container was deleted (`docker rm`), image was pruned, but the named
volume survived. Need to recreate the container.
### Discovery
```bash
# Find surviving volumes
docker volume ls | grep playwright
# → playwright-mcp_npm-cache
# Inspect volume contents for package identity
find /var/lib/docker/volumes/playwright-mcp_npm-cache/_data/ -name "package.json" | head -5
cat /var/lib/docker/volumes/playwright-mcp_npm-cache/_data/_npx/*/node_modules/@playwright/mcp/package.json | grep -E '"name"|"version"'
# → "@playwright/mcp" version "0.0.76"
# Check if base image still exists
docker images | grep playwright
# → mcr.microsoft.com/playwright:latest
```
### Recreation
```bash
docker run -d \
--name playwright-mcp \
-p 8081:8080 \
-v playwright-mcp_npm-cache:/root/.npm \
mcr.microsoft.com/playwright:latest \
npx @playwright/mcp@0.0.76 --port 8080
```
### Verification
```bash
docker logs playwright-mcp | tail -5
# → "Listening on http://localhost:8080"
# → MCP endpoint: http://<host>:8081/mcp
```
### General Pattern
When a container is deleted but its named volume survives:
1. Inspect the volume path for `package.json` or config files
2. Identify the package name and version
3. Check if the base image still exists locally
4. Recreate with `docker run` using the same volume mount and package version
5. The npm cache volume makes `npx` startup instant (no re-download)
## Ghost RBD with Orphaned Data Objects (Header Gone)
### Symptoms
After a failed `rbd rm` or aborted operation, the image enters a ghost
state where:
- `rbd ls <pool>` still lists the image (e.g. `vm-114-disk-0`)
- `rbd info <pool>/vm-114-disk-0``(2) No such file or directory`
- `rbd status <pool>/vm-114-disk-0``No such file or directory`
- `rbd rm <pool>/vm-114-disk-0``image still has watchers` (forever)
- `rbd trash mv``deferred delete error: (2) No such file or directory`
- `rados -p <pool> ls | grep rbd_header.<id>` → empty (header gone)
- `rados -p <pool> ls | grep rbd_data.<prefix>` → thousands of objects!
The image header is already deleted, but the data objects remain and
`rbd ls` shows a stale directory entry. The watcher blocking `rbd rm`
is often a **previous stale `rbd rm` process** (stuck in I/O) or a
blacklisted Ceph client.
### Diagnosis
```bash
# 1. Check for stale rbd rm processes (THE most common watcher)
ps aux | grep 'rbd.*rm' | grep -v grep
# If found: kill <pid>; sleep 35; retry rbd rm
# 2. Check Ceph OSD blacklist for stale clients
ceph osd blacklist ls
# If the blocking client is listed:
ceph osd blacklist rm <client_addr>
# e.g. ceph osd blacklist rm 10.0.20.10:0/3681711461
# 3. Verify data objects are orphaned (header gone)
rados -p <pool> ls | grep rbd_header.<image_id> # empty = header gone
rados -p <pool> ls | grep rbd_data.<prefix> | wc -l # count orphaned objects
```
### Fix: Purge Orphaned Data Objects
When the header is gone and `rbd rm` keeps failing, the data objects
are orphaned and can be purged directly:
```bash
# Find the data object prefix from rbd_data object names
rados -p tm_disks ls | grep rbd_data.378d64d5a67ea8 | head -1
# → rbd_data.378d64d5a67ea8.00000000000121a0
# Purge all orphaned objects (parallel for speed)
rados -p tm_disks ls | grep rbd_data.378d64d5a67ea8 | \
xargs -P8 -I{} rados -p tm_disks rm {}
# Verify cleanup
rados -p tm_disks ls | grep rbd_data.378d64d5a67ea8 | wc -l
# → 0
# ⚠️ This can take a long time — 65k objects × 4 MiB ≈ 341 GiB
# Run in background with notify_on_complete=true
```
⚠️ **Verify the prefix is correct before purging!** Use `rbd info` on
a DIFFERENT image in the same pool to confirm the prefix format
(`rbd_data.<image_id_hash>.<object_index>`). Purging the wrong prefix
destroys live data.
⚠️ **Requires user confirmation** — this is a destructive operation
that bypasses RBD's safety checks. Always ask the user before purging.
### Prevention
- Always `rbd unmap` on ALL nodes after RBD operations
- Kill stale `rbd rm` processes before retrying
- Check `ceph osd blacklist ls` for stale clients
- Don't run `rbd rm` in background — if it hangs, the stale process
becomes the watcher blocking the next attempt
## Competing I/O Forces on Overloaded OSDs
When reweighting an overloaded OSD (e.g. osd.7 at 88%) while concurrent
writes target the same OSD (e.g. RBD copy to an EC pool that includes
that OSD), the net effect can be **increasing** utilization despite
the reweight:
- **Reweight** moves existing data AWAY from the OSD → pressure decreases
- **Concurrent writes** (RBD copy, fsck reads) add NEW data to the OSD → pressure increases
- If write rate > recovery rate, utilization goes UP
Observed: osd.7 dropped from 88% → 80.77% (reweight working), then rose
back to 85% (RBD copy writing new EC chunks to media_ec, which includes
osd.7 as one of 5 HDD OSDs).
### Lesson
Don't reweight an OSD to relieve pressure while simultaneously running
I/O-intensive operations that write to pools containing that OSD. Either:
1. Complete the write operation first, THEN reweight
2. Or reweight first, wait for recovery, THEN start writes
## Session Log — 2026-07-12
- osd.10 now at **96.72%** (up from 80% on July 7, 95% on July 4).
hdd_disk pool at **99.09%**. 7 pools backfillfull.
- Attempted `ceph osd crush reweight osd.10 0.85` then `0.70`
replicated pool PGs started backfilling, but EC pool PGs remained
stuck (same EC4+1 trap as July 7).
- Raised `backfillfull_ratio` to 0.97 — did NOT clear the sticky flag
because osd.10 is still above 96.72% (below 0.97 threshold but the
flag was set earlier and is sticky).
- **User reminded**: "reweight hat beim letzten mal wegen erasure
coding nicht funktioniert" — this is the SAME issue documented
above. The agent should have checked EC profile + OSD count BEFORE
attempting reweight, not after.
- **Lesson encoded**: Pre-reweight checklist is now mandatory:
1. `ceph osd erasure-code-profile ls` — are there EC pools?
2. If yes, `ceph osd erasure-code-profile get <profile>` — get k+m
3. `ceph osd tree | grep <device-class> | wc -l` — count OSDs
4. If OSD count == k+m: DO NOT REWEIGHT. Reset to 1.0 if already
reweighted.
- Confirmed orphans on hdd_disk (~170 GiB reclaimable): base-101,
base-102, vm-143, vm-902, vm-201, vm-202, csi-vol-e596, test-dummy,
vm-104-disk-0@pre-paperless-migration snapshot.
- vm-201/vm-202: `rbd info` returns empty (ghost images, header gone).
Same pattern as the ghost RBD section above — may need `rados rm`
purge of orphaned data objects.
- CT114 (smb-tm-noris) successfully started on EC storage (`media` pool).
RBD copy from 2026-07-06 completed (340/341 GiB). CT config updated
from `tm_disks:vm-114-disk-0``media:vm-114-disk-0`.
- Old `tm_disks/vm-114-disk-0` (341 GiB) deletion: first attempt failed
with "image still has watchers" — a stale `rbd rm` process (PID 2916941)
was the watcher. Killed it, waited 35s, retry succeeded partially but
image entered ghost state (header gone, 65k orphaned data objects,
341 GiB). OSD blacklist also had a stale client. User confirmed
purge of orphaned objects via `rados rm`.
- Seafile fsck completed: 61/55 repos (6 overcount from restart).
- osd.7 reweight reverted from 0.65 → 1.0 (EC4+1 trap).
- osd.10 reweight reverted from 0.5 → 1.0 (same trap).
- PG merge completed: cephfs_data 128→32, cephfs_metadata 128→32.
- Recovery accelerated: osd_max_backfills=5, osd_recovery_op_priority=3.
- backfillfull flag on osd.10 persists despite ratio increase (sticky).
12 PGs remain `backfill_toofull` — waiting for natural clearance.
- CT143 (firecrawl) creation attempted on `media` storage — blocked by
CFS lock timeout + mkfs blocklist (same issue as 2026-07-06).
User instructed to wait for Ceph to stabilize, then retry.
- Playwright MCP container restored from surviving npm cache volume.
- Recovery stagnated at 11.5% misplaced for 7+ hours. Root cause was
EC4+1 reweight trap (2147483647 in up-sets), NOT I/O contention.
After reverting reweights, recovery resumed at 67 MiB/s.
- Recovery progressed to 0.74% misplaced (from 11.5%) after reweight
revert + recovery acceleration (osd_max_backfills=5, op_priority=3).
5 PGs remained `backfill_toofull` on osd.10 (sticky backfillfull flag).
- PG merge for cephfs_data + cephfs_metadata completed (128→32, forced
via `ceph osd pool set <pool> pg_num 32` — worked even with 15 remapped
PGs, reducing remapped count from 18→15).
- User asked about switching EC4+1→EC3+1 to relieve pressure. Answer:
EC profiles are immutable, requires full pool recreation + migration.
User decided to add a new HDD OSD instead — simpler, preserves EC4+1
efficiency, no data migration. See "EC Profile Immutability" section.
- CT143 (firecrawl) creation still blocked — `pvesm alloc` succeeds
(creates RBD image) but `pct create` fails because it calls `mkfs`
internally, which is agent-hardline-blocked. User must run `mkfs.ext4`
manually on the mapped RBD device, then `pct create` succeeds.
- TM backup failure on CT114 (smb-tm-noris): Mac connects, writes bands,
but TM reports RESULT 70 (sparsebundle error). Likely I/O latency from
ongoing Ceph recovery. User chose to wait for Ceph to stabilize.
- Playwright MCP container restored on CT111 from surviving npm-cache
volume (`@playwright/mcp` v0.0.76, `mcr.microsoft.com/playwright:latest`).
- Orphaned RBD cleanup for old `tm_disks/vm-114-disk-0`: ghost image
(header gone, 65k orphaned data objects, 341 GiB). Stale `rbd rm`
process was the watcher. Purge via `rados rm` pending user confirmation.
- PBS offsite backup job expanded: `vmid 106``vmid 106,108,104,99999`.
Initial full backups run manually on correct nodes (CT108 on proxmox6,
CT99999 on proxmox7, CT104 on n5pro). All 3 completed successfully.
- Frigate (CT120) excluded from all backup jobs per user decision.
- Seafile backup strategy analyzed: 6 options compared. Seafile Pro 13.0.19
supports S3 as primary storage backend (not backup target). User
considering S3 backend to move 558 GB blocks out of Ceph. See
`seafile-api` skill → `references/seafile-immich-backup-strategy.md`.
- CT111 (Seafile) backup on noris_s3 showed 0 MB (empty/failed) on 2026-07-07.
Previous valid backup: 558 GB on 2026-07-06.