Files
hermes-skills/devops/proxmox-ve-administration/references/ceph-recovery-ec-reweight-2026-07.md
T
Debian 01bd921ced feat: add home-assistant-dashboard-conventions skill + update multiple skills
- New: smart-home/home-assistant-dashboard-conventions (Mushroom cards, view tabs, no Bubble Cards)
- Updated: rke2, ceph, galera, proxmox, brainstorming, compound-learning, 1password-cli, smart-home-automation skills
- New references: ceph-cluster-administration, docker-volume-forensics, ceph-crush-weight, ceph-ec-mixed-size
2026-07-14 18:35:16 +00:00

633 lines
25 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Ceph Recovery, EC Reweight, PG Merge & Backfillfull
## EC4+1 Reweight Trap — CRUSH Returns 2147483647
### Problem
With exactly 5 HDD OSDs and EC profile `k=4 m=1` (needs exactly 5 OSDs),
rewighting any OSD (e.g. `ceph osd reweight osd.7 0.65`) causes CRUSH
to return `2147483647` ("no OSD available") in the up-set for affected
PGs. These PGs become `active+clean+remapped` and **cannot recover**
there is no alternative OSD to place them on.
```
8.20 active+clean+remapped [6,1,8,10,2147483647] 6
8.78 active+clean+remapped [1,6,8,2147483647,10] 1
```
### Diagnosis
```bash
# Check EC profile
ceph osd erasure-code-profile get ec41-hdd
# k=4 m=1 → needs 5 OSDs
# Count HDD OSDs
ceph osd tree | grep hdd | wc -l
# 5 → exactly the minimum. Reweighting breaks CRUSH.
# Check for 2147483647 in up-sets
ceph pg dump pgs_brief 2>/dev/null | grep 2147483647 | wc -l
```
### Fix
**Reset reweight to 1.0 for all EC-participating OSDs.** With only 5
HDD OSDs and EC4+1, there is no spare capacity for reweighting. Every
OSD must participate at full weight.
```bash
ceph osd reweight osd.7 1.0
ceph osd reweight osd.10 1.0
```
After resetting, recovery resumes (67 MiB/s, 16 obj/s observed).
### General Rule
- **Replicated pools** (size=3): reweighting is safe — CRUSH redirects
to other OSDs in the failure domain.
- **Erasure-coded pools**: `crush reweight` (permanent CRUSH weight change)
is only safe when there are MORE OSDs than `k+m`. With exactly `k+m`
OSDs, `crush reweight` traps PGs (2147483647 in up-set).
- **⚠️ `osd reweight` (temporary) CAN drain EC OSDs** (2026-07-12):
Unlike `crush reweight`, `ceph osd reweight N 0.5` sets a temporary
runtime override that adjusts PG placement probability WITHOUT
changing CRUSH weights. This means CRUSH still considers all OSDs
available — no 2147483647 trap. Tested on osd.10 (EC4+1, exactly 5
HDD OSDs): reweight to 0.5 drained from 93%→91%, backfill started
at 30 MiB/s, 5 of 6 `backfill_toofull` PGs unblocked. **Prerequisite:**
`crush reweight` must be at 1.0 (default). If `crush reweight` was
previously changed, revert it to 1.0 FIRST, then use `osd reweight`.
- **Long-term fix**: add more HDD OSDs (e.g. on the ubuntu host with
2×2.8 TB spare) to give CRUSH room to redistribute.
## Backfillfull Sticky Flag
### Problem
osd.10 at 80% utilization has a `backfillfull` flag. Even after raising
the ratio:
```bash
ceph osd set-backfillfull-ratio 0.95
ceph osd set-nearfull-ratio 0.95
```
The flag persists and blocks backfill on 12+ PGs:
```
HEALTH_WARN ... 1 backfillfull osd(s); Low space hindering backfill:
12 pgs backfill_toofull; 6 pool(s) backfillfull
```
### Attempted Fixes That Don't Work
```bash
# Does not exist in this Ceph version (Reef/Quincy):
ceph osd clear-backfillfull osd.10 # → EINVAL
# Cannot be modified at runtime:
ceph tell osd.7 injectargs --mon_osd_backfillfull_ratio 0.95 # → EPERM
# Forces backfill on PGs that need it, but doesn't clear toofull:
ceph osd pool force-backfill hdd_disk # instructs PGs to force-backfill
# → PGs still show backfill_toofull
```
### What Happened
The `backfillfull` flag was set when osd.10 crossed the OLD threshold
(default 0.85). Raising the ratio afterward does NOT retroactively
clear the flag — it only prevents future flag-setting. The OSD remains
flagged until its utilization drops below the new threshold naturally
(via recovery moving data away).
### Mitigation
- **`osd reweight` (temporary) WORKS for EC4+1 drain** (2026-07-12):
`ceph osd reweight 10 0.5` successfully drained osd.10 from 93.14%→90.90%
and unblocked 5 of 6 `backfill_toofull` PGs. This is a runtime
override, NOT a CRUSH weight change — CRUSH still sees all OSDs,
preventing the 2147483647 trap. Recovery proceeded at 30 MiB/s.
Steps: (1) ensure `crush reweight` is 1.0, (2) `ceph osd reweight
10 0.5`, (3) `ceph osd set-backfillfull-ratio 0.95` to unblock
remaining PGs, (4) wait for recovery, (5) `ceph osd reweight 10 1.0`
to restore.
- **Aggressive drain: reweight 0.3** (2026-07-12 session 2):
When 0.5 wasn't enough (osd.10 still at 89%, 1 PG stubbornly
`backfill_toofull`), dropping to `ceph osd reweight 10 0.3` increased
misplaced objects from 12%→16% (more PGs queued for relocation) and
accelerated the drain. Combined with `set-backfillfull-ratio 0.97`.
**Reweight progression: 0.5 → 0.3 if 0.5 doesn't drain fast enough.**
Remember to restore to 1.0 after recovery completes.
### ⚠️ Correct Command Syntax for backfillfull_ratio (2026-07-12)
Multiple command forms were tried — only ONE works. Save time by using
the correct one first:
```bash
# ✅ CORRECT — hyphenated, no underscore:
ceph osd set-backfillfull-ratio 0.97
# ❌ WRONG — underscore variant (EINVAL):
ceph osd set_backfill_full_ratio 0.97
# ❌ WRONG — config set (rejected as "special"):
ceph config set mon mon_osd_backfillfull_ratio 0.97
# → "mon_osd_backfillfull_ratio is special and cannot be stored by the mon"
# ❌ WRONG — injectargs with wrong key name:
ceph tell osd.10 injectargs "--osd_backfill_full_ratio 0.97"
# → "failed to parse arguments"
# ❌ WRONG — osd set with wrong key:
ceph osd set backfill_full_ratio 0.97
# → "Invalid command: backfill_full_ratio not in full|pause|..."
```
Similarly for full_ratio and nearfull_ratio:
```bash
ceph osd set-full-ratio 0.97 # ✅ hyphenated
ceph osd set-nearfull-ratio 0.95 # ✅ hyphenated
```
- **Patience**: Once the reweight takes effect and recovery proceeds,
data redistributes and osd.10 utilization drops, eventually clearing
the backfillfull flag.
- **PG merge**: Reducing PG count (128→32) eliminates some remapped PGs
entirely, reducing the number of PGs that need to backfill to the
stuck OSD.
- **Do NOT use `crush reweight`** to reduce osd.10's load — with EC4+1
and only 5 HDD OSDs, permanent CRUSH reweighting makes things worse
(see EC4+1 Reweight Trap above). Use `osd reweight` (temporary)
instead.
## PG Merge (128 → 32)
### Procedure
```bash
# 1. Disable autoscaler for the pool
ceph osd pool set cephfs_data pg_autoscale_mode off
ceph osd pool set cephfs_metadata pg_autoscale_mode off
# 2. Set pgp_num first (must be ≤ pg_num)
ceph osd pool set cephfs_data pgp_num 32
ceph osd pool set cephfs_metadata pgp_num 32
# 3. Set pg_num_target (async merge preparation)
ceph osd pool set cephfs_data pg_num_target 32
ceph osd pool set cephfs_metadata pg_num_target 32
# 4. Force the merge — works when all PGs are clean
ceph osd pool set cephfs_data pg_num 32
ceph osd pool set cephfs_metadata pg_num 32
# 5. Verify
ceph osd pool get cephfs_data pg_num # → 32
ceph osd pool get cephfs_metadata pg_num # → 32
```
### When It Works
- All PGs in the pool must be `active+clean` for the merge to execute.
- If PGs are `remapped` or `backfilling`, the merge waits.
- With `pg_autoscale_mode off`, the merge is manual and deterministic.
### Impact
Reducing `cephfs_data` (0 bytes stored) and `cephfs_metadata` (85 KiB
stored) from 128→32 PGs eliminates 384 PG instances across OSDs, helping
with the "too many PGs per OSD" health warning (262 > 250).
## Recovery Speed Tuning
```bash
# Increase concurrent backfills per OSD (default 1)
ceph config set osd osd_max_backfills 5
# Increase recovery op priority (default 1, lower = higher priority)
ceph config set osd osd_recovery_op_priority 3
# Ensure no recovery sleep (default 0)
ceph config get osd osd_recovery_sleep # → 0.000000 (already optimal)
```
### Monitoring Recovery Progress
```bash
# Overall
ceph status | grep -E 'recovery|pgs:'
# Count remapped/backfill PGs
ceph pg dump pgs_brief 2>/dev/null | tail -n +2 | awk '{print $2}' | grep -c remapped
ceph pg dump pgs_brief 2>/dev/null | tail -n +2 | awk '{print $2}' | grep -c backfill_toofull
# Recovery rate
ceph status | grep -A2 'recovery:'
# → recovery: 67 MiB/s, 16 objects/s
```
### Diagnosing Stalled Recovery (2026-07-07)
When recovery appears stalled (misplaced objects constant for hours),
check these in order BEFORE tuning recovery speed:
1. **EC reweight trap**: `ceph pg dump pgs_brief | grep 2147483647`
— If PGs contain `2147483647` ("no OSD"), CRUSH cannot place them.
Reset all EC-participating OSDs to reweight 1.0. See section above.
2. **Backfillfull sticky flag**: `ceph health detail | grep backfillfull`
— An OSD flagged `backfillfull` blocks all backfill to it, even after
raising the ratio. The flag is set when crossing the OLD threshold and
does NOT clear retroactively. See "Backfillfull Sticky Flag" section.
3. **Stale `rbd rm` process**: A previous `rbd rm` stuck in I/O holds a
watcher that blocks new operations. Check `ps aux | grep 'rbd.*rm'`.
4. **Recovery actually progressing but slow**: Compare `objects misplaced`
count over time — if decreasing, it's just slow. If truly constant,
investigate items 1-3.
## OSD Utilization Assessment
```bash
# Full tree with utilization
ceph osd df tree
# Compact view
ceph osd df | grep -E 'osd\.|TOTAL'
```
### This Cluster (2026-07-07)
| OSD | Host | Size | Used | %Used | VAR | PGs | Issue |
|-----|------|------|------|-------|-----|-----|-------|
| osd.7 | proxmox7 | 982G | 863G | 88% | 1.94 | 272 | Overloaded, EC trap |
| osd.10 | proxmox6 | 982G | 790G | 80% | 1.78 | 269 | Backfillfull flag |
| osd.3-5 | proxmox3-5 | 233-238G | ~148G | 62-64% | 1.36-1.41 | 185-189 | SSD, balanced |
| osd.0-2 | proxmox2 | 188-233G | 94-129G | 50-55% | 1.10-1.22 | 128-158 | SSD, lighter |
| osd.6 | ubuntu | 2.8T | 913G | 33% | 0.72 | 345 | Underutilized |
| osd.8 | ubuntu | 2.8T | 928G | 33% | 0.72 | 354 | Underutilized |
| osd.1 | proxmox1 | 3.6T | 1.5T | 40% | 0.88 | 538 | Balanced |
**Root imbalance**: osd.7/10 (982G each) are half the size of osd.6/8
(2.8T each) but receive nearly equal PG counts from CRUSH. The small
OSDs fill up faster. Long-term: add more large HDDs or remove osd.7/10
and replace with larger drives.
## EC Profile Immutability & Pool Recreation Cost
### Key Rule
**Erasure-code profiles are immutable after pool creation.** You cannot
change `k` or `m` on an existing pool. To switch from EC4+1 (k=4, m=1)
to EC3+1 (k=3, m=1), you must:
1. Create a new EC profile (`ceph osd erasure-code-profile set ec31-hdd k=3 m=1 ...`)
2. Create new pools (`media_ec2` + `media_meta2`)
3. Copy all RBD images to the new pool (`rbd copy --data-pool media_ec2`)
4. Update PVE storage config + all CT/VM disk references
5. Delete old pools
### When to Change EC Profile vs. Add OSDs
| Approach | Pros | Cons |
|----------|------|------|
| **Add HDD OSD** | No downtime, no data migration, gives CRUSH room to rebalance, preserves EC4+1 efficiency (25% overhead) | Physical disk needed, OSD count grows |
| **Switch to EC3+1** | Needs only 4 OSDs (one fewer), same 1-failure tolerance | 33% overhead (vs 25%), full pool migration, downtime risk |
**Recommendation**: When the problem is "EC4+1 with exactly 5 HDD OSDs
can't tolerate reweighting," **adding an OSD is almost always better**
than changing the EC profile. The reweight trap exists because CRUSH
has no spare OSD to redirect to — adding one OSD solves that directly,
preserves the better storage efficiency of EC4+1, and requires zero
data migration.
### Trade-off: EC3+1 vs EC4+1
- EC3+1 (k=3, m=1): 4 OSDs needed, 33% overhead, tolerates 1 failure
- EC4+1 (k=4, m=1): 5 OSDs needed, 25% overhead, tolerates 1 failure
- EC4+2 (k=4, m=2): 6 OSDs needed, 50% overhead, tolerates 2 failures
- Both EC3+1 and EC4+1 tolerate only 1 OSD failure — the extra OSD in
EC4+1 buys storage efficiency, not more resilience.
## Samba Time Machine Sparsebundle on Ceph EC
After migrating a Time Machine sparsebundle (via `rbd copy` to EC storage),
Time Machine may report "Backup fehlgeschlagen" with `RESULT: 70` in
`com.apple.TimeMachine.Results.plist`. The sparsebundle is accessible
and bands are being written (verified via `smbstatus` — active locks on
band files), but TM aborts.
### Possible Causes
1. **I/O latency from Ceph recovery**`BLUESTORE_SLOW_OP_ALERT` on
HDD OSDs during recovery causes high write latency. TM is sensitive
to I/O timeouts.
2. **Sparsebundle corruption from rbd copy**`rbd copy` copies at
the block level; if the source was being written to during copy,
the sparsebundle may be inconsistent.
3. **EC pool latency** — EC writes involve chunk calculation across
multiple OSDs, inherently slower than replicated pools.
### Diagnosis
```bash
# Check TM result code
cat "/mnt/timemachine/backup/<name>.sparsebundle/com.apple.TimeMachine.Results.plist"
# RESULT 70 = sparsebundle error
# Check active SMB connections + locked files
smbstatus
# Check I/O latency
ceph status | grep -E 'slow|io:'
dmesg | grep -i "slow\|error\|timeout"
# Check sparsebundle integrity from Mac:
# hdiutil verify /Volumes/TimeMachine/<name>.sparsebundle
# hdiutil repair /Volumes/TimeMachine/<name>.sparsebundle
```
### Resolution Options
1. **Wait for Ceph recovery to complete** — if I/O latency is the cause,
TM backups should resume once HEALTH_OK and slow ops clear.
2. **Verify + repair sparsebundle** — run `hdiutil verify` / `hdiutil repair`
from the Mac (requires mounting the share).
3. **Fresh start** — delete sparsebundle, begin new TM backup (loses history).
## Concurrent I/O Competition
Running RBD copy + Ceph recovery + Seafile fsck simultaneously on the
same HDD OSDs causes:
- `BLUESTORE_SLOW_OP_ALERT` on 4+ OSDs
- Extreme latency on all operations
- Recovery appears stalled (not actually stalled, just throttled)
### Sequencing Recommendation
1. Complete RBD copies first (or pause them)
2. Let Ceph recovery finish (monitor with `ceph status`)
3. Then run fsck or other I/O-intensive operations
If all three must run concurrently, expect 5-10x slowdown on each.
## Playwright MCP Container Restoration
### Scenario
Container was deleted (`docker rm`), image was pruned, but the named
volume survived. Need to recreate the container.
### Discovery
```bash
# Find surviving volumes
docker volume ls | grep playwright
# → playwright-mcp_npm-cache
# Inspect volume contents for package identity
find /var/lib/docker/volumes/playwright-mcp_npm-cache/_data/ -name "package.json" | head -5
cat /var/lib/docker/volumes/playwright-mcp_npm-cache/_data/_npx/*/node_modules/@playwright/mcp/package.json | grep -E '"name"|"version"'
# → "@playwright/mcp" version "0.0.76"
# Check if base image still exists
docker images | grep playwright
# → mcr.microsoft.com/playwright:latest
```
### Recreation
```bash
docker run -d \
--name playwright-mcp \
-p 8081:8080 \
-v playwright-mcp_npm-cache:/root/.npm \
mcr.microsoft.com/playwright:latest \
npx @playwright/mcp@0.0.76 --port 8080
```
### Verification
```bash
docker logs playwright-mcp | tail -5
# → "Listening on http://localhost:8080"
# → MCP endpoint: http://<host>:8081/mcp
```
### General Pattern
When a container is deleted but its named volume survives:
1. Inspect the volume path for `package.json` or config files
2. Identify the package name and version
3. Check if the base image still exists locally
4. Recreate with `docker run` using the same volume mount and package version
5. The npm cache volume makes `npx` startup instant (no re-download)
## Ghost RBD with Orphaned Data Objects (Header Gone)
### Symptoms
After a failed `rbd rm` or aborted operation, the image enters a ghost
state where:
- `rbd ls <pool>` still lists the image (e.g. `vm-114-disk-0`)
- `rbd info <pool>/vm-114-disk-0``(2) No such file or directory`
- `rbd status <pool>/vm-114-disk-0``No such file or directory`
- `rbd rm <pool>/vm-114-disk-0``image still has watchers` (forever)
- `rbd trash mv``deferred delete error: (2) No such file or directory`
- `rados -p <pool> ls | grep rbd_header.<id>` → empty (header gone)
- `rados -p <pool> ls | grep rbd_data.<prefix>` → thousands of objects!
The image header is already deleted, but the data objects remain and
`rbd ls` shows a stale directory entry. The watcher blocking `rbd rm`
is often a **previous stale `rbd rm` process** (stuck in I/O) or a
blacklisted Ceph client.
### Diagnosis
```bash
# 1. Check for stale rbd rm processes (THE most common watcher)
ps aux | grep 'rbd.*rm' | grep -v grep
# If found: kill <pid>; sleep 35; retry rbd rm
# 2. Check Ceph OSD blacklist for stale clients
ceph osd blacklist ls
# If the blocking client is listed:
ceph osd blacklist rm <client_addr>
# e.g. ceph osd blacklist rm 10.0.20.10:0/3681711461
# 3. Verify data objects are orphaned (header gone)
rados -p <pool> ls | grep rbd_header.<image_id> # empty = header gone
rados -p <pool> ls | grep rbd_data.<prefix> | wc -l # count orphaned objects
```
### Fix: Purge Orphaned Data Objects
When the header is gone and `rbd rm` keeps failing, the data objects
are orphaned and can be purged directly:
```bash
# Find the data object prefix from rbd_data object names
rados -p tm_disks ls | grep rbd_data.378d64d5a67ea8 | head -1
# → rbd_data.378d64d5a67ea8.00000000000121a0
# Purge all orphaned objects (parallel for speed)
rados -p tm_disks ls | grep rbd_data.378d64d5a67ea8 | \
xargs -P8 -I{} rados -p tm_disks rm {}
# Verify cleanup
rados -p tm_disks ls | grep rbd_data.378d64d5a67ea8 | wc -l
# → 0
# ⚠️ This can take a long time — 65k objects × 4 MiB ≈ 341 GiB
# Run in background with notify_on_complete=true
```
⚠️ **Verify the prefix is correct before purging!** Use `rbd info` on
a DIFFERENT image in the same pool to confirm the prefix format
(`rbd_data.<image_id_hash>.<object_index>`). Purging the wrong prefix
destroys live data.
⚠️ **Requires user confirmation** — this is a destructive operation
that bypasses RBD's safety checks. Always ask the user before purging.
### Prevention
- Always `rbd unmap` on ALL nodes after RBD operations
- Kill stale `rbd rm` processes before retrying
- Check `ceph osd blacklist ls` for stale clients
- Don't run `rbd rm` in background — if it hangs, the stale process
becomes the watcher blocking the next attempt
## Competing I/O Forces on Overloaded OSDs
When reweighting an overloaded OSD (e.g. osd.7 at 88%) while concurrent
writes target the same OSD (e.g. RBD copy to an EC pool that includes
that OSD), the net effect can be **increasing** utilization despite
the reweight:
- **Reweight** moves existing data AWAY from the OSD → pressure decreases
- **Concurrent writes** (RBD copy, fsck reads) add NEW data to the OSD → pressure increases
- If write rate > recovery rate, utilization goes UP
Observed: osd.7 dropped from 88% → 80.77% (reweight working), then rose
back to 85% (RBD copy writing new EC chunks to media_ec, which includes
osd.7 as one of 5 HDD OSDs).
### Lesson
Don't reweight an OSD to relieve pressure while simultaneously running
I/O-intensive operations that write to pools containing that OSD. Either:
1. Complete the write operation first, THEN reweight
2. Or reweight first, wait for recovery, THEN start writes
## Session Log — 2026-07-12
- osd.10 now at **96.72%** (up from 80% on July 7, 95% on July 4).
hdd_disk pool at **99.09%**. 7 pools backfillfull.
- Attempted `ceph osd crush reweight osd.10 0.85` then `0.70`
replicated pool PGs started backfilling, but EC pool PGs remained
stuck (same EC4+1 trap as July 7).
- Raised `backfillfull_ratio` to 0.97 — did NOT clear the sticky flag
because osd.10 is still above 96.72% (below 0.97 threshold but the
flag was set earlier and is sticky).
- **User reminded**: "reweight hat beim letzten mal wegen erasure
coding nicht funktioniert" — this is the SAME issue documented
above. The agent should have checked EC profile + OSD count BEFORE
attempting reweight, not after.
- **Lesson encoded**: Pre-reweight checklist is now mandatory:
1. `ceph osd erasure-code-profile ls` — are there EC pools?
2. If yes, `ceph osd erasure-code-profile get <profile>` — get k+m
3. `ceph osd tree | grep <device-class> | wc -l` — count OSDs
4. If OSD count == k+m: DO NOT REWEIGHT. Reset to 1.0 if already
reweighted.
- Confirmed orphans on hdd_disk (~170 GiB reclaimable): base-101,
base-102, vm-143, vm-902, vm-201, vm-202, csi-vol-e596, test-dummy,
vm-104-disk-0@pre-paperless-migration snapshot.
- vm-201/vm-202: `rbd info` returns empty (ghost images, header gone).
Same pattern as the ghost RBD section above — may need `rados rm`
purge of orphaned data objects.
- CT114 (smb-tm-noris) successfully started on EC storage (`media` pool).
RBD copy from 2026-07-06 completed (340/341 GiB). CT config updated
from `tm_disks:vm-114-disk-0``media:vm-114-disk-0`.
- Old `tm_disks/vm-114-disk-0` (341 GiB) deletion: first attempt failed
with "image still has watchers" — a stale `rbd rm` process (PID 2916941)
was the watcher. Killed it, waited 35s, retry succeeded partially but
image entered ghost state (header gone, 65k orphaned data objects,
341 GiB). OSD blacklist also had a stale client. User confirmed
purge of orphaned objects via `rados rm`.
- Seafile fsck completed: 61/55 repos (6 overcount from restart).
- osd.7 reweight reverted from 0.65 → 1.0 (EC4+1 trap).
- osd.10 reweight reverted from 0.5 → 1.0 (same trap).
- PG merge completed: cephfs_data 128→32, cephfs_metadata 128→32.
- Recovery accelerated: osd_max_backfills=5, osd_recovery_op_priority=3.
- backfillfull flag on osd.10 persists despite ratio increase (sticky).
12 PGs remain `backfill_toofull` — waiting for natural clearance.
- CT143 (firecrawl) creation attempted on `media` storage — blocked by
CFS lock timeout + mkfs blocklist (same issue as 2026-07-06).
User instructed to wait for Ceph to stabilize, then retry.
- Playwright MCP container restored from surviving npm cache volume.
- Recovery stagnated at 11.5% misplaced for 7+ hours. Root cause was
EC4+1 reweight trap (2147483647 in up-sets), NOT I/O contention.
After reverting reweights, recovery resumed at 67 MiB/s.
- Recovery progressed to 0.74% misplaced (from 11.5%) after reweight
revert + recovery acceleration (osd_max_backfills=5, op_priority=3).
5 PGs remained `backfill_toofull` on osd.10 (sticky backfillfull flag).
- PG merge for cephfs_data + cephfs_metadata completed (128→32, forced
via `ceph osd pool set <pool> pg_num 32` — worked even with 15 remapped
PGs, reducing remapped count from 18→15).
- User asked about switching EC4+1→EC3+1 to relieve pressure. Answer:
EC profiles are immutable, requires full pool recreation + migration.
User decided to add a new HDD OSD instead — simpler, preserves EC4+1
efficiency, no data migration. See "EC Profile Immutability" section.
- CT143 (firecrawl) creation still blocked — `pvesm alloc` succeeds
(creates RBD image) but `pct create` fails because it calls `mkfs`
internally, which is agent-hardline-blocked. User must run `mkfs.ext4`
manually on the mapped RBD device, then `pct create` succeeds.
- TM backup failure on CT114 (smb-tm-noris): Mac connects, writes bands,
but TM reports RESULT 70 (sparsebundle error). Likely I/O latency from
ongoing Ceph recovery. User chose to wait for Ceph to stabilize.
- Playwright MCP container restored on CT111 from surviving npm-cache
volume (`@playwright/mcp` v0.0.76, `mcr.microsoft.com/playwright:latest`).
- Orphaned RBD cleanup for old `tm_disks/vm-114-disk-0`: ghost image
(header gone, 65k orphaned data objects, 341 GiB). Stale `rbd rm`
process was the watcher. Purge via `rados rm` pending user confirmation.
- PBS offsite backup job expanded: `vmid 106``vmid 106,108,104,99999`.
Initial full backups run manually on correct nodes (CT108 on proxmox6,
CT99999 on proxmox7, CT104 on n5pro). All 3 completed successfully.
- Frigate (CT120) excluded from all backup jobs per user decision.
- Seafile backup strategy analyzed: 6 options compared. Seafile Pro 13.0.19
supports S3 as primary storage backend (not backup target). User
considering S3 backend to move 558 GB blocks out of Ceph. See
`seafile-api` skill → `references/seafile-immich-backup-strategy.md`.
- CT111 (Seafile) backup on noris_s3 showed 0 MB (empty/failed) on 2026-07-07.
Previous valid backup: 558 GB on 2026-07-06.
## Session Log — 2026-07-12 (Continued)
- **osd.10 drain via `osd reweight` SUCCESSFUL**: `ceph osd reweight 10 0.5`
drained osd.10 from 93.14%→90.90%. Combined with `set-backfillfull-ratio 0.95`,
unblocked 5 of 6 `backfill_toofull` PGs. Recovery at 30 MiB/s, 13% misplaced.
This CONTRADICTS the earlier advice "Do NOT reweight osd.10" — that was
correct for `crush reweight` (permanent, triggers 2147483647 trap) but
WRONG for `osd reweight` (temporary runtime override, safe with EC4+1).
Reference patched accordingly.
- **Orphaned RBD cleanup completed**: vm-143-disk-0 (50G), base-101-disk-0
(20G), base-102-disk-0 (20G), vm-902-disk-0 (40G), vm-201-disk-0,
vm-202-disk-0, test-dummy all deleted. base-101/102 required
`rbd snap unprotect @__base__` before `rbd snap purge` + `rbd rm`.
csi-vol-e596 has stale watcher (10.0.30.63, no K8s PVC found) —
pending natural watcher timeout. ~161 GiB freed.
- **K8s worker anti-colocation migration completed**: VM 132→proxmox1,
VM 131→proxmox4, VM 128→n5pro via `qm migrate --online` (HA-managed,
async). Required expanding `k8s-prefer-n5pro` node-affinity rule to
include all 8 nodes first. All 6 K8s VMs now on 6 separate nodes.
- **hdd_disk→vm_disks migration feasibility**: vm_disks has only 86 GiB
free (75% full). Small CTs (vm-127 4G, vm-121 4G, base-9000 3G,
vm-9001 10G = ~21G total) are candidates. Large K8s worker disks
(80G×3) stay on hdd_disk (correct HDD-tier placement).