--- name: ceph-cluster-administration description: "Class-level skill for Ceph cluster administration on Proxmox VE. Covers OSD lifecycle (create/destroy/WAL-DB), CRUSH topology (device classes, roots, rules, weights), EC pools (profiles, mixed disk sizes, migration), recovery & rebalancing (tuning, ratios, monitoring), RBD management (orphan cleanup, migration, inspection), pool management (PG count, full recovery, autoscale), and health diagnosis." version: 1.0.0 author: Hermes Agent tags: [ceph, proxmox, bluestore, crush, erasure-coding, rbd, recovery, osd, storage] metadata: hermes: tags: [ceph, proxmox, bluestore, crush, erasure-coding, rbd, recovery, osd, storage] related_skills: [proxmox-ve-administration, rke2-cluster-administration] --- # Ceph Cluster Administration ## Scope **This skill covers the Homelab (self-hosted) Ceph cluster.** Not noris infrastructure. At noris only remote PBS + S3 Storage exist — neither involves Ceph. Do not confuse the two environments. ## Environment | Component | Detail | |-----------|--------| | Environment | **Homelab (self-hosted)** — NOT noris | | PVE Master | 10.0.20.10 (SSH: `root@10.0.20.10`, key `~/.ssh/id_ed25519_proxmox`) | | PVE Nodes (SSH) | All PVE hosts reachable via `10.0.20.x` (storage VLAN), NOT `10.0.30.x` (management VLAN — SSH key rejected). px1=.10, px2=.20, px3=.30, px4=.40, px5=.50, px6=.60, px7(n5pro)=.70 or .91 | | OSDs | 11 total (6 HDD + 5 SSD) | | HDD mix | 4× ~3 TB (osd.1, 6, 8, 11) + 2× ~1 TB (osd.7, 10) | | EC Profile | `ec41-hdd` (k=4, m=1) → migrating to `ec31-hdd-large` (k=3, m=1) | | Pools | `media_ec` (EC), `hdd_disk` (Repl 3), `rbd` (Repl 3), `tm_disks` (Repl 2), `media_meta` (Repl 3), `vm_disks` (Repl 3) | | Ratios | nearfull=0.94, backfillfull=0.96, full=0.97 (failsafe hardcoded 0.97) | | Monitoring | Grafana CT141 (DS UID=PBFA97CFB590B2093), Prometheus, Alertmanager→Telegram | | Ceph Mon Network | `10.0.20.x` (storage VLAN) — NOT `10.0.30.x` (management) | All Ceph commands run from PVE master (**10.0.20.10** — storage network) unless noted. Non-Ceph nodes (e.g. n5pro) need `ceph.conf` + admin keyring copied and `rbd`/`ceph` commands run from a monitor node. ### SSH Access: 10.0.20.x Works, 10.0.30.x Does NOT PVE hosts have **two networks**: storage (10.0.20.x) and management (10.0.30.x). SSH with `~/.ssh/id_ed25519_proxmox` works on **10.0.20.x** but is rejected on **10.0.30.x** (key not in authorized_keys). Always use 10.0.20.x IPs for PVE host SSH: ```bash # CORRECT — storage network ssh -i ~/.ssh/id_ed25519_proxmox root@10.0.20.10 "ceph -s" # WRONG — management network (key rejected) ssh -i ~/.ssh/id_ed25519_proxmox root@10.0.30.10 "ceph -s" # Permission denied ``` PVE host IP mapping (storage network): px1=.10, px2=.20, px3=.30, px4=.40, px5=.50, px6=.60, px7=.70, n5pro=.91. ### Fallback: Ceph CLI via K8s ceph-csi Pods When PVE SSH is unavailable, `ceph` commands can be run from inside K8s `ceph-csi-rbd-nodeplugin` pods (the `csi-rbdplugin` container has `/usr/bin/ceph`). Requires a ceph.conf with monitor addresses and the `csi-rbd-secret` credentials: ```bash # On mgmt-runner (10.0.30.124): POD=$(kubectl get pods -n kube-system --no-headers | grep ceph-csi-rbd-nodeplugin | head -1 | awk '{print $1}') # Get cephx key from K8s secret USER_KEY=$(kubectl get secret csi-rbd-secret -n kube-system -o jsonpath='{.data.userKey}' | base64 -d) # Write ceph.conf into the pod and query kubectl exec -n kube-system $POD -c csi-rbdplugin -- bash -c " cat > /tmp/ceph.conf << 'EOF' [global] auth_cluster_required = cephx auth_service_required = cephx auth_client_required = cephx mon host = 10.0.20.50:6789,10.0.20.70:6789,10.0.20.40:6789 [client.kubernetes] key = $USER_KEY EOF ceph -c /tmp/ceph.conf --name client.kubernetes status " ``` Monitor addresses are in the `ceph-csi-config` ConfigMap: `kubectl get cm ceph-csi-config -n kube-system -o jsonpath='{.data.config\.json}'`. **Pitfall:** Not all monitors are always reachable. Check connectivity first (`bash -c "echo > /dev/tcp/IP/6789"` from a Galera VM or K8s node). Only use reachable monitors in the `mon host` list. **Dual-network pitfall:** PVE hosts have SSH on `10.0.20.x` (storage network, port 22 open) but NOT on `10.0.30.x` (management network, SSH key rejected). Always use `10.0.20.10` for `ssh root@10.0.20.10`. The `10.0.30.x` IPs are for VM bridges, not PVE host management. **Fallback when PVE SSH is unavailable:** Run `ceph` commands from inside a `ceph-csi-rbd-nodeplugin` pod in K8s (see Monitoring section below for the full procedure). ⚠️ **Pitfall: PVE hosts are on 10.0.20.x (storage net), NOT 10.0.30.x (management net).** The 10.0.30.x addresses are for VM bridges, not PVE host SSH. Always use `root@10.0.20.10` for Ceph commands. If SSH to 10.0.30.x PVE hosts fails with "Permission denied", you're on the wrong network — check this skill first, not the SSH key. ### Fallback: Ceph CLI via K8s ceph-csi Pods When PVE host SSH is unavailable (key rotation, host down, network change), `ceph` CLI is available inside the `csi-rbdplugin` container of `ceph-csi-rbd-nodeplugin` pods in K8s. Steps: 1. Find monitor addresses from the `ceph-csi-config` ConfigMap: ```bash kubectl get cm ceph-csi-config -n kube-system -o jsonpath='{.data.config\.json}' # Parse JSON for "monitors" array (e.g. ["10.0.20.10:6789","10.0.20.20:6789",...]) ``` 2. Get cephx credentials from the `csi-rbd-secret`: ```bash kubectl get secret csi-rbd-secret -n kube-system -o jsonpath='{.data.userKey}' | base64 -d ``` 3. Write a temporary ceph.conf into the pod and query: ```bash POD=$(kubectl get pods -n kube-system --no-headers | grep ceph-csi-rbd-nodeplugin | head -1 | awk '{print $1}') # Create /tmp/ceph.conf with [global] mon host = and [client.kubernetes] key = # Then: kubectl exec -n kube-system $POD -c csi-rbdplugin -- ceph -c /tmp/ceph.conf --name client.kubernetes status ``` ⚠️ Not all monitors may be reachable from K8s nodes — test connectivity first. Only use monitors that respond on port 6789. ## Decision Trees ### OSD Creation ``` Is target a standard PVE node? ├─ YES → pveceph osd create /dev/sdX [--wal_dev /dev/nvme0n1pN] │ ⚠️ pveceph CANNOT use LVM device-mapper paths (/dev/dm-*) as WAL/DB │ If WAL/DB is on an LV: use ceph-volume directly (see below) └─ NO (non-PVE node like n5pro) ├─ Has /etc/ceph/ceph.conf symlink to /etc/pve/ceph.conf? │ ├─ NO → rm /etc/ceph/ceph.conf && ln -s /etc/pve/ceph.conf /etc/ceph/ceph.conf │ └─ YES → proceed ├─ WAL/DB on raw partition? → pveceph osd create /dev/sdX --wal_dev /dev/nvme0n1pN └─ WAL/DB on LVM LV? → ceph-volume lvm prepare --bluestore --data /dev/sdX --block.db /dev/pve/wal-db-osd-X (WAL auto-collocates on DB — do NOT specify --block.wal separately) ``` ### Recovery / Rebalance ``` ceph -s shows degraded/misplaced/backfill? ├─ Any OSD > full_ratio (0.97)? │ └─ EMERGENCY: ceph osd set-full-ratio 0.99 → reweight overfull OSD → fix root cause ├─ backfill_toofull PGs > 0? │ └─ Target OSDs above backfillfull_ratio → raise ratio (keep ordering!) or reweight ├─ Recovery speed too slow? │ └─ Tune: osd_max_backfills, osd_recovery_op_priority, osd_recovery_sleep │ ⚠️ ceph config set osd X does NOT propagate to running OSDs │ Use: ceph tell osd.* injectargs "--osd_max_backfills 3" + verify with ceph daemon osd.N config get ├─ Recovery stuck at 0 MiB/s with remapped PGs? │ └─ Check for sentinel 2147483647 in up-sets → EC with OSD count == k+m → reset all weights to 1.0 └─ upmap balancer not helping? └─ Balancer skips when >5% misplaced → wait for recovery, then balancer resumes ``` ### EC Pool with Mixed Disk Sizes ``` EC pool has OSDs of different capacities? ├─ Small OSDs filling disproportionately? │ └─ ROOT CAUSE: EC stores equal-sized chunks regardless of OSD capacity │ ceph osd reweight has marginal effect when k+m ≈ total OSD count │ SOLUTION: │ 1. Define device classes: hdd-large, hdd-small │ 2. Create separate CRUSH root (hdd-root) for all HDDs │ 3. EC rule: take hdd-root~hdd-large (only large OSDs) │ 4. Replicated rule: take hdd-root (all HDDs, no class filter) │ 5. New EC profile with k adjusted to large-OSD count │ 6. Migrate RBD images to new pool (online via rbd migration) └─ Artificially low CRUSH weights on small OSDs? └─ ANTI-PATTERN: causes vicious cycle → reset to true device size after separating classes ``` ## Core Operations ### 1. OSD Lifecycle #### Create OSD with WAL/DB on NVMe (LVM-based) ```bash # 1. Create WAL/DB LV on NVMe lvcreate -L 30G -n wal-db-osd-X pve # 2. Prepare OSD — ONLY --block.db, WAL auto-collocates ceph-volume lvm prepare --bluestore \ --data /dev/sdX \ --block.db /dev/pve/wal-db-osd-X # 3. Activate (use osd_id + fsid from prepare output) ceph-volume lvm activate # 4. Start service systemctl enable --now ceph-osd@ # 5. Set CRUSH host + device class ceph osd crush set osd. root=default host= ceph osd crush set-device-class hdd osd. ``` ⚠️ **NEVER** specify `--block.wal` AND `--block.db` on the same LV — BlueStore opens the device exclusively → "Device or resource busy". #### Destroy and Rebuild OSD ```bash systemctl stop ceph-osd@ ceph osd down ceph osd destroy --yes-i-really-mean-it ceph-volume lvm zap /dev/sdX --destroy # Then recreate as above ``` #### Hot-Plug SATA Drive Detection ```bash # JMB58x controller doesn't auto-scan for h in /sys/class/scsi_host/host{0,1,2,3,4}; do echo "- - -" > ${h}/scan done sleep 2 lsblk -d -o NAME,MODEL,SERIAL,SIZE,TRAN,ROTA,STATE | grep -v rbd | grep -v loop ``` #### Expand BlueFS DB (Online LV Extend + Offline BlueStore Resize) When BlueFS spillover occurs (RocksDB grows beyond DB LV onto HDD), expand the DB LV: ```bash # 1. Extend the LV (online, no OSD stop needed for LVM) lvextend -L +20G pve/ceph-db # 2. Stop the OSD — bluefs-bdev-expand CANNOT run while OSD is active # Error if attempted online: "Device or resource busy" + assert crash systemctl stop ceph-osd@ # 3. Expand BlueFS awareness of the new DB size ceph-bluestore-tool bluefs-bdev-expand --path /var/lib/ceph/osd/ceph- # Output: "1 : Expanding to 0x... (... GiB)" + "size updated" # 4. Restart OSD systemctl start ceph-osd@ # 5. Verify BlueFS sees the new size ceph daemon osd. bluefs stats | head -3 # Should show: "1 : device size 0x...(XX GiB)" # 6. Verify OSD is back up in cluster ceph osd tree | grep "osd." ``` ⚠️ **`bluefs-bdev-expand` MUST be run with OSD stopped.** Running it online crashes with `FAILED ceph_assert(r == 0)` — the block device is locked by the running OSD. The `ceph daemon osd.N` admin socket does NOT have an expand command (unlike config get/set). The only path is: stop OSD → bluestore-tool → start OSD. ⚠️ **Spillover warning clears asynchronously.** After expansion, `BLUEFS_SPILLOVER` health warning persists until BlueFS compacts RocksDB back onto the expanded DB device. This happens in the background over minutes to hours. #### When NOT to Expand (Cosmetic Spillover) **Spillover < 1 MiB with >50% DB free → cosmetic, do NOT expand.** Tiny spillovers (e.g. 128 KiB) are RocksDB metadata overflow that self-resolves on next compaction. Stopping an OSD for `bluefs-bdev-expand` triggers PG degradation and recovery I/O — worse than the warning itself. **Expand only when:** spillover > 1 GiB OR DB device >80% full. This indicates sustained RocksDB growth that won't self-resolve. ### 2. CRUSH Topology #### Device Class Separation ```bash # Assign device classes ceph osd crush set-device-class hdd-large osd.1 osd.6 osd.8 osd.11 ceph osd crush set-device-class hdd-small osd.7 osd.10 # Create separate root ceph osd crush add-bucket hdd-root root # Move HDD hosts to new root (⚠️ TRIGGERS REMAPPING!) ceph osd crush move root=hdd-root ``` #### CRUSH Rules ```bash # EC rule — only large HDDs ceph osd erasure-code-profile set ec31-hdd-large \ k=3 m=1 crush-root=hdd-root crush-device-class=hdd-large \ crush-failure-domain=osd plugin=jerasure technique=reed_sol_van ceph osd crush rule create-erasure media_ec_large ec31-hdd-large # Replicated rule — ALL HDDs (no class filter) ceph osd crush rule create-replicated replicated_hdd_all hdd-root host ``` #### Weight Management | Command | Scope | Persistence | Use Case | |---------|-------|-------------|----------| | `ceph osd reweight N X` | Runtime reweight | Lost on restart | Emergency drain of overfull OSD | | `ceph osd crush reweight osd.N X` | CRUSH weight | Permanent | Match weight to device size | **Correct CRUSH weights by device size:** | Device | Weight | |--------|--------| | 1 TB | ~0.96 | | 2 TB | ~1.82 | | 3 TB | ~2.73 | | 3.6 TiB | ~3.64 | ⚠️ **Anti-pattern**: Lowering CRUSH weight to "protect" a full OSD. This prevents new PGs but doesn't move existing data → vicious cycle. ### 3. EC Pool Migration EC pools cannot change k/m in-place. Must create new pool and migrate. ```bash # 1. Create new EC pool ceph osd pool create media_ec_new 128 128 erasure ec31-hdd-large ceph osd pool set media_ec_new ec_overwrites ceph osd pool application enable media_ec_new rbd # 2. Create new metadata pool ceph osd pool create media_meta_new 32 # 3. Online RBD migration (VMs keep running!) for img in $(rbd ls media_meta); do rbd migration prepare media_meta/${img} media_meta_new/${img} --data-pool media_ec_new rbd migration execute media_meta_new/${img} rbd migration commit media_meta_new/${img} done # 4. Cleanup old pool ceph osd pool delete media_ec --yes-i-really-really-mean-it ceph osd pool rename media_ec_new media_ec ``` Timeline: `prepare` (instant) → `execute` (30-60 min per 500 GiB) → `commit` (atomic switch). ### 4. Recovery Tuning #### Full Recovery Recipe ```bash # 1. Unblock backfill (keep ratio ordering: nearfull < backfillfull < full ≤ 0.97) ceph osd set-nearfull-ratio 0.94 ceph osd set-backfillfull-ratio 0.96 ceph osd set-full-ratio 0.97 # 2. Reweight overfull OSD (emergency, temporary) ceph osd reweight 0.8 # 3. Set recovery params (persistent for future OSDs) ceph config set osd osd_max_backfills 3 ceph config set osd osd_recovery_op_priority 10 ceph config set osd osd_recovery_sleep 0 # 4. Push to running OSDs (config set alone does NOT propagate!) ceph tell osd.* injectargs "--osd_max_backfills 3" # 5. Verify EACH OSD picked up the config ceph daemon osd. config get osd_max_backfills # Must return "3" — if "1", config not applied ``` #### Post-Recovery Cleanup ```bash # Restore defaults ceph config set osd osd_max_backfills 1 ceph tell osd.* injectargs "--osd_max_backfills 1" ceph osd reweight 1.0 # Ratios: keep at 0.94/0.96/0.97 (safe for mixed-size cluster) ``` ### 5. RBD Orphan Cleanup ```bash # 1. Get cluster-wide VM inventory (NOT per-node qm config!) pvesh get /cluster/resources --type vm --output-format json | python3 -c " import sys, json for g in json.load(sys.stdin): print(g['vmid']) " | sort -n | tr '\n' ' ' # 2. Fast orphan check for specific image (O(1)) grep -rl "vm-NNN" /etc/pve/qemu-server/ /etc/pve/lxc/ 2>/dev/null # Exit 0 = referenced. Exit 1 = orphan. # 3. Delete with protected snapshot rbd snap unprotect "hdd_disk/base-NNN-disk-0"@__base__ rbd snap purge "hdd_disk/base-NNN-disk-0" rbd rm "hdd_disk/base-NNN-disk-0" ``` ⚠️ **Always use `pvesh get /cluster/resources --type vm`** for VM existence — per-node `qm config` only sees local VMs → false positives. ⚠️ **Non-Ceph nodes cannot run `rbd` commands** — run from a monitor node (e.g. proxmox7 = 10.0.20.70). ### 6. PG Management ```bash # Reduce PG count (must disable autoscale first!) ceph osd pool set pg_autoscale_mode off ceph osd pool set pgp_num 32 ceph osd pool set pg_num 32 # Merge happens AFTER all PGs reach active+clean # Raise mon_max_pg_per_osd (temporary relief) ceph config set mon mon_max_pg_per_osd 400 ceph tell mon.* injectargs --mon_max_pg_per_osd=400 ``` ## Pitfalls ### Ratio Ordering (HEALTH_ERR) `nearfull < backfillfull < full ≤ osd_failsafe_full_ratio (0.97 hardcoded)`. Violation → `HEALTH_ERR: OSD_OUT_OF_ORDER_FULL`. Always set all three ratios together. ### injectargs Reliability `ceph tell osd.* injectargs` returns empty `{}` which looks like failure but may have succeeded. **Always verify** with `ceph daemon osd.N config get `. Most reliable: `ceph config set` (persistent) + `injectargs` (runtime) + verify. ### EC Reweight with Minimal OSDs **Never reweight OSDs in EC pools when OSD count == k+m.** CRUSH places sentinel `2147483647` ("no OSD available") in up-sets → remapped PGs can never recover. Fix: reset all weights to 1.0. General rule: only reweight when OSD count > k+m. ### pveceph LVM WAL/DB Limitation `pveceph osd create --wal_dev /dev/pve/wal-db-osd-X` fails with "unable to get device info for '/dev/dm-2'". pveceph cannot resolve LVM device-mapper symlinks. Use `ceph-volume lvm prepare` directly with `--block.db`. ### BlueStore WAL/DB Same Device Specifying `--block.wal` AND `--block.db` on the same LV → "Device or resource busy". BlueStore opens WAL exclusively. Solution: specify ONLY `--block.db` — WAL collocates automatically. ### ceph config set Doesn't Propagate to Running OSDs `ceph config set osd osd_max_backfills 3` updates MON config store but running OSDs keep old value. Must use `ceph tell osd.* injectargs` for runtime change, or restart OSDs. Always verify with `ceph daemon osd.N config get`. ### backfillfull Is Derived, Not Settable Pool-level `backfillfull` flag is derived from OSD conditions, not a settable pool flag. `ceph osd pool unset backfillfull` → EINVAL. Fix: raise `backfillfull_ratio` or reweight overfull OSDs. ### Ghost RBD Images Images appearing in `rbd ls` but `rbd info` returns empty → header gone, data objects remain. Need `rados rm` purge, not `rbd rm`. ### pvesm list Fails When Pool Backfillfull `pvesm list hdd_disk` returns empty when Ceph pool is in backfillfull state. Fall back to `rbd ls ` for image listing. ## Monitoring ```bash # Quick health ceph -s # Per-OSD utilization ceph osd df # PG state summary ceph pg dump pgs_brief | awk '{print $2}' | sort | uniq -c | sort -rn # Recovery progress ceph -s | grep -E "recovery:|misplaced" # Balancer status ceph balancer status # Slow ops ceph health detail | grep -A5 "slow" ``` ### Ceph Status via ceph-csi Pod (Fallback When PVE SSH Unavailable) When PVE hosts are unreachable via SSH (key rotation, network issue, host down), but K8s nodes can reach the Ceph monitors, run `ceph` commands from inside a `ceph-csi-rbd-nodeplugin` pod: ```bash # 1. Find the ceph-csi pod (on mgmt-runner or any K8s node with kubectl) POD=$(kubectl get pods -n kube-system --no-headers | \ grep ceph-csi-rbd-nodeplugin | head -1 | awk '{print $1}') # 2. Get monitor addresses from the ceph-csi-config ConfigMap MONS=$(kubectl get cm ceph-csi-config -n kube-system \ -o jsonpath='{.data.config\.json}' | \ python3 -c "import sys,json; print(','.join(json.load(sys.stdin)[0]['monitors']))") # 3. Get the cephx key from the csi-rbd-secret KEY=$(kubectl get secret csi-rbd-secret -n kube-system \ -o jsonpath='{.data.userKey}' | base64 -d) # 4. Write ceph.conf + keyring into the pod, then run ceph commands kubectl exec -n kube-system $POD -c csi-rbdplugin -- bash -c " cat > /tmp/ceph.conf << EOF [global] auth_cluster_required = cephx auth_service_required = cephx auth_client_required = cephx mon host = $MONS [client.kubernetes] key = $KEY " ceph -c /tmp/ceph.conf --name client.kubernetes status ceph -c /tmp/ceph.conf --name client.kubernetes osd df " ``` **Key details:** - The `csi-rbd-plugin` container has `/usr/bin/ceph` installed - Monitor addresses are in the `ceph-csi-config` ConfigMap (may be on a different subnet like `10.0.20.x` — the storage VLAN) - The cephx key is in the `csi-rbd-secret` Secret (field `userKey`) - Not all monitors need to be up — 3 of 7 was sufficient in practice - `client.kubernetes` has limited permissions (reads status, OSD df, pool df) but NOT admin operations (no `ceph osd set-full-ratio` etc.) **Pitfall:** Monitor addresses in `ceph-csi-config` may be on a different subnet (e.g. `10.0.20.x` storage VLAN) than the management network (`10.0.30.x`). K8s pods can usually reach both, but Galera VMs may only see `10.0.30.x`. **Pitfall:** The `ceph -c /tmp/ceph.conf` approach requires writing the config file into the pod. If `kubectl cp` is blocked, use `kubectl exec -- bash -c "cat > /tmp/ceph.conf << EOF ..."` via a heredoc through SSH to the mgmt-runner. ## Recovery Monitoring Cron Pattern ``` Schedule: every 30 minutes Report: PGs remapped/backfilling, % misplaced, OSD utilization, recovery speed Alert: any OSD down or health ERR Announce "COMPLETE" when 0 remapped PGs ``` ## Cross-References Detailed session logs and edge cases in `proxmox-ve-administration` skill references: - `references/ceph-osd-create-hotplug-2026-07.md` — Full OSD hot-plug walkthrough - `references/ceph-ec-mixed-size-optimization-2026-07.md` — Complete EC migration plan - `references/ceph-recovery-acceleration-2026-07.md` — Advanced recovery tuning - `references/ceph-crush-weight-optimization-2026-07.md` — Weight alignment for mixed sizes - `references/ceph-pg-management-2026-07.md` — PG count reduction + EC reweight pitfall - `references/ceph-pool-full-recovery-2026-07.md` — Pool-full diagnosis + OSD preparation - `references/ceph-orphaned-rbd-cleanup-2026-07.md` — Full orphan audit methodology - `references/ceph-recovery-ec-reweight-2026-07.md` — EC reweight stagnation fix - `references/ceph-pg-osd-management-2026-07.md` — PG/OSD rebalancing - `references/ceph-ec-pool-and-service-migration-2026-07.md` — EC pool + service migration - `references/ceph-pg-reduction-and-osd-rebalance-2026-07.md` — PG reduction + rebalance - `references/rbd-pool-migration-2026-07.md` — Offline RBD pool migration - `references/hdd-selection-for-osd-2026-07.md` — Drive selection criteria - `references/bluefs-db-expansion-2026-07.md` — BlueFS DB LV expansion + cosmetic spillover assessment Compound-learning solution docs in `~/docs/solutions/`: - `bug-fixes/2026-07-13-pveceph-wal-lvm-device-mapper-limitation.md` - `bug-fixes/2026-07-13-ceph-ratio-ordering-constraint.md` - `architecture/2026-07-13-ec-mixed-disk-sizes-optimization.md` - `tooling/2026-07-13-injectargs-vs-config-set-reliability.md`