# Ceph Recovery Acceleration — Advanced Tuning — 2026-07-05 ## Context Follow-up to `ceph-pool-full-recovery-2026-07.md`. After unblocking backfill by raising `backfillfull_ratio` and reweighting OSD 7, recovery was progressing but slowly (22 MiB/s, 5 obj/s, only 2 PGs backfilling). This session applied additional tuning to accelerate recovery. ## Additional Recovery Parameters ### `osd_recovery_op_priority` (effective) ```bash # Set globally via MON config ceph config set osd osd_recovery_op_priority 10 # Default: 3. Range: 1-63. Higher = recovery ops preempt client ops. # Verify on running OSD: ceph daemon osd.7 config get osd_recovery_op_priority # Returns: {"osd_recovery_op_priority": "10"} ``` Unlike `osd_max_backfills`, this one DOES apply at runtime via `ceph config set` — no restart needed. OSDs pick it up within seconds. ### `osd_recovery_sleep` (set to 0) ```bash ceph config set osd osd_recovery_sleep 0 # Default: 0.001 (1ms pause between recovery ops). Set to 0 for maximum throughput. ``` ### `osd_recovery_max_chunk` ```bash ceph config set osd osd_recovery_max_chunk 104857600 # Default: 8MiB. Set to 100MiB for larger transfer windows. ``` ### `ceph tell osd.N config set` — Alternative to `injectargs` Newer Ceph (Reef/Squid) supports `ceph tell osd.N config set` as a cleaner alternative to `injectargs`: ```bash # Instead of: ceph tell osd.7 injectargs "--osd_max_backfills 3" # Use: ceph tell osd.7 config set osd_max_backfills 3 ``` However, in testing, `ceph config set osd osd_max_backfills 3` (MON-level) did NOT propagate to running OSDs — the OSDs kept using value 1. The `ceph tell osd.N config set` also didn't reliably stick. The most reliable method remains: ```bash # Most reliable runtime method: ceph tell osd.7 injectargs "--osd_max_backfills 3" # Verify: ceph daemon osd.7 config get osd_max_backfills # Must return "3" — if still "1", the config hasn't been applied. ``` ### UPDATE 2026-07-13: `ceph tell osd.* config set` DOES Work In a follow-up session, `ceph tell osd.* config set osd_max_backfills 3` DID successfully apply to running OSDs. Verification: ```bash ceph tell osd.11 config get osd_max_backfills # Returns: {"osd_max_backfills": "3"} ``` The difference from the earlier session: `ceph config set osd osd_max_backfills 3` (MON-level persistent config) was set FIRST, then `ceph tell osd.* config set` was used to push it to running OSDs. The combination worked. Injectargs returned empty `{}` (appeared to fail) but the config took effect. **Recommendation**: Use BOTH `ceph config set` (persistent for future OSDs) AND `ceph tell osd.* config set` (runtime for current OSDs). Then verify with `ceph tell osd.N config get`. ### `osd_recovery_max_active` — NOT effectively tunable `ceph config set osd osd_recovery_max_active 3` accepts the value but OSDs report 0 (meaning auto/default). `injectargs` also shows empty. This appears to be a Ceph limitation — the parameter is managed internally. Do not waste time trying to force it. ## Full Recovery Tuning Recipe ```bash # 1. Unblock backfill (if backfillfull) ceph osd set-backfillfull-ratio 0.97 ceph osd set-nearfull-ratio 0.97 # 2. Reweight overfull OSD ceph osd crush reweight osd.7 0.50 # 3. Set recovery parameters via MON config (applies to new/restarted OSDs) ceph config set osd osd_max_backfills 3 ceph config set osd osd_recovery_op_priority 10 ceph config set osd osd_recovery_sleep 0 ceph config set osd osd_recovery_max_chunk 104857600 # 4. Force runtime application on each HDD OSD (MON config doesn't propagate to running OSDs) for osd in 1 6 7 8 10; do ceph tell osd.$osd injectargs "--osd_max_backfills 3" ceph tell osd.$osd injectargs "--osd_recovery_op_priority 10" done # 5. Verify each OSD picked up the config for osd in 1 6 7 8 10; do echo "=== osd.$osd ===" ceph daemon osd.$osd config get osd_max_backfills ceph daemon osd.$osd config get osd_recovery_op_priority done ``` ## Recovery Speed Progression | Stage | Recovery Speed | Concurrent Backfills | Notes | |-------|---------------|---------------------|-------| | Before intervention | 0 MiB/s (stuck) | 0 | backfillfull blocks all | | After ratio + reweight | 22-23 MiB/s, 5 obj/s | 1-2 | Backfill unblocked | | After injectargs (max_backfills=3) | 34 MiB/s, 8 obj/s | 2-3 | Even with max=3, Ceph schedules conservatively | | After op_priority=10 + sleep=0 | ~34 MiB/s, 8 obj/s | 2-3 | Priority helps under client I/O load | Even with `osd_max_backfills=3`, expect only 2-3 PGs backfilling simultaneously — Ceph schedules conservatively. The throughput improvement comes from larger transfer windows and reduced pauses, not dramatically more parallel PGs. ## OSD Restart Considerations If `injectargs` doesn't work, restarting the OSD daemon forces it to pick up the MON config: ```bash # ⚠️ MUST restart on the correct node! Use ceph osd find to locate: ceph osd find 7 # Returns JSON with host name and IP # SSH to THAT node and restart: ssh root@ "systemctl restart ceph-osd@7" # ❌ WRONG: restarting on a node that doesn't host the OSD → # "OSD data directory /var/lib/ceph/osd/ceph-7 does not exist; bailing out" # After restart, OSD goes through peering — may take 30-60s # During peering, PGs may show as "peering" or "activating" ``` ## Monitoring Recovery ```bash # Quick status ceph -s | grep -E "osd:|pgs:|recovery:|misplaced" # Detailed PG states ceph -s | grep -A15 "pgs:" # Per-OSD utilization (track if overfull OSD is draining) ceph osd df | grep "^ 7 " # Estimated duration: misplaced_objects / objects_per_second ``` ## Recovery Monitoring Cron Pattern ``` Schedule: every 30 minutes Command: ssh to monitor node, run `ceph status`, report: - PGs remapped/backfilling count - % misplaced objects - OSD 7 utilization - Recovery speed - Alert if any OSD down or health ERR - Announce "COMPLETE" when 0 remapped PGs ``` ## Pitfalls: Setting backfill_full_ratio (2026-07-12) Three commands FAIL when trying to raise the backfill full ratio — only ONE works: ```bash # ✅ CORRECT — CLI command with hyphen (not underscore): ceph osd set-backfillfull-ratio 0.97 # Verify: ceph osd dump | grep backfillfull_ratio → "backfillfull_ratio 0.97" # ❌ WRONG — underscore variant: ceph osd set-backfillfullratio 0.97 # → "no valid command found" (EINVAL) # ❌ WRONG — MON config set: ceph config set mon mon_osd_backfillfull_ratio 0.97 # → "mon_osd_backfillfull_ratio is special and cannot be stored by the mon" # ❌ WRONG — injectargs on specific OSD: ceph tell osd.10 injectargs "--osd_backfill_full_ratio 0.97" # → "failed to parse arguments: --osd_backfill_full_ratio,0.97" (EINVAL) # injectargs only accepts runtime-tunable OSD params, not full-ratio thresholds. ``` The ratio is a MON-level OSD map parameter, not a per-OSD config. It can ONLY be changed via `ceph osd set-backfillfull-ratio` (CLI command, not config set). Similarly: `ceph osd set-nearfull-ratio` and `ceph osd set-full-ratio` for the other thresholds. ## Pitfalls: Pool-Level `backfillfull` Flag (2026-07-13) When pools show the `backfillfull` flag in `ceph osd dump`, attempting to unset it at the pool level FAILS: ```bash # ❌ WRONG — no such pool flag: ceph osd pool unset 6 backfillfull # → "no valid command found; 1 closest matches: osd pool unset noautoscale" # ❌ WRONG — all pool IDs tried, all fail: for pool in 1 5 6 7 8 9 10; do ceph osd pool unset $pool backfillfull; done # All return EINVAL ``` The `backfillfull` flag on a pool is a DERIVED condition — it appears when any OSD in the pool's up/acting set is above the `backfillfull_ratio`. It is NOT a settable pool flag. To clear it, fix the underlying OSD condition: raise `backfillfull_ratio` or reweight the overfull OSD. ```bash # ✅ CORRECT — raise the ratio (clears backfillfull on all pools): ceph osd set-backfillfull-ratio 0.99 # ✅ ALSO — reweight the overfull OSD: ceph osd reweight 7 0.8 ``` ## Pitfalls: `osd_recovery_max_active` Behavior (2026-07-13) `ceph config set osd osd_recovery_max_active 3` accepts the value and `ceph config get osd osd_recovery_max_active` returns 3. However, `ceph tell osd.N config get osd_recovery_max_active` may return empty string — the OSD interprets 0 or empty as "auto/default". Despite this, setting it alongside `osd_max_backfills` appeared to help overall. Don't rely on this parameter alone — use `osd_max_backfills` as the primary lever. ## Pitfalls: `full ratio(s) out of order` (2026-07-13) When adjusting full/nearfull/backfillfull ratios, you MUST preserve the ordering: `nearfull_ratio < backfillfull_ratio < full_ratio <= osd_failsafe_full_ratio`. The `osd_failsafe_full_ratio` is hardcoded at **0.97** and cannot be changed. If you set `full_ratio` above 0.97, the cluster goes to `HEALTH_ERR`: ``` [ERR] OSD_OUT_OF_ORDER_FULL: full ratio(s) out of order osd_failsafe_full_ratio (0.97) < full_ratio (0.98), increased ``` **Safe ratio combinations for a cluster with an OSD at 95%:** ```bash # ✅ CORRECT — all ratios below failsafe (0.97), properly ordered: ceph osd set-nearfull-ratio 0.94 ceph osd set-backfillfull-ratio 0.96 ceph osd set-full-ratio 0.97 # == failsafe, OK # ❌ WRONG — full_ratio above failsafe: ceph osd set-full-ratio 0.98 # → HEALTH_ERR: out of order ``` **With an OSD at 95.2% utilization:** - `nearfull_ratio` 0.94 → OSD at 95.2% shows `nearfull` (WARN, acceptable) - `backfillfull_ratio` 0.96 → OSD at 95.2% is NOT backfillfull → backfill proceeds - `full_ratio` 0.97 → OSD at 95.2% is NOT full → client I/O continues If the OSD drains below 0.94, all flags clear automatically. ## Pitfalls: `ceph osd reweight` vs `ceph osd crush reweight` (2026-07-13) Two DIFFERENT commands that both change OSD weighting: ```bash # Runtime reweight — temporary, lost on OSD restart: ceph osd reweight 7 0.8 # Changes the `reweight` column in `ceph osd df`. Persists until OSD restart # or cluster-wide reweight-by-utilization. Does NOT change CRUSH weight. # Good for emergency drain of an overfull OSD. # CRUSH weight change — permanent, affects PG distribution: ceph osd crush reweight osd.7 0.96 # Changes the CRUSH `weight` column. Survives restarts. Changes how many # PGs the OSD gets assigned going forward. # Use this to match weight to actual device size. ``` **Decision guide:** - Emergency: OSD is 96%+ and blocking recovery → `ceph osd reweight` (fast, temporary) - Permanent: OSD weight doesn't match device size → `ceph osd crush reweight` (slow, proper) - After recovery: Reset `reweight` to 1.0, set `crush weight` to true device size ## Post-Recovery Cleanup After all PGs return to active+clean: ```bash # Restore default ratios ceph osd set-backfillfull-ratio 0.95 ceph osd set-nearfull-ratio 0.93 # Restore default backfill speed ceph config set osd osd_max_backfills 1 ceph tell osd.* injectargs "--osd_max_backfills 1" # Restore OSD weight (if it was reduced) ceph osd crush reweight osd.7 1.0 # or original weight # Verify cluster health ceph health ceph -s ```