Files
memory/patterns/ceph-ssd-wear-ec-pool.md
Dominik Schön bbbe4f8985 feat: patterns/ directory + skill-impact tracker (WikiSkill-inspired)
- New patterns/ directory with 7 initial failure-mode patterns (PAT-001..007)
- skill-impact.md audit trail for skill modifications
- _template.md for future pattern creation
- index.md updated with Patterns section
- log.md entry for this change
- Inspired by arXiv:2608.27454 (WikiSkill)
2026-08-30 11:16:19 +00:00

2.6 KiB

pattern_id, title, category, severity, status, first_observed, last_updated, related_systems, related_solution_docs, related_skills
pattern_id title category severity status first_observed last_updated related_systems related_solution_docs related_skills
PAT-007 Ceph EC pool unusable + SSD wear-level failing — monitor pg states storage high active 2026-07 2026-08-30
ceph-cluster
docs/solutions/architecture/2026-07-12-ceph-ec-pool-no-rebalance-headroom.md
docs/solutions/bug-fixes/2026-07-13-ceph-ratio-ordering-constraint.md
docs/solutions/bug-fixes/2026-07-13-pveceph-wal-lvm-device-mapper-limitation.md
ceph-cluster-administration
ceph-ec-incomplete-pg-recovery

Ceph EC pool unusable + SSD wear-level failing — monitor pg states

Symptom

Multiple Ceph failure modes manifest simultaneously:

  1. EC pool (pool 8/media_ec) becomes unusable — PGs stuck in incomplete state
  2. SSD OSD (osd.3 on px4) reports Wear Leveling Count SMART attribute indicating imminent failure
  3. ceph health shows HEALTH_ERR

Root Cause

These are compounded issues:

  1. EC pool incomplete PGs: After OSD loss, EC pools cannot recover without enough surviving shards. The minimum copies requirement for EC k+m encoding is stricter than replicated pools.
  2. SSD wear-level failure: Consumer-grade SSDs in Ceph OSD duty cycle reach wear limits. Once Wear Leveling Count exceeds threshold, the SSD becomes read-only or unreliable.
  3. Weight imbalance: Unequal OSD weights cause CRUSH placement failures, leaving PGs perpetually remapped.

Mitigation

  1. EC pool recovery: Use the ceph-ec-incomplete-pg-recovery skill — specialized procedure for forcing EC PG recovery after OSD loss
  2. Replace failing SSD: Mark OSD outdestroyzap → physically replace → recreate
  3. Fix weight imbalance: Equalize pg_numpgp_num, set nopgchange=true, reweight OSDs proportionally to disk capacity
  4. Remove empty CRUSH buckets: Clean up stale host entries from decommissioned nodes

Prevention

  • Monitor SMART attributes monthly — alert on wear-level > 80%
  • Never use consumer SSDs for Ceph OSDs in production (use enterprise grade with PLP)
  • Keep EC pool k+m ratio proportional to available OSDs (don't over-provision redundancy)
  • Regular ceph pg dump audits for stuck/unactive PGs
  • Separate PBS onto its own tier (CephFS), away from RBD pools

Evidence

  • Multiple incidents Jul 2026 (OSD 0+2 destruction, OSD 3 wear-level, EC pool unusable)
  • Solution docs: 2026-07-12-ceph-ec-pool-no-rebalance-headroom.md, 2026-07-13-ceph-ratio-ordering-constraint.md
  • MEMORY.md entry: "Ceph:HEALTH_ERR. Pool8/media_ec unusable. px4 SSD(osd.3) WearLevel FAILING."