Files
hermes-skills/devops/proxmox-ve-administration/references/ceph-pool-full-recovery-2026-07.md
T

9.3 KiB

Ceph Pool-Full Diagnosis & Recovery — 2026-07-04

Symptoms

After adding CT 137's 20 GB RBD to media_ec, 6 pools simultaneously showed 100% full (MAX AVAIL = 0 B) despite the HDD class having 6.3 TB free.

Root Cause

A single OSD (osd.10, 982 GB) reached exactly 95% utilization = full_ratio. Ceph blocks ALL writes to PGs mapped to a full OSD. Since nearly every pool has PGs on osd.10, all pools became unwritable — even though osd.6 (2.8 TB) had 2.5 TB free.

Why osd.10 Filled First

CRUSH distributes data by OSD weight. osd.10 and osd.7 are ~1 TB disks; osd.6 and osd.8 are ~2.8 TB disks. The smaller OSDs filled faster because data grew unevenly. The VAR column in ceph osd df showed osd.10 at 2.13x average utilization vs osd.6 at 0.20x.

Diagnosis Commands

# Per-OSD utilization (look for %USE > 90% and high VAR)
ceph osd df

# Pool availability (MAX AVAIL = 0 means full)
ceph df

# Cluster health (look for "full osd" and "pool(s) full")
ceph -s

# Check full ratio threshold
ceph osd dump | grep full_ratio
# Default: full_ratio 0.95, nearfull_ratio 0.93, backfillfull_ratio 0.95

Recovery Procedure

Step 1: Raise full_ratio temporarily

ceph osd set-full-ratio 0.97

This unblocks writes immediately. Pools regain MAX AVAIL > 0. Safe as a temporary measure while rebalancing.

Step 2: Drain overloaded OSDs

# Automatic: reweight all overloaded OSDs at once
ceph osd reweight-by-utilization
# Moves PGs away from OSDs above average utilization threshold

# Manual: target specific OSDs for aggressive draining
ceph osd reweight 10 0.5    # drain to 50% of CRUSH weight
ceph osd reweight 7 0.8     # moderate drain

Step 3: Speed up backfill (optional, moderate)

# Allow 2 backfills per OSD (default: 1)
ceph config set osd osd_max_backfills 2
# Or inject directly:
ceph tell osd.<N> injectargs "--osd_max_backfills=2"

# Reduce recovery sleep (default: 0.1s)
ceph config set osd osd_recovery_sleep 0.05

Tradeoffs of increasing backfill speed:

  • Higher client I/O latency (VMs feel slower)
  • More "slow ops" warnings (we already had 4)
  • Higher CPU/IO load on stressed OSDs
  • May worsen BlueFS spillover

Recommendation: Only double osd_max_backfills (1→2) and halve recovery_sleep (0.1→0.05). Don't go aggressive on production clusters.

Step 4: Verify recovery

# Watch recovery progress
ceph -s | grep -A2 "io:"
# recovery: 32 MiB/s, 8 objects/s

# Track misplaced objects
ceph -s | grep "misplaced"
# 353960/1858594 objects misplaced (19%)

# Estimate time: misplaced_objects / objects_per_second
# Or: misplaced_data_bytes / recovery_rate_bytes_per_second

Step 5: Restore full_ratio after rebalance

# Once OSDs are balanced (< 80% on all), restore default
ceph osd set-full-ratio 0.95

Backfill Blocking Issues

backfill_toofull

If a target OSD is too full to accept backfill data, PGs stay in backfill_toofull state. Fix: lower the reweight of the full OSD further, or raise backfillfull_ratio:

ceph osd set-backfillfull-ratio 0.97

injectargs not sticking

ceph tell osd.N injectargs may not reliably set all parameters. osd_recovery_sleep in particular gets reset to 0. Use ceph config set osd for persistent values, and verify with ceph tell osd.N config get <param>.

CRITICAL: ceph config set osd osd_max_backfills does NOT apply at runtime

Setting ceph config set osd osd_max_backfills 3 updates the MON config store, but already-running OSD daemons keep using their old value (1). The change only takes effect on next OSD restart. Two ways to apply at runtime:

Method A — ceph tell osd.N injectargs (runtime, no restart):

# Apply to each HDD OSD individually
ceph tell osd.7 injectargs "--osd_max_backfills 3"
ceph tell osd.1 injectargs "--osd_max_backfills 3"
ceph tell osd.10 injectargs "--osd_max_backfills 3"
# For OSDs on other nodes, SSH there first or use ceph tell remotely

# ⚠️ Output is CONFUSING — prints all related config vars with empty values like:
# osd_max_backfills = '' osd_recovery_max_active = '' ...
# But the change DID take effect. Verify with:
ceph daemon osd.7 config get osd_max_backfills
# Should return: {"osd_max_backfills": "3"}

Method B — Restart OSD daemon (pick up MON config):

# ⚠️ MUST restart on the correct node! Use ceph osd find to locate:
ceph osd find 7
# Returns JSON with "host": "proxmox7" and IP

# Then SSH to THAT node and restart:
ssh proxmox7 "systemctl restart ceph-osd@7"
# ❌ WRONG: restarting on a node that doesn't host the OSD →
#   "OSD data directory /var/lib/ceph/osd/ceph-7 does not exist; bailing out"

Verification — check runtime value on each OSD:

ceph daemon osd.N config get osd_max_backfills
# Returns actual runtime value. If still "1", the config hasn't been applied.

osd_recovery_max_active cannot be effectively changed

ceph config set osd osd_recovery_max_active 3 accepts the value but OSDs report 0 (meaning auto/default). injectargs also shows empty. This appears to be a Ceph limitation — the parameter is managed internally. Do not waste time trying to force it.

Also raise nearfull_ratio alongside backfillfull_ratio

When an OSD is above 93%, it triggers nearfull warnings which add additional pool-level flags (nearfull on pools). If you only raise backfillfull_ratio to 0.97 but leave nearfull_ratio at 0.93, pools still show warnings and recovery may be partially blocked. Raise both together:

ceph osd set-backfillfull-ratio 0.97
ceph osd set-nearfull-ratio 0.97
# Verify:
ceph osd dump | grep -E "nearfull_ratio|backfillfull_ratio"

Recovery speed progression

After applying all tweaks, recovery speed increases progressively as OSDs pick up new config:

Stage Recovery Speed Notes
Before intervention 0 MiB/s (stuck) backfillfull blocks all backfill
After ratio increase + reweight 22-23 MiB/s, 5 obj/s 1-2 PGs backfilling
After osd_max_backfills=3 via injectargs 34 MiB/s, 8 obj/s Still only 2-3 PGs simultaneous

Even with osd_max_backfills=3, expect only 2-3 PGs backfilling simultaneously — Ceph schedules conservatively. The throughput improvement comes from larger transfer windows, not more parallel PGs.

Recovery monitoring cron pattern

Set up a recurring check to track progress and get alerted when complete:

Schedule: every 30 minutes
Command: ssh to monitor node, run `ceph status`, report:
  - PGs remapped/backfilling count
  - % misplaced objects
  - OSD 7 utilization
  - Recovery speed
  - Alert if any OSD down or health ERR
  - Announce "COMPLETE" when 0 remapped PGs

Estimating Backfill Duration

Metric Formula Example
By objects misplaced_objects ÷ objs/sec 353,960 ÷ 8 = ~12h
By bytes misplaced_bytes ÷ bytes/sec 418 GB ÷ 32 MiB/s = ~3.6h
Realistic Take the higher estimate, factor in PG scheduling 4-8h

Only a few PGs backfill simultaneously (controlled by osd_max_backfills). 192 PGs in backfill_wait won't all start at once.


OSD Preparation on Proxmox Nodes — 2026-07-04

Hardware Assessment Pattern

Before adding HDD OSDs to a node, verify:

  1. Available SATA/SAS ports:

    lspci | grep -i "sas\|sata\|nvme\|raid\|hba"
    ls /sys/class/ata_port/
    dmesg | grep -i "ahci\|ata[0-9]"
    

    ⚠️ AHCI controllers may show N ports but only implement a subset. Check dmesg for "ports implemented (port mask 0xN)". An OptiPlex 3070 showed 5 SATA ports but only port 0x1 (1 port) was implemented.

  2. Solution for limited ports: Install a PCIe HBA card (LSI 9211-8i / 9207-8i in IT mode) to connect additional SATA/SAS drives.

  3. Existing OSD inventory:

    ceph osd tree | grep -A20 "host <hostname>"
    ceph-volume lvm list
    

WAL/DB Space Preparation

Identifying available flash space

# LVM VG free space on system disk
vgs pve
lvs pve

# Check if thin pool is unused (safe to remove)
pvesm list local-lvm    # if storage doesn't exist, thin pool is orphaned
lvdisplay pve/data      # check allocation %

Repurposing unused thin pool for WAL/DB

If pve/data thin pool exists but is 0% allocated and not registered as a PVE storage, it can be removed to free NVMe space:

# Remove unused thin pool
lvremove pve/data
lvremove pve/data_tmeta
lvremove pve/data_tdata
# Or simply: lvremove pve/data (removes associated meta)

# Now VG has ~141 GB free for WAL/DB volumes
# Create DB LVs for new OSDs
lvcreate -L 30G -n osd-db-<id> pve

Rule: No loopback files for DB/WAL

Per user policy: Ceph DB/WAL must use real partitions or LVs, never loopback files. Options:

  • Repurpose swap partition (if RAM is adequate)
  • Shrink root LV (risky)
  • Remove unused thin pools
  • Add a physical SSD/NVMe

Node Checklist for Adding OSDs

  1. ☐ Verify physical SATA/SAS ports available (or install HBA)
  2. ☐ Verify disks are SMART-clean (smartctl -a /dev/sdX)
  3. ☐ Prepare WAL/DB space on flash (real partition/LV, no loopback)
  4. ☐ Create DB LVs: lvcreate -L 30G -n osd-db-N pve
  5. ☐ Create OSDs: ceph-volume lvm create --data /dev/sdX --block.db pve/osd-db-N
  6. ☐ Verify OSD joins cluster: ceph osd tree
  7. ☐ Check rebalance: ceph -s
  8. ☐ Set osd_max_backfills back to 1 after rebalance completes