Files
hermes-skills/devops/proxmox-ve-administration/references/ceph-pg-reduction-and-osd-rebalance-2026-07.md
T

159 lines
5.4 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Ceph PG Reduction & OSD Rebalancing — 2026-07-06
## Context
Cluster had `TOO_MANY_PGS` warning (260 > 250 max). Two CephFS pools with
128 PGs each held almost no data (cephfs_data: 0 B, cephfs_metadata: 85 KiB).
osd.7 (982 GB HDD) was at 88% utilization (VAR 1.95) while osd.6/osd.8
(2.8 TB each) sat at 33%. Both issues addressed in one session.
## PG Reduction
### Prerequisites
1. **Disable autoscale before manual PG changes** — If `pg_autoscale_mode` is
`on`, the autoscaler reverts manual `pg_num`/`pgp_num` changes immediately:
```bash
ceph osd pool set <pool> pg_autoscale_mode off
```
2. **All PGs must be `active+clean`** — PG merge (reducing pg_num) won't start
while recovery/backfill is in progress. Check with:
```bash
ceph pg dump pgs_brief | tail -n +2 | awk '{print $2}' | grep -v 'active+clean'
# Should return nothing
```
### Reducing PGs
Set `pgp_num` first, then `pg_num`:
```bash
ceph osd pool set <pool> pgp_num <new>
ceph osd pool set <pool> pg_num <new>
```
⚠️ **pg_num may not decrease immediately.** Ceph sets `pg_num_target` and
`pgp_num_target` in the pool metadata and defers the actual merge until all
PGs are clean and no recovery is in progress. Verify the target is set:
```bash
ceph osd pool ls detail | grep <pool>
# Look for: pg_num 128 pgp_num 128 pg_num_target 32 pgp_num_target 32
```
The merge happens automatically once conditions are met. No further action needed.
### Which Pools to Reduce
Pools with tiny data but high PG counts are prime candidates:
| Pool | PGs | Data | Action |
|------|-----|------|--------|
| cephfs_data | 128 | 0 B | → 32 (96 PGs × 3 replicas = 288 instances saved) |
| cephfs_metadata | 128 | 85 KiB | → 32 (same) |
Reducing both from 128→32 saves 576 PG instances across 10 OSDs (~58/OSD),
bringing the average from ~260 to ~205 — well under the 250 limit.
### mon_max_pg_per_osd Workaround (Temporary)
If you need the warning gone immediately while waiting for PG merge:
```bash
ceph config set mon mon_max_pg_per_osd 300
```
⚠️ This config may not take effect immediately — the Mons may need a restart
to pick up the new value. `ceph tell mon.* injectargs` reported the value as
"not observed, change may require restart." This is a band-aid; the real fix
is reducing PG counts.
## OSD Reweight for Rebalancing
### Identifying Imbalanced OSDs
```bash
ceph osd df
# Look for: %USE > 85% and VAR > 1.5 (significantly above average)
```
osd.7 at 88% (VAR 1.95) vs osd.6/osd.8 at 33% (VAR 0.72) — clear imbalance.
### Reweighting
```bash
ceph osd reweight osd.7 0.65
# Was 0.80 (already reduced from 1.0 in a prior session)
# Lower reweight → CRUSH assigns fewer PGs → data migrates away
```
Recovery speed: ~88 MiB/s, 22 obj/s. At ~50 GiB to move, estimated 10-15 min.
### Monitoring Rebalance
```bash
# Overall recovery progress
ceph status | grep -A2 'recovery:'
# OSD utilization trend
ceph osd df | grep osd.7
# Watch %USE drop and VAR approach 1.0
# Misplaced objects
ceph status | grep misplaced
```
### Post-Rebalance
Once osd.7 drops below ~75% and VAR approaches 1.0, consider raising the
reweight back slightly to prevent underutilization:
```bash
ceph osd reweight osd.7 0.75 # moderate value between 0.65 and 0.80
```
## Interaction Between PG Merge and Recovery
PG merge (pg_num decrease) is deferred while recovery/backfill is in progress.
This creates a dependency chain:
1. OSD reweight triggers data migration (recovery)
2. Recovery must complete (all PGs `active+clean`)
3. Only then does PG merge execute
4. After merge, PG-per-OSD count drops → `TOO_MANY_PGS` warning clears
Timeline: recovery (~15 min) → PG merge (~5 min) → health check clears.
## Pitfalls
1. **Autoscaler reverts manual changes** — Always `pg_autoscale_mode off` before
setting `pg_num`/`pgp_num` manually. Otherwise the autoscaler sees the
reduction and immediately resets to its calculated optimum.
2. **pg_num decrease is deferred, not instant** — Unlike increasing pg_num
(which splits PGs immediately), decreasing merges PGs and requires all PGs
to be clean. Don't expect immediate results during active recovery.
3. **mon_max_pg_per_osd may need Mon restart** — `ceph config set mon
mon_max_pg_per_osd 300` updates the config store but Mons may not pick it
up without a restart. Don't rely on this as a permanent fix.
4. **Reweight during active fsck amplifies slow ops** — If a heavy I/O workload
(e.g., Seafile `seaf-fsck --repair` reading from `hdd_disk` pool) is running
concurrently, the OSD reweight recovery adds more I/O pressure. Expect
`BLUESTORE_SLOW_OP_ALERT` warnings until both finish.
5. **Reweight can INCREASE OSD usage when concurrent writes target the same
OSD** — Counterintuitive but observed: osd.7 was reweighted to 0.65 to drain
data (88% → 80%), but then rose back to 85% during an RBD copy to `media_ec`.
Cause: the RBD copy wrote NEW EC chunks to all HDD OSDs including osd.7,
and the new write rate exceeded the drain rate from reweight. The reweight
reduces CRUSH placement probability for NEW PGs, but existing active writes
to already-placed PGs continue. Net effect: usage goes UP if new writes
outpace the recovery drain. Resolution: either pause the write workload
until reweight recovery completes, or accept that the reweight will only
take effect after the write workload ends. Do not interpret rising usage
as a failed reweight — it's a transient interaction.