4.6 KiB
Ceph PG Management & OSD Rebalancing
PG Count Reduction (too many PGs per OSD)
Symptom
ceph health reports: too many PGs per OSD (259 > max 250)
Procedure
# 1. Disable autoscale on target pools (CRITICAL — autoscaler overwrites manual values)
ceph osd pool set <pool> pg_autoscale_mode off
# 2. Set pgp_num FIRST (accepts immediately)
ceph osd pool set <pool> pgp_num <new_value>
# 3. Set pg_num (may not take effect immediately)
ceph osd pool set <pool> pg_num <new_value>
Pitfalls
-
Autoscaler overrides manual pg_num: If
pg_autoscale_mode=on, settingpg_numappears to succeed but the value reverts. Must setofffirst. -
pg_num won't reduce during recovery: If PGs are not all
active+clean(e.g. OSD reweight recovery in progress),pg_numstays at the old value. Thepg_num_targetis set (visible inceph osd pool ls detail) and the merge executes automatically once recovery completes. -
mon_max_pg_per_osddoes NOT hot-reload: Setting it viaceph config set mon mon_max_pg_per_osd 400updates the config but the MONs continue using the old value.ceph tell mon.* injectargsalso failed to apply it. The health warning persists until either:- The PG merge completes (actual PG count drops), or
- MONs are restarted (picks up new config value)
-
Reducing from 128→32 on near-empty pools is safe: Pools with 0 B or 85 KiB of data (e.g. unused cephfs_data, cephfs_metadata) can be reduced aggressively. 128→32 saves 576 PG instances across 3 replicas = significant.
Choosing Target PG Count
| Pool stored data | Recommended PGs |
|---|---|
| 0 B – 1 GiB | 32 |
| 1–100 GiB | 64–128 |
| 100 GiB–1 TiB | 128–256 |
| >1 TiB | 256–512 |
Rule: each PG should hold ~100 GiB of data for optimal balance.
OSD Reweight (Rebalancing Full OSDs)
Symptom
One OSD at 85-88% utilization while others are at 30-40%. VAR >1.5 in
ceph osd df.
Procedure
# Gradually reduce weight (default 1.0, lower = less data assigned)
ceph osd reweight osd.<N> 0.65
# Monitor progress
ceph osd df | grep osd.<N>
# Watch %USE decrease over time as data migrates away
Pitfalls
-
Competing I/O defeats reweight: If new writes target the same OSD's pool (e.g. RBD copy writing to
media_ecHDD pool while reweighting an HDD OSD), the OSD fills faster than reweight empties it. Observed: osd.7 went from 80% → 85% during a 341 GiB RBD copy despite reweight=0.65. The reweight only wins after the concurrent write workload stops. -
Recovery is slow under I/O contention: With concurrent fsck + RBD copy + reweight recovery all hitting HDD OSDs, expect
BLUESTORE_SLOW_OP_ALERTand recovery rates of 20-90 MiB/s instead of peak. This is temporary — speeds recover after the competing workload finishes. -
Reweight is gradual by design: CRUST gradually moves PGs away from the reweighted OSD. Don't expect instant results — monitor over 10-60 minutes depending on data volume.
Concurrent I/O Awareness (Critical)
Multiple Ceph-intensive operations running simultaneously cause cascading slowness:
| Operation | I/O Profile |
|---|---|
| OSD reweight recovery | Random read + sequential write across OSDs |
| RBD copy (--data-pool) | Sequential read source + EC write target |
| Seafile fsck --repair | Random read across HDD pool |
| PG merge (post-recovery) | Metadata-heavy, blocks on clean state |
Recommendation: Run these sequentially when possible. If concurrent, expect 3-5x slowdown on all operations and temporary HEALTH_WARN from slow ops. All operations complete eventually — patience is the fix.
Quick Reference: This Cluster's Pools
| Pool | PGs | Size | Data Class | Notes |
|---|---|---|---|---|
| cephfs_data | 128→32 | 3 | HDD | 0 B stored — massively over-provisioned |
| cephfs_metadata | 128→32 | 3 | HDD | 85 KiB stored — massively over-provisioned |
| vm_disks | 128 | 3 | SSD | VM root disks |
| rbd | 32 | 3 | SSD | Small images |
| hdd_disk | 128 | 3 | HDD | CT root disks |
| tm_disks | 128 | 2 | HDD | Legacy, size=2 (NO fault tolerance) |
| media_ec | 128 | 5 (EC4+1) | HDD | EC data pool |
| media_meta | 32 | 3 | HDD | EC metadata pool |
| .mgr | 1 | 3 | HDD | Manager daemon |
Session Log — 2026-07-06
Reduced cephfs_data and cephfs_metadata PGs 128→32 (pg_num_target set, merge pending recovery completion). Reweighted osd.7 from 0.80→0.65 to rebalance from 88% utilization. Both operations competed with a 341 GiB RBD copy (tm_disks→media_ec) and Seafile fsck, causing slow ops and extended timelines. PG warning persisted until recovery completed.