Files
hermes-skills/devops/proxmox-ve-administration/references/ceph-pg-osd-management-2026-07.md
T

126 lines
4.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Ceph PG Management & OSD Rebalancing
## PG Count Reduction (too many PGs per OSD)
### Symptom
`ceph health` reports: `too many PGs per OSD (259 > max 250)`
### Procedure
```bash
# 1. Disable autoscale on target pools (CRITICAL — autoscaler overwrites manual values)
ceph osd pool set <pool> pg_autoscale_mode off
# 2. Set pgp_num FIRST (accepts immediately)
ceph osd pool set <pool> pgp_num <new_value>
# 3. Set pg_num (may not take effect immediately)
ceph osd pool set <pool> pg_num <new_value>
```
### Pitfalls
- **Autoscaler overrides manual pg_num**: If `pg_autoscale_mode=on`, setting
`pg_num` appears to succeed but the value reverts. Must set `off` first.
- **pg_num won't reduce during recovery**: If PGs are not all `active+clean`
(e.g. OSD reweight recovery in progress), `pg_num` stays at the old value.
The `pg_num_target` is set (visible in `ceph osd pool ls detail`) and the
merge executes automatically once recovery completes.
- **`mon_max_pg_per_osd` does NOT hot-reload**: Setting it via
`ceph config set mon mon_max_pg_per_osd 400` updates the config but the
MONs continue using the old value. `ceph tell mon.* injectargs` also
failed to apply it. The health warning persists until either:
- The PG merge completes (actual PG count drops), or
- MONs are restarted (picks up new config value)
- **Reducing from 128→32 on near-empty pools is safe**: Pools with 0 B or
85 KiB of data (e.g. unused cephfs_data, cephfs_metadata) can be reduced
aggressively. 128→32 saves 576 PG instances across 3 replicas = significant.
### Choosing Target PG Count
| Pool stored data | Recommended PGs |
|-----------------|-----------------|
| 0 B 1 GiB | 32 |
| 1100 GiB | 64128 |
| 100 GiB1 TiB | 128256 |
| >1 TiB | 256512 |
Rule: each PG should hold ~100 GiB of data for optimal balance.
## OSD Reweight (Rebalancing Full OSDs)
### Symptom
One OSD at 85-88% utilization while others are at 30-40%. `VAR` >1.5 in
`ceph osd df`.
### Procedure
```bash
# Gradually reduce weight (default 1.0, lower = less data assigned)
ceph osd reweight osd.<N> 0.65
# Monitor progress
ceph osd df | grep osd.<N>
# Watch %USE decrease over time as data migrates away
```
### Pitfalls
- **Competing I/O defeats reweight**: If new writes target the same OSD's
pool (e.g. RBD copy writing to `media_ec` HDD pool while reweighting
an HDD OSD), the OSD fills faster than reweight empties it. Observed:
osd.7 went from 80% → 85% during a 341 GiB RBD copy despite reweight=0.65.
The reweight only wins after the concurrent write workload stops.
- **Recovery is slow under I/O contention**: With concurrent fsck + RBD
copy + reweight recovery all hitting HDD OSDs, expect `BLUESTORE_SLOW_OP_ALERT`
and recovery rates of 20-90 MiB/s instead of peak. This is temporary —
speeds recover after the competing workload finishes.
- **Reweight is gradual by design**: CRUST gradually moves PGs away from
the reweighted OSD. Don't expect instant results — monitor over
10-60 minutes depending on data volume.
## Concurrent I/O Awareness (Critical)
Multiple Ceph-intensive operations running simultaneously cause cascading
slowness:
| Operation | I/O Profile |
|-----------|------------|
| OSD reweight recovery | Random read + sequential write across OSDs |
| RBD copy (--data-pool) | Sequential read source + EC write target |
| Seafile fsck --repair | Random read across HDD pool |
| PG merge (post-recovery) | Metadata-heavy, blocks on clean state |
**Recommendation**: Run these sequentially when possible. If concurrent,
expect 3-5x slowdown on all operations and temporary HEALTH_WARN from
slow ops. All operations complete eventually — patience is the fix.
## Quick Reference: This Cluster's Pools
| Pool | PGs | Size | Data Class | Notes |
|------|-----|------|------------|-------|
| cephfs_data | 128→32 | 3 | HDD | 0 B stored — massively over-provisioned |
| cephfs_metadata | 128→32 | 3 | HDD | 85 KiB stored — massively over-provisioned |
| vm_disks | 128 | 3 | SSD | VM root disks |
| rbd | 32 | 3 | SSD | Small images |
| hdd_disk | 128 | 3 | HDD | CT root disks |
| tm_disks | 128 | 2 | HDD | Legacy, size=2 (NO fault tolerance) |
| media_ec | 128 | 5 (EC4+1) | HDD | EC data pool |
| media_meta | 32 | 3 | HDD | EC metadata pool |
| .mgr | 1 | 3 | HDD | Manager daemon |
## Session Log — 2026-07-06
Reduced cephfs_data and cephfs_metadata PGs 128→32 (pg_num_target set,
merge pending recovery completion). Reweighted osd.7 from 0.80→0.65 to
rebalance from 88% utilization. Both operations competed with a 341 GiB
RBD copy (tm_disks→media_ec) and Seafile fsck, causing slow ops and
extended timelines. PG warning persisted until recovery completed.