- New: smart-home/home-assistant-dashboard-conventions (Mushroom cards, view tabs, no Bubble Cards) - Updated: rke2, ceph, galera, proxmox, brainstorming, compound-learning, 1password-cli, smart-home-automation skills - New references: ceph-cluster-administration, docker-volume-forensics, ceph-crush-weight, ceph-ec-mixed-size
5.2 KiB
Ceph CRUSH Weight Optimization for Mixed OSD Sizes — 2026-07-13
Problem
When a cluster has OSDs of different sizes (1 TB and 3 TB), CRUSH weights should be proportional to device size. But in practice, operators sometimes manually lower CRUSH weights on small/full OSDs to reduce incoming data — which creates a vicious cycle:
- Small OSD fills up → operator lowers CRUSH weight
- Lower weight → fewer new PGs assigned → old data stays
- OSD stays full → operator lowers weight further
- Large OSDs (especially new ones) remain underutilized
Diagnosis
# Show all HDD OSDs with weight, reweight, utilization, PG count
ceph osd df tree | grep hdd
# Key columns to compare:
# WEIGHT — CRUSH weight (should ≈ size_in_TB)
# REWEIGHT — runtime reweight (should be 1.0 normally)
# %USE — utilization (should be roughly equal across OSDs)
# PGS — PG count (should be proportional to weight)
Symptoms of misaligned weights
| Signal | Meaning |
|---|---|
| WEIGHT << SIZE (e.g. 0.3 for 1 TB) | Artificially suppressed weight |
| %USE varies wildly (27% vs 94%) | Data not distributed proportionally |
| PGS count disproportional (150 vs 450) | CRUSH assigns PGs by weight, not size |
| New large OSD barely fills (2% after hours) | Weight correct but recovery slow |
Solution
Step 1: Identify correct CRUSH weights
Set CRUSH weight = device size in TiB (Ceph convention):
| Device Size | Correct CRUSH Weight |
|---|---|
| 1 TB (982 GiB) | ~0.96 |
| 2 TB | ~1.82 |
| 3 TB (2.8 TiB) | ~2.73 |
| 3.6 TiB | ~3.64 |
# Check current weights
ceph osd tree | grep hdd
# Compare WEIGHT column to SIZE in: ceph osd df
Step 2: Reset artificial weights (after recovery!)
⚠️ Timing matters: If a rebalance is already running, changing CRUSH weights triggers ADDITIONAL remapping. Either:
- Wait for current rebalance to finish, THEN adjust weights, OR
- Adjust weights now and accept a longer combined rebalance
# Reset CRUSH weight to match device size
ceph osd crush reweight osd.7 0.96 # was 0.50, should be ~0.96 for 1 TB
ceph osd crush reweight osd.10 0.96 # was 0.30, should be ~0.96 for 1 TB
# Also reset runtime reweight to 1.0 (if it was lowered for emergency drain)
ceph osd reweight 7 1.0
ceph osd reweight 10 1.0
Step 3: Let the upmap balancer handle fine-tuning
The upmap balancer module optimizes PG placement without full remapping:
# Check balancer status
ceph balancer status
# If "no_optimization_needed": false, the balancer is actively working
# If "optimize_result" mentions "too many objects misplaced": wait for
# current recovery to drop below 5% misplaced, then balancer kicks in
The balancer cannot run while >5% objects are misplaced (it skips to avoid compounding recovery load). Once the initial rebalance from weight changes completes, the balancer will fine-tune PG distribution automatically.
Step 4: Optional — reweight-by-utilization for dynamic balancing
# Automatically reweight OSDs based on utilization (temporary reweights)
ceph osd reweight-by-utilization
# Or with custom threshold:
ceph osd test-reweight-by-utilization 120 # dry run, shows what would change
ceph osd reweight-by-utilization 120 # apply (120 = 1.2x average = overfull)
⚠️ Use cautiously — this changes runtime reweights, not CRUSH weights. The effect is temporary and can interact with ongoing recovery.
Mixed-Size Cluster Example
Real-world cluster with 1 TB and 3 TB HDDs:
| OSD | Size | Old Weight | New Weight | Old %Used | Expected %Used |
|---|---|---|---|---|---|
| osd.1 | 3.6 TiB | 3.64 | 3.64 (OK) | 33% | ~33% |
| osd.6 | 2.8 TiB | 2.76 | 2.76 (OK) | 27% | ~27% |
| osd.8 | 2.8 TiB | 2.73 | 2.73 (OK) | 27% | ~27% |
| osd.11 | 2.8 TiB | 2.76 | 2.76 (OK) | 2% | ~27% (filling) |
| osd.7 | 982 GiB | 0.50 | 0.96 | 94% | ~33% |
| osd.10 | 982 GiB | 0.30 | 0.96 | 57% | ~33% |
After resetting osd.7 and osd.10 to their true weights, CRUSH will assign them proportionally more PGs, and data will distribute evenly across all HDDs regardless of size. The 3 TB OSDs will naturally hold ~3× the data of 1 TB OSDs.
Pitfalls
Don't lower weight to "protect" a full OSD
Lowering CRUSH weight on a full OSD prevents new PGs from landing there,
but does NOT move existing data away. The OSD stays full. Meanwhile, the
reduced weight means the OSD contributes less to the cluster's apparent
capacity, causing nearfull/backfillfull warnings on remaining OSDs.
Instead: use ceph osd reweight (runtime, temporary) for emergency drain,
then fix the root cause (add capacity, redistribute, or accept the OSD
size and set proper CRUSH weight).
Balancer won't help during active recovery
The upmap balancer skips when >5% objects are misplaced. Don't expect
it to optimize placement while a major rebalance is in progress. Wait for
recovery to complete, then let the balancer run.
CRUSH weight changes trigger remapping
Every ceph osd crush reweight causes CRUSH to recompute PG placements.
On a cluster with existing data, this triggers backfill. Schedule weight
changes during maintenance windows or combine with planned capacity additions.