Files
hermes-skills/devops/proxmox-ve-administration/references/ceph-crush-weight-optimization-2026-07.md
Debian 01bd921ced feat: add home-assistant-dashboard-conventions skill + update multiple skills
- New: smart-home/home-assistant-dashboard-conventions (Mushroom cards, view tabs, no Bubble Cards)
- Updated: rke2, ceph, galera, proxmox, brainstorming, compound-learning, 1password-cli, smart-home-automation skills
- New references: ceph-cluster-administration, docker-volume-forensics, ceph-crush-weight, ceph-ec-mixed-size
2026-07-14 18:35:16 +00:00

5.2 KiB
Raw Permalink Blame History

Ceph CRUSH Weight Optimization for Mixed OSD Sizes — 2026-07-13

Problem

When a cluster has OSDs of different sizes (1 TB and 3 TB), CRUSH weights should be proportional to device size. But in practice, operators sometimes manually lower CRUSH weights on small/full OSDs to reduce incoming data — which creates a vicious cycle:

  1. Small OSD fills up → operator lowers CRUSH weight
  2. Lower weight → fewer new PGs assigned → old data stays
  3. OSD stays full → operator lowers weight further
  4. Large OSDs (especially new ones) remain underutilized

Diagnosis

# Show all HDD OSDs with weight, reweight, utilization, PG count
ceph osd df tree | grep hdd

# Key columns to compare:
# WEIGHT  — CRUSH weight (should ≈ size_in_TB)
# REWEIGHT — runtime reweight (should be 1.0 normally)
# %USE    — utilization (should be roughly equal across OSDs)
# PGS     — PG count (should be proportional to weight)

Symptoms of misaligned weights

Signal Meaning
WEIGHT << SIZE (e.g. 0.3 for 1 TB) Artificially suppressed weight
%USE varies wildly (27% vs 94%) Data not distributed proportionally
PGS count disproportional (150 vs 450) CRUSH assigns PGs by weight, not size
New large OSD barely fills (2% after hours) Weight correct but recovery slow

Solution

Step 1: Identify correct CRUSH weights

Set CRUSH weight = device size in TiB (Ceph convention):

Device Size Correct CRUSH Weight
1 TB (982 GiB) ~0.96
2 TB ~1.82
3 TB (2.8 TiB) ~2.73
3.6 TiB ~3.64
# Check current weights
ceph osd tree | grep hdd
# Compare WEIGHT column to SIZE in: ceph osd df

Step 2: Reset artificial weights (after recovery!)

⚠️ Timing matters: If a rebalance is already running, changing CRUSH weights triggers ADDITIONAL remapping. Either:

  • Wait for current rebalance to finish, THEN adjust weights, OR
  • Adjust weights now and accept a longer combined rebalance
# Reset CRUSH weight to match device size
ceph osd crush reweight osd.7 0.96   # was 0.50, should be ~0.96 for 1 TB
ceph osd crush reweight osd.10 0.96  # was 0.30, should be ~0.96 for 1 TB

# Also reset runtime reweight to 1.0 (if it was lowered for emergency drain)
ceph osd reweight 7 1.0
ceph osd reweight 10 1.0

Step 3: Let the upmap balancer handle fine-tuning

The upmap balancer module optimizes PG placement without full remapping:

# Check balancer status
ceph balancer status

# If "no_optimization_needed": false, the balancer is actively working
# If "optimize_result" mentions "too many objects misplaced": wait for
# current recovery to drop below 5% misplaced, then balancer kicks in

The balancer cannot run while >5% objects are misplaced (it skips to avoid compounding recovery load). Once the initial rebalance from weight changes completes, the balancer will fine-tune PG distribution automatically.

Step 4: Optional — reweight-by-utilization for dynamic balancing

# Automatically reweight OSDs based on utilization (temporary reweights)
ceph osd reweight-by-utilization
# Or with custom threshold:
ceph osd test-reweight-by-utilization 120  # dry run, shows what would change
ceph osd reweight-by-utilization 120       # apply (120 = 1.2x average = overfull)

⚠️ Use cautiously — this changes runtime reweights, not CRUSH weights. The effect is temporary and can interact with ongoing recovery.

Mixed-Size Cluster Example

Real-world cluster with 1 TB and 3 TB HDDs:

OSD Size Old Weight New Weight Old %Used Expected %Used
osd.1 3.6 TiB 3.64 3.64 (OK) 33% ~33%
osd.6 2.8 TiB 2.76 2.76 (OK) 27% ~27%
osd.8 2.8 TiB 2.73 2.73 (OK) 27% ~27%
osd.11 2.8 TiB 2.76 2.76 (OK) 2% ~27% (filling)
osd.7 982 GiB 0.50 0.96 94% ~33%
osd.10 982 GiB 0.30 0.96 57% ~33%

After resetting osd.7 and osd.10 to their true weights, CRUSH will assign them proportionally more PGs, and data will distribute evenly across all HDDs regardless of size. The 3 TB OSDs will naturally hold ~3× the data of 1 TB OSDs.

Pitfalls

Don't lower weight to "protect" a full OSD

Lowering CRUSH weight on a full OSD prevents new PGs from landing there, but does NOT move existing data away. The OSD stays full. Meanwhile, the reduced weight means the OSD contributes less to the cluster's apparent capacity, causing nearfull/backfillfull warnings on remaining OSDs.

Instead: use ceph osd reweight (runtime, temporary) for emergency drain, then fix the root cause (add capacity, redistribute, or accept the OSD size and set proper CRUSH weight).

Balancer won't help during active recovery

The upmap balancer skips when >5% objects are misplaced. Don't expect it to optimize placement while a major rebalance is in progress. Wait for recovery to complete, then let the balancer run.

CRUSH weight changes trigger remapping

Every ceph osd crush reweight causes CRUSH to recompute PG placements. On a cluster with existing data, this triggers backfill. Schedule weight changes during maintenance windows or combine with planned capacity additions.