Files
hermes-skills/devops/proxmox-ve-administration/references/ceph-crush-weight-optimization-2026-07.md
T
Debian 01bd921ced feat: add home-assistant-dashboard-conventions skill + update multiple skills
- New: smart-home/home-assistant-dashboard-conventions (Mushroom cards, view tabs, no Bubble Cards)
- Updated: rke2, ceph, galera, proxmox, brainstorming, compound-learning, 1password-cli, smart-home-automation skills
- New references: ceph-cluster-administration, docker-volume-forensics, ceph-crush-weight, ceph-ec-mixed-size
2026-07-14 18:35:16 +00:00

145 lines
5.2 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Ceph CRUSH Weight Optimization for Mixed OSD Sizes — 2026-07-13
## Problem
When a cluster has OSDs of different sizes (1 TB and 3 TB), CRUSH weights
should be proportional to device size. But in practice, operators sometimes
manually lower CRUSH weights on small/full OSDs to reduce incoming data —
which creates a vicious cycle:
1. Small OSD fills up → operator lowers CRUSH weight
2. Lower weight → fewer new PGs assigned → old data stays
3. OSD stays full → operator lowers weight further
4. Large OSDs (especially new ones) remain underutilized
## Diagnosis
```bash
# Show all HDD OSDs with weight, reweight, utilization, PG count
ceph osd df tree | grep hdd
# Key columns to compare:
# WEIGHT — CRUSH weight (should ≈ size_in_TB)
# REWEIGHT — runtime reweight (should be 1.0 normally)
# %USE — utilization (should be roughly equal across OSDs)
# PGS — PG count (should be proportional to weight)
```
### Symptoms of misaligned weights
| Signal | Meaning |
|--------|---------|
| WEIGHT << SIZE (e.g. 0.3 for 1 TB) | Artificially suppressed weight |
| %USE varies wildly (27% vs 94%) | Data not distributed proportionally |
| PGS count disproportional (150 vs 450) | CRUSH assigns PGs by weight, not size |
| New large OSD barely fills (2% after hours) | Weight correct but recovery slow |
## Solution
### Step 1: Identify correct CRUSH weights
Set CRUSH weight = device size in TiB (Ceph convention):
| Device Size | Correct CRUSH Weight |
|-------------|---------------------|
| 1 TB (982 GiB) | ~0.96 |
| 2 TB | ~1.82 |
| 3 TB (2.8 TiB) | ~2.73 |
| 3.6 TiB | ~3.64 |
```bash
# Check current weights
ceph osd tree | grep hdd
# Compare WEIGHT column to SIZE in: ceph osd df
```
### Step 2: Reset artificial weights (after recovery!)
⚠️ **Timing matters**: If a rebalance is already running, changing CRUSH
weights triggers ADDITIONAL remapping. Either:
- Wait for current rebalance to finish, THEN adjust weights, OR
- Adjust weights now and accept a longer combined rebalance
```bash
# Reset CRUSH weight to match device size
ceph osd crush reweight osd.7 0.96 # was 0.50, should be ~0.96 for 1 TB
ceph osd crush reweight osd.10 0.96 # was 0.30, should be ~0.96 for 1 TB
# Also reset runtime reweight to 1.0 (if it was lowered for emergency drain)
ceph osd reweight 7 1.0
ceph osd reweight 10 1.0
```
### Step 3: Let the upmap balancer handle fine-tuning
The `upmap` balancer module optimizes PG placement without full remapping:
```bash
# Check balancer status
ceph balancer status
# If "no_optimization_needed": false, the balancer is actively working
# If "optimize_result" mentions "too many objects misplaced": wait for
# current recovery to drop below 5% misplaced, then balancer kicks in
```
The balancer cannot run while >5% objects are misplaced (it skips to avoid
compounding recovery load). Once the initial rebalance from weight changes
completes, the balancer will fine-tune PG distribution automatically.
### Step 4: Optional — `reweight-by-utilization` for dynamic balancing
```bash
# Automatically reweight OSDs based on utilization (temporary reweights)
ceph osd reweight-by-utilization
# Or with custom threshold:
ceph osd test-reweight-by-utilization 120 # dry run, shows what would change
ceph osd reweight-by-utilization 120 # apply (120 = 1.2x average = overfull)
```
⚠️ Use cautiously — this changes runtime reweights, not CRUSH weights.
The effect is temporary and can interact with ongoing recovery.
## Mixed-Size Cluster Example
Real-world cluster with 1 TB and 3 TB HDDs:
| OSD | Size | Old Weight | New Weight | Old %Used | Expected %Used |
|-----|------|-----------|------------|-----------|----------------|
| osd.1 | 3.6 TiB | 3.64 | 3.64 (OK) | 33% | ~33% |
| osd.6 | 2.8 TiB | 2.76 | 2.76 (OK) | 27% | ~27% |
| osd.8 | 2.8 TiB | 2.73 | 2.73 (OK) | 27% | ~27% |
| osd.11 | 2.8 TiB | 2.76 | 2.76 (OK) | 2% | ~27% (filling) |
| osd.7 | 982 GiB | **0.50** | **0.96** | 94% | ~33% |
| osd.10 | 982 GiB | **0.30** | **0.96** | 57% | ~33% |
After resetting osd.7 and osd.10 to their true weights, CRUSH will assign
them proportionally more PGs, and data will distribute evenly across all
HDDs regardless of size. The 3 TB OSDs will naturally hold ~3× the data
of 1 TB OSDs.
## Pitfalls
### Don't lower weight to "protect" a full OSD
Lowering CRUSH weight on a full OSD prevents new PGs from landing there,
but does NOT move existing data away. The OSD stays full. Meanwhile, the
reduced weight means the OSD contributes less to the cluster's apparent
capacity, causing `nearfull`/`backfillfull` warnings on remaining OSDs.
Instead: use `ceph osd reweight` (runtime, temporary) for emergency drain,
then fix the root cause (add capacity, redistribute, or accept the OSD
size and set proper CRUSH weight).
### Balancer won't help during active recovery
The `upmap` balancer skips when >5% objects are misplaced. Don't expect
it to optimize placement while a major rebalance is in progress. Wait for
recovery to complete, then let the balancer run.
### CRUSH weight changes trigger remapping
Every `ceph osd crush reweight` causes CRUSH to recompute PG placements.
On a cluster with existing data, this triggers backfill. Schedule weight
changes during maintenance windows or combine with planned capacity additions.