feat: add home-assistant-dashboard-conventions skill + update multiple skills
- New: smart-home/home-assistant-dashboard-conventions (Mushroom cards, view tabs, no Bubble Cards) - Updated: rke2, ceph, galera, proxmox, brainstorming, compound-learning, 1password-cli, smart-home-automation skills - New references: ceph-cluster-administration, docker-volume-forensics, ceph-crush-weight, ceph-ec-mixed-size
This commit is contained in:
@@ -0,0 +1,546 @@
|
||||
---
|
||||
name: ceph-cluster-administration
|
||||
description: "Class-level skill for Ceph cluster administration on Proxmox VE. Covers OSD lifecycle (create/destroy/WAL-DB), CRUSH topology (device classes, roots, rules, weights), EC pools (profiles, mixed disk sizes, migration), recovery & rebalancing (tuning, ratios, monitoring), RBD management (orphan cleanup, migration, inspection), pool management (PG count, full recovery, autoscale), and health diagnosis."
|
||||
version: 1.0.0
|
||||
author: Hermes Agent
|
||||
tags: [ceph, proxmox, bluestore, crush, erasure-coding, rbd, recovery, osd, storage]
|
||||
metadata:
|
||||
hermes:
|
||||
tags: [ceph, proxmox, bluestore, crush, erasure-coding, rbd, recovery, osd, storage]
|
||||
related_skills: [proxmox-ve-administration, rke2-cluster-administration]
|
||||
---
|
||||
|
||||
# Ceph Cluster Administration
|
||||
|
||||
## Scope
|
||||
|
||||
**This skill covers the Homelab (self-hosted) Ceph cluster.** Not noris infrastructure. At noris only remote PBS + S3 Storage exist — neither involves Ceph. Do not confuse the two environments.
|
||||
|
||||
## Environment
|
||||
|
||||
| Component | Detail |
|
||||
|-----------|--------|
|
||||
| Environment | **Homelab (self-hosted)** — NOT noris |
|
||||
| PVE Master | 10.0.20.10 (SSH: `root@10.0.20.10`, key `~/.ssh/id_ed25519_proxmox`) |
|
||||
| PVE Nodes (SSH) | All PVE hosts reachable via `10.0.20.x` (storage VLAN), NOT `10.0.30.x` (management VLAN — SSH key rejected). px1=.10, px2=.20, px3=.30, px4=.40, px5=.50, px6=.60, px7(n5pro)=.70 or .91 |
|
||||
| OSDs | 11 total (6 HDD + 5 SSD) |
|
||||
| HDD mix | 4× ~3 TB (osd.1, 6, 8, 11) + 2× ~1 TB (osd.7, 10) |
|
||||
| EC Profile | `ec41-hdd` (k=4, m=1) → migrating to `ec31-hdd-large` (k=3, m=1) |
|
||||
| Pools | `media_ec` (EC), `hdd_disk` (Repl 3), `rbd` (Repl 3), `tm_disks` (Repl 2), `media_meta` (Repl 3), `vm_disks` (Repl 3) |
|
||||
| Ratios | nearfull=0.94, backfillfull=0.96, full=0.97 (failsafe hardcoded 0.97) |
|
||||
| Monitoring | Grafana CT141 (DS UID=PBFA97CFB590B2093), Prometheus, Alertmanager→Telegram |
|
||||
| Ceph Mon Network | `10.0.20.x` (storage VLAN) — NOT `10.0.30.x` (management) |
|
||||
|
||||
All Ceph commands run from PVE master (**10.0.20.10** — storage network) unless noted. Non-Ceph nodes (e.g. n5pro) need `ceph.conf` + admin keyring copied and `rbd`/`ceph` commands run from a monitor node.
|
||||
|
||||
### SSH Access: 10.0.20.x Works, 10.0.30.x Does NOT
|
||||
|
||||
PVE hosts have **two networks**: storage (10.0.20.x) and management (10.0.30.x). SSH with `~/.ssh/id_ed25519_proxmox` works on **10.0.20.x** but is rejected on **10.0.30.x** (key not in authorized_keys). Always use 10.0.20.x IPs for PVE host SSH:
|
||||
|
||||
```bash
|
||||
# CORRECT — storage network
|
||||
ssh -i ~/.ssh/id_ed25519_proxmox root@10.0.20.10 "ceph -s"
|
||||
|
||||
# WRONG — management network (key rejected)
|
||||
ssh -i ~/.ssh/id_ed25519_proxmox root@10.0.30.10 "ceph -s" # Permission denied
|
||||
```
|
||||
|
||||
PVE host IP mapping (storage network): px1=.10, px2=.20, px3=.30, px4=.40, px5=.50, px6=.60, px7=.70, n5pro=.91.
|
||||
|
||||
### Fallback: Ceph CLI via K8s ceph-csi Pods
|
||||
|
||||
When PVE SSH is unavailable, `ceph` commands can be run from inside K8s `ceph-csi-rbd-nodeplugin` pods (the `csi-rbdplugin` container has `/usr/bin/ceph`). Requires a ceph.conf with monitor addresses and the `csi-rbd-secret` credentials:
|
||||
|
||||
```bash
|
||||
# On mgmt-runner (10.0.30.124):
|
||||
POD=$(kubectl get pods -n kube-system --no-headers | grep ceph-csi-rbd-nodeplugin | head -1 | awk '{print $1}')
|
||||
|
||||
# Get cephx key from K8s secret
|
||||
USER_KEY=$(kubectl get secret csi-rbd-secret -n kube-system -o jsonpath='{.data.userKey}' | base64 -d)
|
||||
|
||||
# Write ceph.conf into the pod and query
|
||||
kubectl exec -n kube-system $POD -c csi-rbdplugin -- bash -c "
|
||||
cat > /tmp/ceph.conf << 'EOF'
|
||||
[global]
|
||||
auth_cluster_required = cephx
|
||||
auth_service_required = cephx
|
||||
auth_client_required = cephx
|
||||
mon host = 10.0.20.50:6789,10.0.20.70:6789,10.0.20.40:6789
|
||||
[client.kubernetes]
|
||||
key = $USER_KEY
|
||||
EOF
|
||||
ceph -c /tmp/ceph.conf --name client.kubernetes status
|
||||
"
|
||||
```
|
||||
|
||||
Monitor addresses are in the `ceph-csi-config` ConfigMap: `kubectl get cm ceph-csi-config -n kube-system -o jsonpath='{.data.config\.json}'`.
|
||||
|
||||
**Pitfall:** Not all monitors are always reachable. Check connectivity first (`bash -c "echo > /dev/tcp/IP/6789"` from a Galera VM or K8s node). Only use reachable monitors in the `mon host` list.
|
||||
|
||||
**Dual-network pitfall:** PVE hosts have SSH on `10.0.20.x` (storage
|
||||
network, port 22 open) but NOT on `10.0.30.x` (management network,
|
||||
SSH key rejected). Always use `10.0.20.10` for `ssh root@10.0.20.10`.
|
||||
The `10.0.30.x` IPs are for VM bridges, not PVE host management.
|
||||
|
||||
**Fallback when PVE SSH is unavailable:** Run `ceph` commands from
|
||||
inside a `ceph-csi-rbd-nodeplugin` pod in K8s (see Monitoring section
|
||||
below for the full procedure).
|
||||
|
||||
⚠️ **Pitfall: PVE hosts are on 10.0.20.x (storage net), NOT 10.0.30.x (management net).** The 10.0.30.x addresses are for VM bridges, not PVE host SSH. Always use `root@10.0.20.10` for Ceph commands. If SSH to 10.0.30.x PVE hosts fails with "Permission denied", you're on the wrong network — check this skill first, not the SSH key.
|
||||
|
||||
### Fallback: Ceph CLI via K8s ceph-csi Pods
|
||||
|
||||
When PVE host SSH is unavailable (key rotation, host down, network change), `ceph` CLI is available inside the `csi-rbdplugin` container of `ceph-csi-rbd-nodeplugin` pods in K8s. Steps:
|
||||
|
||||
1. Find monitor addresses from the `ceph-csi-config` ConfigMap:
|
||||
```bash
|
||||
kubectl get cm ceph-csi-config -n kube-system -o jsonpath='{.data.config\.json}'
|
||||
# Parse JSON for "monitors" array (e.g. ["10.0.20.10:6789","10.0.20.20:6789",...])
|
||||
```
|
||||
2. Get cephx credentials from the `csi-rbd-secret`:
|
||||
```bash
|
||||
kubectl get secret csi-rbd-secret -n kube-system -o jsonpath='{.data.userKey}' | base64 -d
|
||||
```
|
||||
3. Write a temporary ceph.conf into the pod and query:
|
||||
```bash
|
||||
POD=$(kubectl get pods -n kube-system --no-headers | grep ceph-csi-rbd-nodeplugin | head -1 | awk '{print $1}')
|
||||
# Create /tmp/ceph.conf with [global] mon host = <reachable monitors> and [client.kubernetes] key = <userKey>
|
||||
# Then: kubectl exec -n kube-system $POD -c csi-rbdplugin -- ceph -c /tmp/ceph.conf --name client.kubernetes status
|
||||
```
|
||||
|
||||
⚠️ Not all monitors may be reachable from K8s nodes — test connectivity first. Only use monitors that respond on port 6789.
|
||||
|
||||
## Decision Trees
|
||||
|
||||
### OSD Creation
|
||||
|
||||
```
|
||||
Is target a standard PVE node?
|
||||
├─ YES → pveceph osd create /dev/sdX [--wal_dev /dev/nvme0n1pN]
|
||||
│ ⚠️ pveceph CANNOT use LVM device-mapper paths (/dev/dm-*) as WAL/DB
|
||||
│ If WAL/DB is on an LV: use ceph-volume directly (see below)
|
||||
└─ NO (non-PVE node like n5pro)
|
||||
├─ Has /etc/ceph/ceph.conf symlink to /etc/pve/ceph.conf?
|
||||
│ ├─ NO → rm /etc/ceph/ceph.conf && ln -s /etc/pve/ceph.conf /etc/ceph/ceph.conf
|
||||
│ └─ YES → proceed
|
||||
├─ WAL/DB on raw partition? → pveceph osd create /dev/sdX --wal_dev /dev/nvme0n1pN
|
||||
└─ WAL/DB on LVM LV? → ceph-volume lvm prepare --bluestore --data /dev/sdX --block.db /dev/pve/wal-db-osd-X
|
||||
(WAL auto-collocates on DB — do NOT specify --block.wal separately)
|
||||
```
|
||||
|
||||
### Recovery / Rebalance
|
||||
|
||||
```
|
||||
ceph -s shows degraded/misplaced/backfill?
|
||||
├─ Any OSD > full_ratio (0.97)?
|
||||
│ └─ EMERGENCY: ceph osd set-full-ratio 0.99 → reweight overfull OSD → fix root cause
|
||||
├─ backfill_toofull PGs > 0?
|
||||
│ └─ Target OSDs above backfillfull_ratio → raise ratio (keep ordering!) or reweight
|
||||
├─ Recovery speed too slow?
|
||||
│ └─ Tune: osd_max_backfills, osd_recovery_op_priority, osd_recovery_sleep
|
||||
│ ⚠️ ceph config set osd X does NOT propagate to running OSDs
|
||||
│ Use: ceph tell osd.* injectargs "--osd_max_backfills 3" + verify with ceph daemon osd.N config get
|
||||
├─ Recovery stuck at 0 MiB/s with remapped PGs?
|
||||
│ └─ Check for sentinel 2147483647 in up-sets → EC with OSD count == k+m → reset all weights to 1.0
|
||||
└─ upmap balancer not helping?
|
||||
└─ Balancer skips when >5% misplaced → wait for recovery, then balancer resumes
|
||||
```
|
||||
|
||||
### EC Pool with Mixed Disk Sizes
|
||||
|
||||
```
|
||||
EC pool has OSDs of different capacities?
|
||||
├─ Small OSDs filling disproportionately?
|
||||
│ └─ ROOT CAUSE: EC stores equal-sized chunks regardless of OSD capacity
|
||||
│ ceph osd reweight has marginal effect when k+m ≈ total OSD count
|
||||
│ SOLUTION:
|
||||
│ 1. Define device classes: hdd-large, hdd-small
|
||||
│ 2. Create separate CRUSH root (hdd-root) for all HDDs
|
||||
│ 3. EC rule: take hdd-root~hdd-large (only large OSDs)
|
||||
│ 4. Replicated rule: take hdd-root (all HDDs, no class filter)
|
||||
│ 5. New EC profile with k adjusted to large-OSD count
|
||||
│ 6. Migrate RBD images to new pool (online via rbd migration)
|
||||
└─ Artificially low CRUSH weights on small OSDs?
|
||||
└─ ANTI-PATTERN: causes vicious cycle → reset to true device size after separating classes
|
||||
```
|
||||
|
||||
## Core Operations
|
||||
|
||||
### 1. OSD Lifecycle
|
||||
|
||||
#### Create OSD with WAL/DB on NVMe (LVM-based)
|
||||
|
||||
```bash
|
||||
# 1. Create WAL/DB LV on NVMe
|
||||
lvcreate -L 30G -n wal-db-osd-X pve
|
||||
|
||||
# 2. Prepare OSD — ONLY --block.db, WAL auto-collocates
|
||||
ceph-volume lvm prepare --bluestore \
|
||||
--data /dev/sdX \
|
||||
--block.db /dev/pve/wal-db-osd-X
|
||||
|
||||
# 3. Activate (use osd_id + fsid from prepare output)
|
||||
ceph-volume lvm activate <osd_id> <osd_uuid>
|
||||
|
||||
# 4. Start service
|
||||
systemctl enable --now ceph-osd@<osd_id>
|
||||
|
||||
# 5. Set CRUSH host + device class
|
||||
ceph osd crush set osd.<id> <weight> root=default host=<hostname>
|
||||
ceph osd crush set-device-class hdd osd.<id>
|
||||
```
|
||||
|
||||
⚠️ **NEVER** specify `--block.wal` AND `--block.db` on the same LV — BlueStore opens the device exclusively → "Device or resource busy".
|
||||
|
||||
#### Destroy and Rebuild OSD
|
||||
|
||||
```bash
|
||||
systemctl stop ceph-osd@<id>
|
||||
ceph osd down <id>
|
||||
ceph osd destroy <id> --yes-i-really-mean-it
|
||||
ceph-volume lvm zap /dev/sdX --destroy
|
||||
# Then recreate as above
|
||||
```
|
||||
|
||||
#### Hot-Plug SATA Drive Detection
|
||||
|
||||
```bash
|
||||
# JMB58x controller doesn't auto-scan
|
||||
for h in /sys/class/scsi_host/host{0,1,2,3,4}; do
|
||||
echo "- - -" > ${h}/scan
|
||||
done
|
||||
sleep 2
|
||||
lsblk -d -o NAME,MODEL,SERIAL,SIZE,TRAN,ROTA,STATE | grep -v rbd | grep -v loop
|
||||
```
|
||||
|
||||
#### Expand BlueFS DB (Online LV Extend + Offline BlueStore Resize)
|
||||
|
||||
When BlueFS spillover occurs (RocksDB grows beyond DB LV onto HDD), expand the DB LV:
|
||||
|
||||
```bash
|
||||
# 1. Extend the LV (online, no OSD stop needed for LVM)
|
||||
lvextend -L +20G pve/ceph-db
|
||||
|
||||
# 2. Stop the OSD — bluefs-bdev-expand CANNOT run while OSD is active
|
||||
# Error if attempted online: "Device or resource busy" + assert crash
|
||||
systemctl stop ceph-osd@<id>
|
||||
|
||||
# 3. Expand BlueFS awareness of the new DB size
|
||||
ceph-bluestore-tool bluefs-bdev-expand --path /var/lib/ceph/osd/ceph-<id>
|
||||
# Output: "1 : Expanding to 0x... (... GiB)" + "size updated"
|
||||
|
||||
# 4. Restart OSD
|
||||
systemctl start ceph-osd@<id>
|
||||
|
||||
# 5. Verify BlueFS sees the new size
|
||||
ceph daemon osd.<id> bluefs stats | head -3
|
||||
# Should show: "1 : device size 0x...(XX GiB)"
|
||||
|
||||
# 6. Verify OSD is back up in cluster
|
||||
ceph osd tree | grep "osd.<id>"
|
||||
```
|
||||
|
||||
⚠️ **`bluefs-bdev-expand` MUST be run with OSD stopped.** Running it online crashes with `FAILED ceph_assert(r == 0)` — the block device is locked by the running OSD. The `ceph daemon osd.N` admin socket does NOT have an expand command (unlike config get/set). The only path is: stop OSD → bluestore-tool → start OSD.
|
||||
|
||||
⚠️ **Spillover warning clears asynchronously.** After expansion, `BLUEFS_SPILLOVER` health warning persists until BlueFS compacts RocksDB back onto the expanded DB device. This happens in the background over minutes to hours.
|
||||
|
||||
#### When NOT to Expand (Cosmetic Spillover)
|
||||
|
||||
**Spillover < 1 MiB with >50% DB free → cosmetic, do NOT expand.** Tiny spillovers (e.g. 128 KiB) are RocksDB metadata overflow that self-resolves on next compaction. Stopping an OSD for `bluefs-bdev-expand` triggers PG degradation and recovery I/O — worse than the warning itself.
|
||||
|
||||
**Expand only when:** spillover > 1 GiB OR DB device >80% full. This indicates sustained RocksDB growth that won't self-resolve.
|
||||
|
||||
### 2. CRUSH Topology
|
||||
|
||||
#### Device Class Separation
|
||||
|
||||
```bash
|
||||
# Assign device classes
|
||||
ceph osd crush set-device-class hdd-large osd.1 osd.6 osd.8 osd.11
|
||||
ceph osd crush set-device-class hdd-small osd.7 osd.10
|
||||
|
||||
# Create separate root
|
||||
ceph osd crush add-bucket hdd-root root
|
||||
|
||||
# Move HDD hosts to new root (⚠️ TRIGGERS REMAPPING!)
|
||||
ceph osd crush move <hostname> root=hdd-root
|
||||
```
|
||||
|
||||
#### CRUSH Rules
|
||||
|
||||
```bash
|
||||
# EC rule — only large HDDs
|
||||
ceph osd erasure-code-profile set ec31-hdd-large \
|
||||
k=3 m=1 crush-root=hdd-root crush-device-class=hdd-large \
|
||||
crush-failure-domain=osd plugin=jerasure technique=reed_sol_van
|
||||
ceph osd crush rule create-erasure media_ec_large ec31-hdd-large
|
||||
|
||||
# Replicated rule — ALL HDDs (no class filter)
|
||||
ceph osd crush rule create-replicated replicated_hdd_all hdd-root host
|
||||
```
|
||||
|
||||
#### Weight Management
|
||||
|
||||
| Command | Scope | Persistence | Use Case |
|
||||
|---------|-------|-------------|----------|
|
||||
| `ceph osd reweight N X` | Runtime reweight | Lost on restart | Emergency drain of overfull OSD |
|
||||
| `ceph osd crush reweight osd.N X` | CRUSH weight | Permanent | Match weight to device size |
|
||||
|
||||
**Correct CRUSH weights by device size:**
|
||||
|
||||
| Device | Weight |
|
||||
|--------|--------|
|
||||
| 1 TB | ~0.96 |
|
||||
| 2 TB | ~1.82 |
|
||||
| 3 TB | ~2.73 |
|
||||
| 3.6 TiB | ~3.64 |
|
||||
|
||||
⚠️ **Anti-pattern**: Lowering CRUSH weight to "protect" a full OSD. This prevents new PGs but doesn't move existing data → vicious cycle.
|
||||
|
||||
### 3. EC Pool Migration
|
||||
|
||||
EC pools cannot change k/m in-place. Must create new pool and migrate.
|
||||
|
||||
```bash
|
||||
# 1. Create new EC pool
|
||||
ceph osd pool create media_ec_new 128 128 erasure ec31-hdd-large
|
||||
ceph osd pool set media_ec_new ec_overwrites
|
||||
ceph osd pool application enable media_ec_new rbd
|
||||
|
||||
# 2. Create new metadata pool
|
||||
ceph osd pool create media_meta_new 32
|
||||
|
||||
# 3. Online RBD migration (VMs keep running!)
|
||||
for img in $(rbd ls media_meta); do
|
||||
rbd migration prepare media_meta/${img} media_meta_new/${img} --data-pool media_ec_new
|
||||
rbd migration execute media_meta_new/${img}
|
||||
rbd migration commit media_meta_new/${img}
|
||||
done
|
||||
|
||||
# 4. Cleanup old pool
|
||||
ceph osd pool delete media_ec --yes-i-really-really-mean-it
|
||||
ceph osd pool rename media_ec_new media_ec
|
||||
```
|
||||
|
||||
Timeline: `prepare` (instant) → `execute` (30-60 min per 500 GiB) → `commit` (atomic switch).
|
||||
|
||||
### 4. Recovery Tuning
|
||||
|
||||
#### Full Recovery Recipe
|
||||
|
||||
```bash
|
||||
# 1. Unblock backfill (keep ratio ordering: nearfull < backfillfull < full ≤ 0.97)
|
||||
ceph osd set-nearfull-ratio 0.94
|
||||
ceph osd set-backfillfull-ratio 0.96
|
||||
ceph osd set-full-ratio 0.97
|
||||
|
||||
# 2. Reweight overfull OSD (emergency, temporary)
|
||||
ceph osd reweight <osd_id> 0.8
|
||||
|
||||
# 3. Set recovery params (persistent for future OSDs)
|
||||
ceph config set osd osd_max_backfills 3
|
||||
ceph config set osd osd_recovery_op_priority 10
|
||||
ceph config set osd osd_recovery_sleep 0
|
||||
|
||||
# 4. Push to running OSDs (config set alone does NOT propagate!)
|
||||
ceph tell osd.* injectargs "--osd_max_backfills 3"
|
||||
|
||||
# 5. Verify EACH OSD picked up the config
|
||||
ceph daemon osd.<N> config get osd_max_backfills
|
||||
# Must return "3" — if "1", config not applied
|
||||
```
|
||||
|
||||
#### Post-Recovery Cleanup
|
||||
|
||||
```bash
|
||||
# Restore defaults
|
||||
ceph config set osd osd_max_backfills 1
|
||||
ceph tell osd.* injectargs "--osd_max_backfills 1"
|
||||
ceph osd reweight <osd_id> 1.0
|
||||
# Ratios: keep at 0.94/0.96/0.97 (safe for mixed-size cluster)
|
||||
```
|
||||
|
||||
### 5. RBD Orphan Cleanup
|
||||
|
||||
```bash
|
||||
# 1. Get cluster-wide VM inventory (NOT per-node qm config!)
|
||||
pvesh get /cluster/resources --type vm --output-format json | python3 -c "
|
||||
import sys, json
|
||||
for g in json.load(sys.stdin): print(g['vmid'])
|
||||
" | sort -n | tr '\n' ' '
|
||||
|
||||
# 2. Fast orphan check for specific image (O(1))
|
||||
grep -rl "vm-NNN" /etc/pve/qemu-server/ /etc/pve/lxc/ 2>/dev/null
|
||||
# Exit 0 = referenced. Exit 1 = orphan.
|
||||
|
||||
# 3. Delete with protected snapshot
|
||||
rbd snap unprotect "hdd_disk/base-NNN-disk-0"@__base__
|
||||
rbd snap purge "hdd_disk/base-NNN-disk-0"
|
||||
rbd rm "hdd_disk/base-NNN-disk-0"
|
||||
```
|
||||
|
||||
⚠️ **Always use `pvesh get /cluster/resources --type vm`** for VM existence — per-node `qm config` only sees local VMs → false positives.
|
||||
|
||||
⚠️ **Non-Ceph nodes cannot run `rbd` commands** — run from a monitor node (e.g. proxmox7 = 10.0.20.70).
|
||||
|
||||
### 6. PG Management
|
||||
|
||||
```bash
|
||||
# Reduce PG count (must disable autoscale first!)
|
||||
ceph osd pool set <pool> pg_autoscale_mode off
|
||||
ceph osd pool set <pool> pgp_num 32
|
||||
ceph osd pool set <pool> pg_num 32
|
||||
# Merge happens AFTER all PGs reach active+clean
|
||||
|
||||
# Raise mon_max_pg_per_osd (temporary relief)
|
||||
ceph config set mon mon_max_pg_per_osd 400
|
||||
ceph tell mon.* injectargs --mon_max_pg_per_osd=400
|
||||
```
|
||||
|
||||
## Pitfalls
|
||||
|
||||
### Ratio Ordering (HEALTH_ERR)
|
||||
|
||||
`nearfull < backfillfull < full ≤ osd_failsafe_full_ratio (0.97 hardcoded)`. Violation → `HEALTH_ERR: OSD_OUT_OF_ORDER_FULL`. Always set all three ratios together.
|
||||
|
||||
### injectargs Reliability
|
||||
|
||||
`ceph tell osd.* injectargs` returns empty `{}` which looks like failure but may have succeeded. **Always verify** with `ceph daemon osd.N config get <param>`. Most reliable: `ceph config set` (persistent) + `injectargs` (runtime) + verify.
|
||||
|
||||
### EC Reweight with Minimal OSDs
|
||||
|
||||
**Never reweight OSDs in EC pools when OSD count == k+m.** CRUSH places sentinel `2147483647` ("no OSD available") in up-sets → remapped PGs can never recover. Fix: reset all weights to 1.0. General rule: only reweight when OSD count > k+m.
|
||||
|
||||
### pveceph LVM WAL/DB Limitation
|
||||
|
||||
`pveceph osd create --wal_dev /dev/pve/wal-db-osd-X` fails with "unable to get device info for '/dev/dm-2'". pveceph cannot resolve LVM device-mapper symlinks. Use `ceph-volume lvm prepare` directly with `--block.db`.
|
||||
|
||||
### BlueStore WAL/DB Same Device
|
||||
|
||||
Specifying `--block.wal` AND `--block.db` on the same LV → "Device or resource busy". BlueStore opens WAL exclusively. Solution: specify ONLY `--block.db` — WAL collocates automatically.
|
||||
|
||||
### ceph config set Doesn't Propagate to Running OSDs
|
||||
|
||||
`ceph config set osd osd_max_backfills 3` updates MON config store but running OSDs keep old value. Must use `ceph tell osd.* injectargs` for runtime change, or restart OSDs. Always verify with `ceph daemon osd.N config get`.
|
||||
|
||||
### backfillfull Is Derived, Not Settable
|
||||
|
||||
Pool-level `backfillfull` flag is derived from OSD conditions, not a settable pool flag. `ceph osd pool unset <pool> backfillfull` → EINVAL. Fix: raise `backfillfull_ratio` or reweight overfull OSDs.
|
||||
|
||||
### Ghost RBD Images
|
||||
|
||||
Images appearing in `rbd ls` but `rbd info` returns empty → header gone, data objects remain. Need `rados rm` purge, not `rbd rm`.
|
||||
|
||||
### pvesm list Fails When Pool Backfillfull
|
||||
|
||||
`pvesm list hdd_disk` returns empty when Ceph pool is in backfillfull state. Fall back to `rbd ls <pool>` for image listing.
|
||||
|
||||
## Monitoring
|
||||
|
||||
```bash
|
||||
# Quick health
|
||||
ceph -s
|
||||
|
||||
# Per-OSD utilization
|
||||
ceph osd df
|
||||
|
||||
# PG state summary
|
||||
ceph pg dump pgs_brief | awk '{print $2}' | sort | uniq -c | sort -rn
|
||||
|
||||
# Recovery progress
|
||||
ceph -s | grep -E "recovery:|misplaced"
|
||||
|
||||
# Balancer status
|
||||
ceph balancer status
|
||||
|
||||
# Slow ops
|
||||
ceph health detail | grep -A5 "slow"
|
||||
```
|
||||
|
||||
### Ceph Status via ceph-csi Pod (Fallback When PVE SSH Unavailable)
|
||||
|
||||
When PVE hosts are unreachable via SSH (key rotation, network issue,
|
||||
host down), but K8s nodes can reach the Ceph monitors, run `ceph`
|
||||
commands from inside a `ceph-csi-rbd-nodeplugin` pod:
|
||||
|
||||
```bash
|
||||
# 1. Find the ceph-csi pod (on mgmt-runner or any K8s node with kubectl)
|
||||
POD=$(kubectl get pods -n kube-system --no-headers | \
|
||||
grep ceph-csi-rbd-nodeplugin | head -1 | awk '{print $1}')
|
||||
|
||||
# 2. Get monitor addresses from the ceph-csi-config ConfigMap
|
||||
MONS=$(kubectl get cm ceph-csi-config -n kube-system \
|
||||
-o jsonpath='{.data.config\.json}' | \
|
||||
python3 -c "import sys,json; print(','.join(json.load(sys.stdin)[0]['monitors']))")
|
||||
|
||||
# 3. Get the cephx key from the csi-rbd-secret
|
||||
KEY=$(kubectl get secret csi-rbd-secret -n kube-system \
|
||||
-o jsonpath='{.data.userKey}' | base64 -d)
|
||||
|
||||
# 4. Write ceph.conf + keyring into the pod, then run ceph commands
|
||||
kubectl exec -n kube-system $POD -c csi-rbdplugin -- bash -c "
|
||||
cat > /tmp/ceph.conf << EOF
|
||||
[global]
|
||||
auth_cluster_required = cephx
|
||||
auth_service_required = cephx
|
||||
auth_client_required = cephx
|
||||
mon host = $MONS
|
||||
[client.kubernetes]
|
||||
key = $KEY
|
||||
"
|
||||
ceph -c /tmp/ceph.conf --name client.kubernetes status
|
||||
ceph -c /tmp/ceph.conf --name client.kubernetes osd df
|
||||
"
|
||||
```
|
||||
|
||||
**Key details:**
|
||||
- The `csi-rbd-plugin` container has `/usr/bin/ceph` installed
|
||||
- Monitor addresses are in the `ceph-csi-config` ConfigMap (may be on
|
||||
a different subnet like `10.0.20.x` — the storage VLAN)
|
||||
- The cephx key is in the `csi-rbd-secret` Secret (field `userKey`)
|
||||
- Not all monitors need to be up — 3 of 7 was sufficient in practice
|
||||
- `client.kubernetes` has limited permissions (reads status, OSD df,
|
||||
pool df) but NOT admin operations (no `ceph osd set-full-ratio` etc.)
|
||||
|
||||
**Pitfall:** Monitor addresses in `ceph-csi-config` may be on a
|
||||
different subnet (e.g. `10.0.20.x` storage VLAN) than the management
|
||||
network (`10.0.30.x`). K8s pods can usually reach both, but Galera VMs
|
||||
may only see `10.0.30.x`.
|
||||
|
||||
**Pitfall:** The `ceph -c /tmp/ceph.conf` approach requires writing
|
||||
the config file into the pod. If `kubectl cp` is blocked, use
|
||||
`kubectl exec -- bash -c "cat > /tmp/ceph.conf << EOF ..."` via a
|
||||
heredoc through SSH to the mgmt-runner.
|
||||
|
||||
## Recovery Monitoring Cron Pattern
|
||||
|
||||
```
|
||||
Schedule: every 30 minutes
|
||||
Report: PGs remapped/backfilling, % misplaced, OSD utilization, recovery speed
|
||||
Alert: any OSD down or health ERR
|
||||
Announce "COMPLETE" when 0 remapped PGs
|
||||
```
|
||||
|
||||
## Cross-References
|
||||
|
||||
Detailed session logs and edge cases in `proxmox-ve-administration` skill references:
|
||||
- `references/ceph-osd-create-hotplug-2026-07.md` — Full OSD hot-plug walkthrough
|
||||
- `references/ceph-ec-mixed-size-optimization-2026-07.md` — Complete EC migration plan
|
||||
- `references/ceph-recovery-acceleration-2026-07.md` — Advanced recovery tuning
|
||||
- `references/ceph-crush-weight-optimization-2026-07.md` — Weight alignment for mixed sizes
|
||||
- `references/ceph-pg-management-2026-07.md` — PG count reduction + EC reweight pitfall
|
||||
- `references/ceph-pool-full-recovery-2026-07.md` — Pool-full diagnosis + OSD preparation
|
||||
- `references/ceph-orphaned-rbd-cleanup-2026-07.md` — Full orphan audit methodology
|
||||
- `references/ceph-recovery-ec-reweight-2026-07.md` — EC reweight stagnation fix
|
||||
- `references/ceph-pg-osd-management-2026-07.md` — PG/OSD rebalancing
|
||||
- `references/ceph-ec-pool-and-service-migration-2026-07.md` — EC pool + service migration
|
||||
- `references/ceph-pg-reduction-and-osd-rebalance-2026-07.md` — PG reduction + rebalance
|
||||
- `references/rbd-pool-migration-2026-07.md` — Offline RBD pool migration
|
||||
- `references/hdd-selection-for-osd-2026-07.md` — Drive selection criteria
|
||||
- `references/bluefs-db-expansion-2026-07.md` — BlueFS DB LV expansion + cosmetic spillover assessment
|
||||
|
||||
Compound-learning solution docs in `~/docs/solutions/`:
|
||||
- `bug-fixes/2026-07-13-pveceph-wal-lvm-device-mapper-limitation.md`
|
||||
- `bug-fixes/2026-07-13-ceph-ratio-ordering-constraint.md`
|
||||
- `architecture/2026-07-13-ec-mixed-disk-sizes-optimization.md`
|
||||
- `tooling/2026-07-13-injectargs-vs-config-set-reliability.md`
|
||||
@@ -0,0 +1,103 @@
|
||||
# BlueFS DB Expansion — Session Log (2026-07-13)
|
||||
|
||||
## Scenario
|
||||
|
||||
osd.7 (982 GiB HDD, proxmox7) triggered `BLUEFS_SPILLOVER` warning — RocksDB grew beyond its 50 GiB DB LV on NVMe and spilled onto the slow HDD device. DB was 50 GiB, only 2.2 GiB actually used, but BlueFS spillover is about the RocksDB *logical* size exceeding the DB device, not physical usage.
|
||||
|
||||
## Diagnosis
|
||||
|
||||
```bash
|
||||
# Check spillover target
|
||||
ceph health detail | grep -A2 SPILLOVER
|
||||
# Shows which OSD spilled and how much
|
||||
|
||||
# Check DB device size and usage
|
||||
ceph daemon osd.7 bluefs stats
|
||||
# Line 1: DB device size + usage
|
||||
# Line 2: Slow (HDD) device size + usage
|
||||
# If DB device is full → spillover to slow device
|
||||
|
||||
# Check DB LV and VG free space
|
||||
ssh root@<osd-host> 'lvs --units g | grep ceph-db && vgs --units g | grep pve'
|
||||
```
|
||||
|
||||
## Assessment Heuristic
|
||||
|
||||
**Spillover < 1 MiB with >50% DB free → cosmetic, do NOT expand.**
|
||||
The spillover is often a tiny RocksDB metadata overflow that will self-resolve on compaction. Stopping an OSD for `bluefs-bdev-expand` causes unnecessary downtime (PGs degraded, recovery triggered).
|
||||
|
||||
**Spillover > 1 GiB OR DB device >80% full → expand.**
|
||||
This indicates sustained RocksDB growth that won't self-resolve.
|
||||
|
||||
## Expansion Procedure (Verified Working)
|
||||
|
||||
osd.7 on proxmox7 (10.0.20.70), DB LV = `pve/ceph-db`:
|
||||
|
||||
```bash
|
||||
# 1. Extend LV (online, no OSD stop needed for LVM)
|
||||
lvextend -L +20G pve/ceph-db
|
||||
# Output: "Size of logical volume pve/ceph-db changed from 50.00 GiB to 70.00 GiB"
|
||||
|
||||
# 2. Stop OSD — bluefs-bdev-expand CANNOT run while OSD is active
|
||||
systemctl stop ceph-osd@7
|
||||
sleep 2
|
||||
|
||||
# 3. Expand BlueFS awareness
|
||||
ceph-bluestore-tool bluefs-bdev-expand --path /var/lib/ceph/osd/ceph-7
|
||||
# Output:
|
||||
# "1 : device size 0x1180000000(70 GiB) : using 0x89000000(2.1 GiB)"
|
||||
# "Expanding DB/WAL..."
|
||||
# "1 : Expanding to 0x1180000000(70 GiB)"
|
||||
# "1 : size updated to 0x1180000000(70 GiB)"
|
||||
|
||||
# 4. Restart OSD
|
||||
systemctl start ceph-osd@7
|
||||
|
||||
# 5. Wait for OSD to come up in cluster (10-15 seconds)
|
||||
ceph osd tree | grep "osd.7"
|
||||
# Should show "up"
|
||||
|
||||
# 6. Verify BlueFS sees new size
|
||||
ceph daemon osd.7 bluefs stats | head -3
|
||||
# Should show: "1 : device size 0x1180000000(70 GiB)"
|
||||
```
|
||||
|
||||
## What Happened After
|
||||
|
||||
- **osd.7 spillover cleared** — warning moved to osd.8 (128 KiB, cosmetic)
|
||||
- **osd.7 utilization dropped** from 86.1% → 78.9% (backfill continued, OSD healthier)
|
||||
- **osd.8 spillover assessed**: 128 KiB spill, 28 GiB DB free → NO action needed
|
||||
- **BlueFS spillover warning persists** in `ceph health` until RocksDB compacts back onto DB device — background process, minutes to hours
|
||||
|
||||
## Pitfall: Online bluefs-bdev-expand Crashes
|
||||
|
||||
Running `ceph-bluestore-tool bluefs-bdev-expand` while OSD is active:
|
||||
```
|
||||
bdev(...) open open got: (16) Device or resource busy
|
||||
bluestore(...) _lock_fsid failed to lock fsid (is another ceph-osd still running?)
|
||||
FAILED ceph_assert(r == 0)
|
||||
```
|
||||
The tool tries to open the block device exclusively, which the running OSD already holds. Must stop OSD first.
|
||||
|
||||
## Pitfall: ceph daemon admin socket Has No Expand Command
|
||||
|
||||
`ceph daemon osd.7 help` shows bluefs commands but NO `expand` or `bdev-expand` option. Available bluefs commands: `bluefs stats`, `bluefs files list`, `bluefs debug_inject_read_zeros`. The only path to expand is the offline `ceph-bluestore-tool`.
|
||||
|
||||
## osd.8 Case Study — When NOT to Expand
|
||||
|
||||
osd.8 on host `ubuntu` (10.0.30.100):
|
||||
- DB LV: `ceph-db-vg/osd-8-db` = 30 GiB on Crucial MX300 SSD (`sda3`)
|
||||
- DB usage: 2.0 GiB of 30 GiB (93% free)
|
||||
- Spillover: 128 KiB (trivial)
|
||||
- VG free: 35.67 GiB across two PVs (`sda3` 1.89 GiB free, `sda4` 33.78 GiB free)
|
||||
|
||||
**Decision: NO expansion.** 128 KiB spillover with 28 GiB DB free is cosmetic. Stopping the OSD would trigger PG degradation and recovery I/O for zero practical benefit. The spillover will likely self-resolve on next RocksDB compaction.
|
||||
|
||||
## Cross-Host SSH Pattern
|
||||
|
||||
OSDs span multiple PVE hosts. The Ceph commands (`ceph -s`, `ceph osd df`) run from any monitor node (10.0.20.10). But `ceph-bluestore-tool` and `systemctl` must run on the OSD's host. Use `ceph osd metadata N | grep hostname` to find the host, then SSH there:
|
||||
|
||||
```bash
|
||||
HOST=$(ssh root@10.0.20.10 "ceph osd metadata 7 | python3 -c \"import sys,json; print(json.load(sys.stdin)['hostname'])\"")
|
||||
ssh root@$HOST 'systemctl stop ceph-osd@7 && ceph-bluestore-tool bluefs-bdev-expand --path /var/lib/ceph/osd/ceph-7 && systemctl start ceph-osd@7'
|
||||
```
|
||||
Reference in New Issue
Block a user