8.2 KiB
ZFS-to-Ceph Migration Analysis (2026-07-03)
Context
User wants to decommission osd.1 (USB-connected Hitachi HDD on proxmox1, slow BlueStore ops) and replace it with disks from the ZFS pool on 10.0.30.100. Initial assumption was "9×2.7TB" but actual inventory revealed 4×3TB in RAIDZ2 + 2 free disks + 2 already in Ceph.
Disk Inventory (10.0.30.100, 2026-07-03)
| Disk | Model | Serial | Size | POH | Age | SMART | Role | Verdict |
|---|---|---|---|---|---|---|---|---|
| sda | Crucial MX300 SSD | 164914EEB886 | 275GB | 81,871 | 9.3y | ✅ 0/0/0 | Boot SSD (swap→DB for osd.8) | Keep — sda3 now DB device |
| sdb | WD30EFRX Red+ | WCC4N4VRJ5UU | 3TB | 81,786 | 9.3y | ✅ 0/0/0, CRC=3 | ZFS pool | Oldest in pool, first offline candidate |
| sdc | WD30EZRX Green | WMC1T2256499 | 3TB | 89,262 | 10.2y | ⛔ 371 pending, 243 uncorr | Free | REJECT — defective |
| sdd | WD30EFRX Red+ | WCC4N4SVR8JL | 3TB | 57,925 | 6.6y | ⚠️ 16 pending, 15 uncorr, Self-test FAIL | Free | REJECT — confirmed read failure |
| sde | WD30EFRX Red+ | WMC4N2399595 | 3TB | 104,327 | 11.9y | ✅ 0/0/0 | Ceph osd.8 | Active, very old. NOW HAS DB ON SSD |
| sdf | WD30EFRX Red+ | WMC4N2400243 | 3TB | 98,475 | 11.2y | ✅ 0/0/0 | Ceph osd.9 | Active, very old. NO DB device |
| sdg | WD30EFZX Purple | WX12DA01FUA7 | 3TB | 30,901 | 3.5y | ✅ 0/0/0 | ZFS pool | Best candidate |
| sdh | WD30EFZX Purple | WX12DA0R1NFH | 3TB | 30,901 | 3.5y | ✅ 0/0/0 | ZFS pool | Best candidate |
| sdi | WD30EFZX Purple | WX22DA0D7EP3 | 3TB | 30,900 | 3.5y | ✅ 0/0/0 | ZFS pool | Best candidate |
SMART Self-Test History (sdd)
Num Test_Description Status Remaining LifeTime(hours) LBA_of_first_error
# 1 Short offline Completed: read failure 90% 57925 62939576
# 2 Short offline Completed without error 00% 50673 -
# 3 Extended offline Completed: read failure 10% 50659 63009384
# 4 Short offline Completed without error 00% 50651 -
Two read failures at LBA ~63M (~32GB region), 7000h apart. Progressive degradation.
badblocks Test (sdd)
- Command:
badblocks -wsv -b 4096 /dev/sdd - Started: 2026-07-03 12:25:31 CEST
- Killed by user request at ~3.4% (Pattern 1/4: 0xaa)
- Errors found: 0 (but test was aborted early)
- Log:
/tmp/badblocks_sdd.logon 10.0.30.100 - Verdict: Inconclusive (early abort), but SMART self-test failure already disqualifies disk
ZFS Pool Details
Pool: pool01_n2_redundant
Type: RAIDZ2 (2 parity)
Disks: sdg, sdh, sdi, sdb (4× 2.73TB)
Raw: 10.9TB
Usable: ~5.46TB (2 data + 2 parity)
Used: 3.63TB (66%)
Free: 1.87TB
Fragmentation: 57%
Health: ONLINE, 0 errors
Dedup: 1.12x
Scrub: Monthly, last 2026-06-28, 0 errors
Datasets: Docker volumes, TimeMachine (~951GB), proxmox_nfs (~963GB), HomeAssistant.
Ceph State (2026-07-03)
- Health: HEALTH_WARN (BLUESTORE_SLOW_OP_ALERT on osd.0 & osd.1)
- 10 OSDs up/in, 12TB raw, 5.1TB used, 7.0TB free
- osd.1: USB Hitachi HDD 3.6TiB on proxmox1 (10.0.20.10), slow ops
- 7 pools active
Migration Strategies Evaluated
Option A: One-shot migration (recommended)
- Migrate all 3.63TB ZFS data → Ceph (has 7.0TB free)
zpool destroy→ all 4 disks freed- Create 4 new Ceph OSDs on n5pro
- Purge osd.1
Option B: Phased with ZFS degradation (CHOSEN)
zpool offline pool sdb(oldest disk) → wipe → Ceph OSD (+3TB)- ZFS now degraded (1 parity left), migrate data to Ceph
zpool destroy→ wipe sdg/sdh/sdi → 3 more Ceph OSDs (+9TB)
Chosen: Option B — allows incremental Ceph growth while ZFS data is migrated gradually.
Rejected: 2-disk offline
Offlining 2 disks from RAIDZ2 leaves 0 redundancy. Too risky during active data migration.
OSD DB/WAL Acceleration Analysis (2026-07-03)
Which OSDs Have/Lack DB Devices
| OSD | Host | HDD | DB on Flash? | Flash Available | Free on Flash |
|---|---|---|---|---|---|
| osd.1 | proxmox1 | sdc 3.6TB | ❌ | ✅ 1TB Toshiba NVMe | 848 GB — BUT osd.1 being decommissioned |
| osd.7 | proxmox7 | sda 982GB | ✅ 50GB ceph-db | ✅ 256GB Toshiba NVMe | 110 GB |
| osd.8 | ubuntu(.30.100) | sde 2.7TB | ✅ NOW 30GB on sda3 | 256GB Crucial SSD | 0 (swap consumed) |
| osd.9 | ubuntu(.30.100) | sdf 2.7TB | ❌ | Same SSD, no space | ❌ Needs root shrink or new SSD |
| osd.10 | proxmox6 | sda 982GB | ✅ 50GB ceph-db | ✅ 256GB Toshiba NVMe | 110 GB |
DB Device Addition: osd.8 (2026-07-03, SUCCESSFUL)
Method: Repurposed 32GB swap partition (sda3) on boot SSD as LVM-backed DB device.
Steps taken:
swapoff /dev/sda3(1.5GB used, 6.4GB RAM free — safe)wipefs -a /dev/sda3pvcreate /dev/sda3→vgcreate ceph-db-vg /dev/sda3→lvcreate -L 30G -n osd-8-db ceph-db-vgsystemctl stop ceph-osd@8ceph-volume lvm new-db --osd-id 8 --osd-fsid <fsid> --target ceph-db-vg/osd-8-dbsystemctl start ceph-osd@8
Result: osd.8 up with block.db → /dev/dm-2, devices = sda,sde. BlueFS spillover warning (1.1GB on HDD, 231MB on SSD) — expected to clear as compaction migrates data.
Key Lesson: ceph-bluestore-tool vs ceph-volume
ceph-bluestore-tool --command bluefs-bdev-new-db FAILS on Ceph Squid 19.2.x with "No valid bdev label found" — tried raw partition, LVM LV, with/without zap, as root and ceph user. The working tool is ceph-volume lvm new-db (requires LVM-formatted target, not raw partition).
Key Lessons
- RAIDZ cannot shrink —
zpool removedoesn't work on RAIDZ vdevs. Onlyzpool offline(temporary) orzpool destroy(permanent) can free disks. - SMART PASSED ≠ healthy — sdd passed overall SMART assessment but had confirmed self-test read failures. Always check self-test log, not just overall health.
- Assumptions about disk counts are dangerous — "9×2.7TB" was wrong; actual was 4×3TB in ZFS + 2 free + 2 in Ceph. Always verify with
lsblk+zpool statusbefore planning. - badblocks -wsv is destructive and slow — 4 patterns × ~5h per 3TB = ~20h total. Only run on disks committed to repurposing. Can be killed early if SMART already disqualifies.
- SSH key selection — 10.0.30.100 accepts
~/.ssh/id_ed25519_proxmox(root), NOT the Galera key from 1Password. - ceph-bluestore-tool bluefs-bdev-new-db is broken on Squid — Use
ceph-volume lvm new-dbinstead. Requires LVM LV as target, not raw partition. - BlueFS spillover after adding DB is normal — Existing RocksDB SSTs stay on HDD, new writes go to flash. Warning clears as compaction progresses.
- Swap can be safely repurposed for DB — If RAM is adequate (>4GB free with swap nearly empty), disabling swap to reclaim flash for Ceph DB is viable. Consider zram as swap replacement.
- ZFS has no defrag command — Fragmentation (e.g. 57%) only affects freespace map efficiency, not file integrity. The ONLY way to "defrag" ZFS is: copy data elsewhere -> destroy pool -> recreate -> copy back. Irrelevant if migrating off ZFS entirely.
- RAIDZ disk removal is permanent degradation — Pulling 1 disk from RAIDZ2 drops to RAIDZ1-equivalent (1 parity). ZFS marks the disk UNAVAIL but it stays in the vdev config.
zpool detachdoes NOT work on RAIDZ members. Pool continues degraded until destroyed. - zfs destroy -r enters D-state on degraded pools — On RAIDZ2 with defective disks,
zfs destroy -rcan hang in uninterruptible I/O (D-state) for 10+ minutes. Cannot be killed. Plan accordingly -- run in background. - ZFS async free lag — After
rm -rfon ZFS,zfs listshows OLD usage for minutes/hours.du -shshows actual live data. The gap is TXG-async free, resolves automatically. Don't panic ifzfs listsays 290GB butdusays 3GB. - Background monitor processes block zfs destroy — Processes with
cwdinside the dataset (including your own polling loops) preventzfs destroy. Uselsof +D <path>andfuser -mv <path>to find ALL blockers, kill them first. - n5pro WAL/DB prep pattern — Remove unused
pve/datathin pool (0% allocated, not in storage.cfg) vialvremove pve/data, thenlvcreate -L 30G -n wal-db-osd-X pvefor each new OSD. Verified on n5pro (10.0.20.91, 128GB NVMe, gained 70GB for 2x 30GB DB LVs).