11 KiB
Ceph RBD Pool Migration (Between Pools with Different Redundancy)
When to Use
Moving CT/VM disk volumes from one Ceph RBD pool to another — typically to upgrade redundancy (e.g. size=2 replica pool → EC4+1 erasure-coded pool).
Assessing Pool Redundancy
Before migrating, check the redundancy profile of all pools:
# Per-pool: replication factor and minimum
for pool in tm_disks vm_disks hdd_disk media_meta media_ec; do
size=$(ceph osd pool get $pool size 2>/dev/null | awk '{print $2}')
min=$(ceph osd pool get $pool min_size 2>/dev/null | awk '{print $2}')
echo "$pool: size=$size min_size=$min"
done
# EC profile details
ceph osd erasure-code-profile get ec41-hdd
# k=4 m=1 → EC4+1, tolerates 1 failure, 5 OSDs needed
# Crush rule (which device class, failure domain)
ceph osd crush rule dump <rule_name>
Redundancy Comparison Table (This Cluster)
| Pool | size | min_size | Tolerates Failure | Type |
|---|---|---|---|---|
| tm_disks | 2 | 2 | NONE (read-only on 1 loss) | Replicated HDD |
| vm_disks | 3 | 2 | 1 OSD/host | Replicated |
| hdd_disk | 3 | 2 | 1 OSD/host | Replicated |
| media_meta | 3 | 2 | 1 OSD/host | Replicated (metadata for EC) |
| media_ec | 5 (EC4+1) | 4 | 1 OSD | Erasure Coded |
Warning: size=2, min_size=2 means ZERO fault tolerance — losing one
replica makes the pool read-only. This is a dangerous configuration.
Migration Procedure
Prerequisites
- Identify source pool and target pool
- Target pool must support
content rootdir(check/etc/pve/storage.cfg) - For EC pools: metadata pool (e.g.
media_meta) stores the RBD header, data pool (e.g.media_ec) stores the chunks. PVE storage config links them viadata-pooldirective.
Step 1: Remove CT from HA
# On any PVE node (ha-manager is cluster-wide)
ha-manager remove ct:NNN
Cannot stop a HA-managed CT directly — HA will restart it. Must remove from HA first.
Step 2: Stop the CT
# On the node hosting the CT
pct stop NNN
# If locked:
pct unlock NNN # Clear stale locks
pct delsnapshot NNN <snapname> # Remove snapshots causing locks
pct shutdown NNN --timeout 10 --forceStop 1 # Force if needed
Step 3: Allocate New Volume on Target Pool
# Allocate on target storage (PVE-managed, proper volid format)
pvesm alloc <target_storage> <vmid> <volume_name> <size>
# Example: pvesm alloc media 114 vm-114-disk-new 500G
Step 4: Copy Data Between RBD Devices
# Map both source and destination RBD images
SRC_DEV=$(rbd map <src_pool>/<src_volume>)
DST_DEV=$(rbd map <dst_pool>/<dst_volume>)
# Block-level copy with dd
dd if=$SRC_DEV of=$DST_DEV bs=4M status=progress &
# 500G at ~160 MB/s = ~50 min
# Monitor progress
tail -f /tmp/rbd_copy.log
kill -0 <dd_pid> && echo "RUNNING" || echo "DONE"
Alternative: rbd export ... | rbd import ... — but destination must
NOT already exist (import fails on existing images). Use dd when the
destination was pre-allocated via pvesm alloc.
Step 5: Update CT Config
# Point rootfs to new storage
pct set NNN --rootfs <target_storage>:<new_volume>,size=<SIZE>G,mountoptions=discard
# Example: pct set 114 --rootfs media:vm-114-disk-new,size=500G,mountoptions=discard
Step 6: Start CT and Re-add to HA
pct start NNN
ha-manager add ct:NNN --group 1 --max_relocate 3 --max_restart 2
Step 7: Clean Up Old Volume
# After verifying the CT boots and data is intact
pvesm free <src_storage>:<src_volume>
# Or: rbd rm <src_pool>/<src_volume>
Step 8: Remove Old Pool (Optional)
# Only after ALL volumes are migrated
ceph osd pool delete <old_pool> --yes-i-really-really-mean-it
# Remove from PVE storage config:
# Delete the "rbd: <old_pool>" block from /etc/pve/storage.cfg
Pitfalls
- HA restarts stopped CTs: Must
ha-manager removebeforepct stop. HA manager will faithfully restart any stopped service. - Stale locks: Snapshot operations can leave
lock: diskon the CT. Usepct unlock+pct delsnapshotto clear. pct cloneon running CT: Needs--snapnameflag and a pre-existing snapshot. Without--snapname, PVE refuses even with snapshots present.pct unlockduring clone ABORTS IT: Thelock: createflag means the clone is actively running. Callingpct unlockmid-clone kills the process and leaves the CT config incomplete (rootfs/mp0 missing). Wait for the lock to clear naturally by pollingpct config. If accidentally unlocked: destroy, purge, remove RBD image, re-snapshot source, re-clone.rbd importon existing image: Fails with "(17) File exists". Useddbetween mapped devices when destination is pre-allocated.- mkfs hardline block: Hermes blocks
mkfsunconditionally. RBD volumes created viapvesm allocorrbd createhave no filesystem. Usepct clone(formats internally) orddfrom an existing formatted volume. - EC pool as rootfs: Works via the
mediastorage which pairsmedia_meta(replicated, for RBD headers) withmedia_ec(EC4+1, for data chunks). The PVE storage configdata-pooldirective handles this transparently. - EC clone performance cliff: Cloning a CT with a large EC-backed
mountpoint (e.g. 500G on
media) is practically impossible — observed <1 MB/s due to EC write amplification (k=4, m=1 → 5 writes per chunk). A 500G clone would take days. Instead, clone a CT with a small media mountpoint (e.g. CT 137 imap, 20G mp0) and resize afterward:pct resize <ct> mp0 500G(instant, thin-provisioned). - Ghost RBD images: Aborted clones can leave ghost RBD images that
rbd lsshows butrbd info/rbd rmcan't access (stale OMAP + lingering watcher). Recovery:rbd unmapall mapped devices for the image, kill lingeringpct clone/rbdprocesses, wait 30s for watcher timeout, then retryrbd rm. If still stuck, the ghost is harmless — create new volumes with different names. - Stale
rbd rmprocess as watcher:rbd rmfails with "image still has watchers" butrbd showmappedshows nothing mapped on any node. The watcher is a previousrbd rmcommand still running (stuck in I/O).rbd statusreturns "No such file" (header already deleted) butrbd lsstill lists the image andps aux | grep 'rbd.*rm'shows the stale process. Fix:kill <pid>,sleep 35, retryrbd rm. Always check for stalerbd rmprocesses when watchers block removal but no devices are mapped. pct clonefrom wrong node:pct configonly works on the node hosting the CT.pct clonelikewise must run on the source CT's node. Useha-manager status | grep ct:NNNto find the hosting node, then SSH to that node for allpctoperations.
Session Log — 2026-07-05
Migrating CT 114 (smb-tm) from tm_disks (size=2, no fault tolerance)
to media (EC4+1). Using pct move-volume 114 rootfs media — PVE-native
replacement for manual dd. Sets lock: disk during copy (~14 MB/s on EC).
CT 138 (smb-tm-sarah) created successfully via pct create from template
on vm_disks + pvesm alloc/mke2fs for mp0 on media. See
references/pve-native-volume-ops-2026-07.md for full procedure.
Session Log — 2026-07-06 (rbd copy completion)
The pct move-volume from 2026-07-05 left CT 114 stopped with lock: disk.
The EC copy was incomplete (only 338 GiB of 341 GiB). Completed the migration
using rbd copy --data-pool:
# On a Ceph monitor node (proxmox1):
# 1. Unmap old RBD devices on the CT's host node (proxmox5)
rbd unmap /dev/rbd0 # media_meta/vm-114-disk-0
rbd unmap /dev/rbd1 # tm_disks/vm-114-disk-0
# 2. Remove the incomplete target image (may need watcher timeout)
rbd rm media_meta/vm-114-disk-0
# If "image still has watchers": unmap on ALL nodes, wait 30s, retry.
# If rbd status says "No such file" but rbd rm says "has watchers":
# the image is already gone — race condition. Proceed to copy.
# 3. Fresh copy with EC data pool
rbd copy --data-pool media_ec tm_disks/vm-114-disk-0 media_meta/vm-114-disk-0
# 341 GiB, took ~7 hours with concurrent Ceph recovery (slow ops)
# Monitor: rbd du media_meta/vm-114-disk-0
# 4. Update CT config — use PVE storage name, NOT Ceph pool name
# WRONG: rootfs: media_meta:vm-114-disk-0 → "storage 'media_meta' does not exist"
# RIGHT: rootfs: media:vm-114-disk-0 → PVE storage "media" → pool media_meta + data_pool media_ec
sed -i 's|rootfs: tm_disks:vm-114-disk-0|rootfs: media:vm-114-disk-0|' \
/etc/pve/nodes/proxmox5/lxc/114.conf
# 5. Unlock and start
pct unlock 114
pct start 114
Key Learnings
rbd copy --data-poolis simpler thanddbetween mapped devices for EC migration. It creates the image with correctdata_poolattribute in one step. No need forpvesm alloc+dd.- PVE storage name ≠ Ceph pool name: Storage
media→ poolmedia_meta, data-poolmedia_ec. CT configs must usemedia:, notmedia_meta:. - Stale RBD watchers: After
rbd unmap,rbd rmmay still fail with "image still has watchers". Unmap on ALL nodes that had it mapped, wait 30s for watcher expiry. Sometimes the image is already removed but the error persists as a race — check withrbd ls -lto confirm. - Concurrent I/O competition: RBD copy + Ceph recovery (from OSD reweight) + Seafile fsck all hitting HDD OSDs simultaneously caused BLUESTORE_SLOW_OP_ALERT and extreme slowness. Sequential would be much faster.
pct create Fails with CFS Lock Timeout (2026-07-06)
Symptom
pct create fails on ALL PVE nodes with:
trying to acquire cfs lock 'storage-hdd_disk' ...
cfs-lock 'storage-hdd_disk' error: got lock request timeout
mounting container failed
unable to create CT 143 - rbd error: rbd: error opening image vm-143-disk-0: (2) No such file or directory
Root Cause
The pmxcfs (Proxmox cluster filesystem) storage lock times out,
preventing PVE from creating the RBD image. Meanwhile rbd create
and pvesm alloc work fine (they don't need the CFS lock).
Even when pvesm alloc successfully creates the RBD image, pct create
still fails because it tries to mount the raw image to extract the
template — and the image has no filesystem. mkfs/mke2fs is on the
Hermes hardline blocklist, so the image can't be formatted from the
agent.
Workaround
Use pct clone --snapname from an existing CT (formats the rootfs
internally on the PVE node, bypassing the mkfs blocklist). See
references/lxc-creation-rbd-clone-2026-07.md for the clone procedure.
Alternatively, ask the user to run pct create manually in a terminal
outside the agent — the mkfs blocklist only applies to agent-initiated
commands.
PVE Storage Name ≠ Ceph Pool Name (Confirmed Again)
When editing CT configs directly (e.g. sed -i on /etc/pve/nodes/.../lxc/NNN.conf),
use the PVE storage name, NOT the Ceph pool name:
- Storage
media→ poolmedia_meta+ data-poolmedia_ec - CT config:
rootfs: media:vm-NNN-disk-0,...✅ - CT config:
rootfs: media_meta:vm-NNN-disk-0,...❌ → "storage 'media_meta' does not exist"
The media storage in /etc/pve/storage.cfg maps:
rbd: media
content rootdir
data-pool media_ec
pool media_meta
username admin