SSD Nodes Learn Hosting plans →
Guides Matt ConnorBy Matt Connor

Remove a disk from an LVM volume group

Move data off an LVM physical volume with pvmove, then vgreduce and pvremove before you detach the disk, in the order that keeps a mounted filesystem alive.

What removing a disk from an LVM volume group means

To remove a disk from an LVM volume group, you move its data onto another device, take the now empty device out of the group, wipe the LVM label off it, and only then detach it from the server. LVM (logical volume manager) hands out storage in small chunks taken from every disk in the group, so a filesystem you think of as "one volume" can have pieces sitting on the disk you are about to pull. LVM keeps no second copy of those pieces by default. Detaching first is how people lose a filesystem that was working a minute earlier.

If the filesystem is staying and only the disk is leaving, the whole job runs online. pvmove copies data while the filesystem is mounted and serving traffic. No reboot and no rescue mode. The risky part is not the technology. It is the ordering, and the fact that every command here takes a device path and none of them asks you twice.

What a physical volume, volume group and logical volume actually are

These are three layers of one stack, and whether you can "just detach it" depends entirely on which layer your disk is carrying.

  • Physical volume (PV). A disk or a partition with an LVM label written to the front of it. pvcreate writes that label. The space behind it is divided into physical extents, which are fixed size chunks, 4 MiB by default.
  • Volume group (VG). A pool built from one or more PVs. The VG owns all of their extents together and does not care which disk any single extent came from.
  • Logical volume (LV). A slice handed out of the pool. This is the block device your filesystem lives on, and it appears as /dev/<vg>/<lv> or /dev/mapper/<vg>-<lv>.
  • The mapping between the layers is the whole problem. One LV's extents can sit on any PV in the group, and LVM will spread a single LV across several disks, because that is what a pool is for.

So the first job is to find out what the departing disk is holding. There are four situations, and they need different amounts of work.

  1. The disk carries an LVM label but was never added to a volume group. Nothing points at it. Run pvremove, then detach.
  2. The disk is a member of a volume group but holds no allocated extents. Run vgreduce, then pvremove, then detach.
  3. The disk holds extents belonging to logical volumes you are keeping. Those extents have to move to another PV first, which is what pvmove is for.
  4. The logical volume is going away along with the disk. Then unmount it and delete its /etc/fstab line before lvremove, or the next boot stops at an emergency prompt.

Plain LVM mirrors nothing unless you built a mirrored or RAID logical volume on purpose. Situation 3 with the disk already detached is therefore not a removal procedure. It is a restore from backup. If you are still planning the change, the reverse of this job, attaching a block storage volume and adding it to a volume group, is worth re-reading first, because how the disk went in decides how much work getting it out will be.

Say which device is which before you type

Start by listing what the machine has, and by asking LVM which of those devices it considers a physical volume.

lsblk -o NAME,SIZE,TYPE,FSTYPE,MOUNTPOINTS
sudo pvs

Then set the paths into shell variables once, so you type each of them a single time. Replace the quoted text with the real paths from the listing above before you run anything else.

VG="the volume group name from pvs"
OLD="the path of the disk that is leaving"
NEW="the path of the replacement disk"

Print those variables back and read them out loud against the lsblk listing before every destructive step. Two volumes of the same size from the same provider are exactly the pair that ends badly. If there is any doubt about which physical device a path points at, match the model and serial number of the device to the one you mean instead of trusting the order in which the kernel found them.

Most LVM commands accept a global dry run flag, --test, which works out the change and reports it without writing metadata. Rehearse each destructive step with it before you run the real one. Confirm the flag in man pvmove, man vgreduce and man lvremove on your own system, because the option lists shift between lvm2 releases.

Back up the group description as well. sudo vgcfgbackup writes the current layout to /etc/lvm/backup/, and vgcfgrestore can put a damaged description back. It restores metadata only, so it cannot recover data whose disk has gone.

Step 1: confirm whether this PV still holds extents

This is the question that decides everything after it, so ask the disk rather than your memory of what you did months ago.

sudo pvs -o pv_name,vg_name,pv_size,pv_free,pv_used
sudo pvdisplay "$OLD"

Read two things out of that. First, does the path in $OLD show a volume group name? An empty group column means the device is a PV that no group ever claimed, which is situation 1. Second, how much of it is used? Free space equal to total size means nothing is allocated from this device, so you can skip straight to step 4. Anything allocated means logical volumes have extents here, and they have to move. pvdisplay tells you the same story counted in PEs (physical extents), as an allocated count and a free count.

Next, ask which logical volumes are involved, and whether the rest of the group can take their data.

sudo lvs -a -o lv_name,vg_name,lv_size,seg_pe_ranges
sudo vgs -o vg_name,pv_count,vg_size,vg_free

In the lvs output, the segment ranges name the PV that each piece of each LV sits on, so look for the path in $OLD and note every LV name beside it. Those are the volumes that will be rewritten underneath while they stay mounted. In the vgs output, compare the group's free space against the amount allocated on $OLD. If the remaining disks in the group already have room for it, you need no replacement device and can go to step 3. If they do not, step 2 is where you fix that.

One check worth making before any of this: if the reason for the change is that a filesystem filled up, confirm the space is genuinely in use. df and du disagreeing about a full filesystem usually means deleted files are still held open by a running process, and moving storage around does not fix that.

Step 2: extend the group onto the replacement device

Attach the replacement volume at your provider, confirm which path it arrived as, then label it and add it to the group.

sudo apt install lvm2   # Ubuntu and Debian ship the tools in this package
sudo pvcreate "$NEW"
sudo vgextend "$VG" "$NEW"
sudo vgs -o vg_name,pv_count,vg_size,vg_free

pvcreate writes the LVM label and nothing else. vgextend adds that device's extents to the pool. Neither command touches data in the existing logical volumes, because allocation only changes when you ask for it.

LVM accepts a whole raw device, so pvcreate /dev/sdb works. Putting one partition across the disk costs a single extra step and leaves a signature that other tools understand, so the next person running lsblk sees a purpose instead of an unlabelled disk. Use a GPT partition table rather than MBR for anything you might grow past 2 TiB. Size the replacement with headroom too, since the advertised size and the size the tools report are not the same number, and LVM metadata takes a little more off the top. A replacement that is exactly as large as the original on paper can end up slightly smaller in extents.

The last command in that block asks the question that gates step 3: does the group now hold more free extents than are allocated on $OLD? Until it does, pvmove has nowhere to put the data and will decline the job.

Step 3: pvmove the extents, online and resumable

sudo pvmove --test "$OLD"
sudo pvmove -i 30 "$OLD"

Here is the mechanism, because it explains why this is safe on a mounted filesystem. For each segment that has to move, pvmove creates a temporary mirror inside the logical volume, copies the extents to their new home, waits for both sides to be in sync, then drops the old side and rewrites the group metadata. All of that happens in the device mapper layer below the filesystem, so the filesystem never sees a gap and applications keep reading and writing throughout. The -i 30 flag prints progress every 30 seconds. To choose the destination yourself instead of letting LVM pick, name it: sudo pvmove "$OLD" "$NEW".

  • It is slow, because it copies every allocated extent at the speed of the slower device. On a network attached volume that is network speed. Run it inside tmux or screen so a dropped SSH session cannot kill it.
  • It competes with your workload for disk I/O. Start it when the server is quiet, or move one volume at a time with sudo pvmove -n <lvname> "$OLD" to spread the load over several windows.
  • It is resumable. If the machine reboots or the process dies part way through, run sudo pvmove with no arguments to continue the unfinished moves. sudo pvmove --abort cancels instead, and leaves every logical volume pointing at one consistent copy of its data.
  • Some layouts refuse to move in place. Devices under a cache layer or a thin pool, and certain RAID segment types on older lvm2 releases, are not candidates. If pvmove declines, read what it says and check man pvmove for the restriction. Adding --force does not make an unsupported layout supported.

When it finishes, ask step 1's question again, about $OLD specifically. Is anything still allocated on that device? Allocated extents left behind mean a move that has not finished rather than one that failed quietly, and sudo pvmove with no arguments will carry on with it.

Step 4: vgreduce the device out of the group

sudo vgreduce --test "$VG" "$OLD"
sudo vgreduce "$VG" "$OLD"

vgreduce drops the device from the group's member list and rewrites the volume group metadata onto the PVs that remain. It refuses while the device still has allocated extents, because agreeing would cut pieces out of a live logical volume. Treat that refusal as the safety check working, and go back to step 3.

Do not reach for vgreduce --removemissing at this point. That flag exists for a group whose disk has already disappeared, and combined with --force it deletes every logical volume that had extents on the missing device. It is a repair tool for an accident that has already happened.

Step 5: pvremove the label, then detach

sudo pvremove "$OLD"
sudo wipefs -a "$OLD"
sudo lvmdevices --deldev "$OLD"   # only on systems using /etc/lvm/devices/system.devices

pvremove erases the LVM label, so the device stops being a physical volume. It refuses while the device still belongs to a volume group, which is why step 4 comes first. wipefs -a clears the signatures that are left, including any partition table, so the disk does not present half an identity to the next tool that scans it. On lvm2 releases that use a devices file rather than a filter, remove the path from that file too, or the file goes on naming a device that is not there.

Now detach the volume at your provider, and confirm afterwards with lsblk and sudo pvs that the path has gone and the group lists only its remaining members. Because you removed the device from the metadata before removing it from the machine, the group activates completely at the next boot with nothing to warn about. Detaching first produces the opposite: a group that activates partially, a warning on every LVM command, and any logical volume with extents on the missing disk unusable until you repair the group.

If the logical volume is going away, fstab comes first

Sometimes the disk is leaving because the thing living on it is finished. Then you are removing a logical volume, and the order matters more here than anywhere else in this guide. Unmount, delete the /etc/fstab line, and only then lvremove.

sudo umount /path/to/mountpoint
sudo swapoff "/dev/$VG/<lvname>"   # only if this LV was swap
sudo nano /etc/fstab               # delete the line naming this LV
sudo systemctl daemon-reload
sudo lvremove "$VG/<lvname>"

That order exists because systemd turns every /etc/fstab line into a mount unit at boot. Remove the logical volume while its line is still in the file and the unit waits for a device that will never appear, so local-fs.target fails and the boot stops at an emergency prompt asking for the root password on the console. On a rented server that means hunting for the provider's console before the machine is reachable at all. Running systemctl daemon-reload after editing fstab makes systemd forget the generated unit straight away, so nothing tries to remount the volume in the seconds before you delete it.

If umount reports the filesystem as busy, find the process holding it with sudo fuser -vm /path/to/mountpoint or sudo lsof +f -- /path/to/mountpoint, and stop that process properly. Avoid umount -l. A lazy unmount detaches the mount from the namespace while open files keep working, which means you can lvremove storage that something is still writing to. Check the less obvious users of the path as well: NFS exports in /etc/exports, backup and rsync jobs in cron, and container storage, since a bind mount in a compose file points at a host path that will quietly start filling the root filesystem once the logical volume under it is gone.

When the group has nowhere to put the data

If you cannot attach a replacement device and the other PVs have no room, the data has to get smaller before the disk can leave. That is a filesystem job before it is an LVM job, and the filesystem decides whether it is possible at all.

  • ext4 shrinks, but only while unmounted. Unmount the volume, run sudo e2fsck -f on it, then sudo resize2fs <lv> <new size>, then sudo lvreduce -L <new size> "$VG/<lvname>". sudo lvreduce --resizefs -L <new size> "$VG/<lvname>" does both halves through fsadm, with the same rules underneath.
  • XFS does not shrink, at any version. The only route is a backup, a smaller new logical volume, and a restore into it.

The rule that catches people: the logical volume must never end up smaller than the filesystem it carries. Shrink the filesystem first, then the volume, and leave the volume slightly larger than the filesystem asked for. lvreduce asks for confirmation because doing this in the wrong order truncates the end of the filesystem, and growing the volume back afterwards does not undo it. Once the volume is smaller, the freed extents return to the group, and step 3 has somewhere to move the data to.

The order, in one list

  1. Ask what the PV holds, with pvs, pvdisplay and lvs -a -o lv_name,seg_pe_ranges.
  2. Give the group room, either by vgextend onto a replacement device or by freeing extents elsewhere.
  3. pvmove the extents off the departing device, online, resumable, rehearsed with --test.
  4. vgreduce the device out of the group.
  5. pvremove the label and wipefs what is left.
  6. Detach at the provider, then confirm with lsblk and pvs.

When the logical volume is going too, the first four lines change to: unmount, edit /etc/fstab, systemctl daemon-reload, lvremove.

FAQ

Can I just detach the block volume if I do not need what is on it?

Only when nothing else needs it, and you check that rather than assume it. Ask pvs -o pv_name,pv_used,pv_free and pvdisplay whether that physical volume still has allocated extents, and lvs -a -o lv_name,seg_pe_ranges which logical volumes those extents belong to. If another logical volume has pieces on the disk, detaching it breaks that volume, because plain LVM keeps no second copy of an extent anywhere. Move the extents with pvmove, take the device out of the group with vgreduce, clear the label with pvremove, and then the detach is uneventful.

Does pvmove work while the filesystem is mounted, and what if it is interrupted?

Yes, that is what it is built for. pvmove creates a temporary mirror below the filesystem, copies the extents, then switches the logical volume to the new location inside device mapper, so applications keep reading and writing the whole time. If a reboot or a dropped session interrupts it, run sudo pvmove with no arguments to resume the unfinished moves, or sudo pvmove --abort to stop and leave every volume on one consistent copy. Start it inside tmux or screen, because it copies every allocated extent and takes as long as your disks take.

Why did my server stop booting after I removed a logical volume?

Because the volume's /etc/fstab line was still in the file. systemd generates a mount unit from every fstab entry, that unit waits for a device that no longer exists, local-fs.target fails, and the boot ends at an emergency prompt on the console. Recovery means reaching the provider's console or a rescue image, deleting the stale line, and rebooting. Avoid it by unmounting, deleting the fstab line, running systemctl daemon-reload, and only then running lvremove. Adding nofail to the fstab lines of non-root data mounts makes the same mistake survivable next time.

Why does vgreduce refuse to remove the disk from my volume group?

Usually because that physical volume still has allocated extents on it, so removing it would cut pieces out of a live logical volume. Check with pvs -o pv_name,pv_used,pv_free and pvdisplay against the device, then finish or resume the pvmove and try again. The other common cause is a mistyped path that names a device in a different group, which is why setting the paths in shell variables and reading them back before each step is worth the two seconds it costs.