ZFS settings you can't change later
Which ZFS settings you can't change later on a VPS: ashift, vdev layout and creation only properties, how to read them back, and what a rebuild costs.
Which ZFS settings can you not change later?
The ZFS settings you can't change later all share one thing: they are fixed at create time, either when a vdev (virtual device, the unit a pool is built from) is made or when a dataset is made. ashift is fixed for the life of a vdev. The vdev layout, meaning how many top-level vdevs the pool stripes across and in what redundancy shape, cannot be undone for raidz and can only sometimes be undone for the rest. Dataset properties such as encryption and volblocksize are set by zfs create and never move after that.
Everything else is a property you can change this afternoon. The catch is what "change" means. Setting recordsize or compression affects blocks written after the change. The data already on disk keeps the old behaviour until something rewrites it.
On a rented VPS this matters more than it does in a rack. You usually get one virtual disk, no spare bay to migrate into, and a monthly bandwidth allowance that decides how painful a rebuild is. So the create-time decisions deserve twenty minutes before you run zpool create.
What can you actually set on a VPS image?
Almost no provider image ships a ZFS root filesystem. You get ext4 or xfs on one virtual disk, and you build a pool on a second volume you attach. Check what you have before planning anything.
zfs version
lsblk -o NAME,SIZE,ROTA,MODEL,SERIAL
sudo zpool statuszfs version prints two lines, the userland version and the loaded kernel module version, something like zfs-2.2.2-0ubuntu9 followed by zfs-kmod-2.2.2-0ubuntu9. A missing second line means the module is not loaded. On Ubuntu 24.04 the package is zfsutils-linux:
sudo apt update && sudo apt install -y zfsutils-linuxTwo limits come from the platform rather than from you. On container virtualisation that shares the host kernel, such as LXC or OpenVZ, you cannot load a kernel module at all, so ZFS is not available on that plan. And on a plan with one virtual disk, your pool is a single-device vdev. ZFS still checksums every block, so it will tell you the data is wrong, but with one copy it has nothing to repair from. zpool status names the damaged file and counts the errors, and that is all it can do. Set copies=2 at dataset creation if the data matters more than the space, and keep real backups regardless, because a snapshot on the same pool is not a backup.
Which disk you put the pool on is its own decision, and it changes throughput by an order of magnitude: read local NVMe against network attached block storage before you commit.
Why ashift 12 is the safe choice on a virtual disk
ashift is the base 2 logarithm of the smallest block ZFS aligns its I/O to. ashift=9 means 512 bytes, ashift=12 means 4 KiB, ashift=13 means 8 KiB. The default is 0, which the manual describes as "auto-detect using the kernel's block layer and a ZFS internal exception list".
Auto-detect is where VPS disks cause trouble. A virtio or virtual SCSI disk usually reports 512 byte logical and 512 byte physical sectors, because the hypervisor presents a generic geometry instead of passing through what the NVMe underneath really does. ZFS believes it and picks ashift=9. The device then turns every 4 KiB write into a read, modify, write cycle internally, and ZFS allocates metadata in 512 byte units, so the pool does more physical work per logical write than it needs to. The device never reports this. You only see it as write latency that does not match the hardware you are paying for.
Set the value yourself:
sudo zpool create -o ashift=12 -O compression=lz4 tank /dev/disk/by-id/virtio-abcdef123456Use a path under /dev/disk/by-id/. Kernel names like /dev/vdb are assigned in discovery order, so attaching another volume can renumber them, and a pool whose config points at the wrong name imports with a device missing.
The cost of ashift=12 on a device that really is 512 bytes is wasted space: every allocation rounds up to 4 KiB, which hurts on a pool full of very small files. That is the whole downside, and it is small next to the alternative. ashift=13 is right only when you know the device writes in 8 KiB units.
Read the manual sentence that catches people: "Changing this value will not modify any existing vdev, not even on disk replacement." The pool property named ashift is the value used for vdevs added in future. It does not describe the vdev you already have.
How do you read the ashift you actually got?
Three commands, answering three different questions.
zpool get ashift tank
sudo zpool get all tank all-vdevs | grep ashift
sudo zdb -C tank | grep ashiftThe first prints the pool property, and 0 there is normal: it means vdevs added later will auto-detect. The second prints the read-only ashift vdev property for each vdev, which the manual defines as "the physical sector size of this vdev expressed as the power of two". The third reads the on-disk pool configuration and prints a line such as ashift: 12 for each vdev. Trust the second and the third.
If those report 9 on a pool you meant to build at 12, nothing you set now will fix it. The vdev keeps that value until the vdev itself is gone, so read the rebuild section below.
recordsize is per dataset, and only touches new blocks
recordsize is the largest logical block a file in that dataset can use. The default is 128 KiB. A file smaller than the record size is stored in a single block sized to fit it, so the property is a ceiling rather than a fixed stripe width.
The manual is direct about the half that surprises people: "Changing the file system's recordsize affects only files created afterward; existing files are unaffected." The property is changeable, and changing it does nothing to the data you already hold.
zfs get -r recordsize tank
sudo zfs set recordsize=16K tank/postgresWhy match it to a workload? Because a record is the unit of read, modify, write. A database writing an 8 KiB page inside a 128 KiB record makes ZFS read the whole 128 KiB record, apply the change, compress it, and write a new 128 KiB record elsewhere. That is sixteen times the intended write. Matching recordsize to the page size that PostgreSQL or InnoDB uses removes that amplification. In the other direction, a dataset holding video files or backup archives does less metadata work at recordsize=1M.
To apply a new recordsize to existing data you have to rewrite it. Copying the files works. So does sending the dataset into a new one. OpenZFS 2.3 and later also ship a zfs rewrite subcommand, documented as zfs rewrite [-rvx] [-l length] [-o offset] file|directory, which rewrites blocks in place under the current properties. Check man zfs-rewrite on your build before you rely on it. Watch free space while it runs: if a snapshot still references the old blocks, you now store both copies, which is the same accounting that makes snapshots and clones look free until they suddenly are not.
For zvols the matching property is volblocksize, and that one is final: "The blocksize cannot be changed once the volume has been written." A zvol handed to a virtual machine at the wrong block size is a create-time mistake with no property fix.
vdev layout is the decision you cannot undo
A pool stripes across its top-level vdevs. Adding one widens the stripe, and from then on the pool depends on every vdev in it.
zpool attach adds a device to an existing vdev. On a single disk that creates a mirror. On a raidz vdev, in OpenZFS 2.3 and later, it starts a raidz expansion. That expansion has limits worth knowing before you plan around it. The manual states that "expansion does not change the number of failures that can be tolerated without data loss (e.g. a RAID-Z2 is still a RAID-Z2 after expansion)", and that "old blocks retain their old data-to-parity ratio", so you gain less usable space than the arithmetic suggests until the old data is rewritten.
zpool remove takes a top-level vdev back out, but only in a narrow case. The manual lists the removable kinds as "hot spare, cache, log, and both mirrored and non-redundant primary top-level vdevs, including dedup and special vdevs", and requires that the pool holds no top-level raidz or draid vdev, that all top-level vdevs have the same ashift, and that keys for all encrypted datasets are loaded. Two of those conditions are the trap. One raidz vdev anywhere in the pool makes every top-level vdev permanent. Mismatched ashift does the same, which is one more reason to set ashift by hand the first time. Nothing converts raidz1 into raidz2, and nothing shrinks a raidz vdev.
The mistake that costs people a pool on a VPS is one word:
sudo zpool add -n tank /dev/disk/by-id/virtio-second-volume
sudo zpool attach tank virtio-first-volume virtio-second-volumeadd makes the new volume a second top-level vdev, so the pool stripes across both and losing either one loses everything on it. attach makes it a mirror of the device you name. Run add with -n first, which displays the configuration that would be used without applying it, and read the tree it prints before you drop the flag. If ZFS refuses with a mismatched replication level message, that refusal is correct, and -f is almost never the right response. If you are weighing mirrored pairs against parity for a rented pool, the RAID 10 tradeoff for VPS storage covers why mirrors stay popular for anything that takes random writes.
What does fixing one of these actually cost?
The fix for a wrong ashift, or for a vdev layout you cannot undo, is the same procedure: build a new pool and move the data with zfs send. On a VPS with one disk the new pool cannot live beside the old one, so the data goes to a second host and comes back.
sudo zfs snapshot -r tank@migrate
sudo zfs send -Rc tank@migrate | ssh user@second-host "sudo zfs recv -s -F backup/tank"-R sends the dataset tree with its properties and snapshots. -c sends blocks in the compressed form they already have on disk, so the link carries the compressed size. -s on the receiving side saves a resume token when the stream breaks, and zfs send -t <token> continues from it instead of starting again, which matters when the transfer is measured in days.
Time is set by the link, not by the pool. The rows below are arithmetic for 2 TB of data at 80 percent of line rate, not a measurement:
The data behind this chart
[
{
"label": "100 Mbit link",
"throughput_mb_s": 10,
"hours_for_2tb": 55.6
},
{
"label": "250 Mbit link",
"throughput_mb_s": 25,
"hours_for_2tb": 22.2
},
{
"label": "1 Gbit link",
"throughput_mb_s": 100,
"hours_for_2tb": 5.6
},
{
"label": "2.5 Gbit link",
"throughput_mb_s": 250,
"hours_for_2tb": 2.2
}
]On a 100 Mbit port that is 55.6 hours at 10 MB/s, and you pay it twice, once out and once back, so budget more than four days for the round trip plus the pool rebuild in between. On a 1 Gbit link the same 2 TB moves in 5.6 hours each way. Check the meter as well: 2 TB out and 2 TB back is 4 TB counted against a monthly transfer allowance that is often 2 TB or 4 TB on a small plan.
Keep the downtime separate from the transfer. Send the full stream while services run, then stop them, take a second snapshot, and send an incremental with zfs send -Rci tank@migrate tank@final. The incremental carries only what changed during the first copy, so the outage is minutes rather than days.
Where dedup is still a trap
Dedup is a per-dataset property you can turn on and off, which makes it look reversible. Turning the property off does not undo the table it built.
Every unique block written under dedup=on gets an entry in the pool's dedup table. That table lives in the pool and is cached in ARC (adaptive replacement cache, the in-memory cache ZFS keeps). When it stops fitting in memory, each write turns into random reads to check the table, and write latency collapses. Plan against the manual's figure: "It is generally recommended that you have at least 1.25 GiB of RAM per 1 TiB of storage when you enable deduplication." A 4 GB VPS runs out of that budget around 3 TB of pool, and that RAM is the same RAM ARC was using to keep reads fast.
Setting dedup=off stops new blocks entering the table. Entries already there stay until the blocks they describe are freed or rewritten, so the pool keeps paying for a decision you have reversed. Look at what you have with sudo zpool status -D tank, which prints "a histogram of deduplication statistics, showing the allocated (physically present on disk) and referenced (logically referenced in the pool) block counts and sizes by reference count".
Turn on compression instead. compression=lz4 or compression=zstd costs CPU you can measure and no RAM you cannot get back, and it is changeable at any time for new blocks. Dedup earns its keep on a narrow set of workloads, such as many near-identical virtual machine images. The honest test before enabling it is sudo zdb -S tank, which simulates the dedup table for the data you already have and prints the ratio you would get.
When does a full pool start costing write speed?
ZFS never overwrites a block in place. A change to an existing file allocates new space, writes there, and frees the old block afterwards. So a pool near full is slow at writing, and it is slow in a way no defragmentation command can fix, because ZFS has no such command.
Two things happen as capacity climbs. Free space becomes scattered, so the allocator spends longer finding a contiguous run for each write. And ZFS holds back slop space, a reserve of one thirty-second of the pool by default, that no write may touch. That reserve is part of why the space you can actually use is less than the space reported free.
zpool list -v
zfs list -o name,used,avail,refer,usedbysnapshotszpool list prints CAP and FRAG. FRAG describes fragmentation of free space, not of your files, so watch its trend rather than one reading. The practical line is 80 percent: past that, write latency rises noticeably, and past 90 percent it gets bad. Snapshots are the usual reason a pool fills with nobody adding data, so usedbysnapshots is the column to read first. A regular scrub schedule for a small pool gives you a second, independent signal about pool health while you watch capacity.
There is no shrinking your way out on a VPS. Pool size is plan size, so size it at buy time. The advertised number is not the usable number either, for the same base 2 against base 10 reason a 500 GB disk shows 465 GB, and ZFS then takes parity, slop and metadata on top of that. Work through a storage VPS checklist before committing to a plan you will live in for a year.
The twenty minutes before zpool create
- Set
ashift=12explicitly unless you have measured a reason for another value. - Name devices by their
/dev/disk/by-id/path, never/dev/vdb. - Set
compression=lz4orzstdat the pool root with-O, and leavededupoff. - Decide the vdev layout for the pool you will have in two years, because
zpool addis easy andzpool removeoften is not possible at all. - Create one dataset per workload, so
recordsizecan differ per workload later. - Create zvols at the right
volblocksizethe first time, because that property has no later fix. - If the pool may ever be imported by an older ZFS release, create it with
-o compatibility=, since "by default all supported features are enabled on the new pool" and an enabled feature cannot be disabled.
That last point bites across platforms. Feature flags are one-way, so a pool created on the newest OpenZFS can refuse to import on an older release. It is worth checking when your rescue system is not the same vintage as your server, and worth checking twice when moving pools between FreeBSD and Linux.
FAQ
Can I change ashift after creating a ZFS pool?
No. ashift is fixed per vdev when that vdev is created. The pool property of the same name only sets the value for vdevs added later, and the manual says plainly that "changing this value will not modify any existing vdev, not even on disk replacement". Read what you actually have with sudo zdb -C tank | grep ashift or sudo zpool get all tank all-vdevs | grep ashift. If it is wrong, the only fix is a new pool plus a zfs send of the data into it.
Does changing recordsize rewrite my existing files?
No. "Changing the file system's recordsize affects only files created afterward; existing files are unaffected." To apply a new record size you must rewrite the data: copy the files, send the dataset into a new one, or use zfs rewrite on OpenZFS 2.3 and later. Check free space first, because any snapshot holding the old blocks means you store both versions until that snapshot is destroyed.
Can I remove a vdev I added to a ZFS pool by mistake?
Sometimes. zpool remove handles hot spare, cache, log, and mirrored or non-redundant top-level vdevs, but only when the pool holds no top-level raidz or draid vdev, all top-level vdevs share the same ashift, and every encryption key is loaded. A single raidz vdev in the pool makes all of them permanent. Run zpool add -n before every add so you see the resulting layout before it exists.
Should I turn on ZFS dedup on a small VPS?
In almost every case, no. Plan on "at least 1.25 GiB of RAM per 1 TiB of storage", which a 4 GB VPS cannot supply past a few terabytes, and the table keeps costing you after dedup=off because existing entries remain until their blocks are freed or rewritten. Use compression=lz4 or compression=zstd instead. If you believe your data is a genuine dedup case, run sudo zdb -S tank first and read the simulated ratio it prints.