Fix NVML driver/library version mismatch
nvidia-smi fails with Driver/library version mismatch when an upgrade replaces the NVIDIA libraries under a loaded kernel module. Diagnose it, then fix it.
Why nvidia-smi reports a driver/library version mismatch
nvidia-smi prints Failed to initialize NVML: Driver/library version mismatch because the NVIDIA kernel module in memory and the NVML (NVIDIA management library) file on disk come from two different driver releases. NVML compares its own version string against the version the loaded module reports, and it stops when the two differ. The card is not involved in that check. An upgrade replaced the userspace files while the old module stayed loaded, so the two halves of one driver no longer agree.
Recent drivers print the number they found, which shortens the diagnosis:
Failed to initialize NVML: Driver/library version mismatch
NVML library version: 560.35A kernel module is loaded once and stays in memory until something unloads it or the machine reboots. apt can replace nvidia.ko and libnvidia-ml.so.560.35 on disk in under a second, but it cannot swap the module that is already running and holding GPU state. So every fix has the same shape: get the old modules out of memory, then let the new ones load from disk.
Compare the loaded module with the installed driver
Two commands answer the question. The first reads the version out of the running module, the second reads the version of the module file on disk.
cat /proc/driver/nvidia/version
modinfo -F version nvidia/proc/driver/nvidia/version exists only while the module is loaded, and it prints a line like NVRM version: NVIDIA UNIX x86_64 Kernel Module 550.120 Wed Jun 10 09:12:44 UTC 2026. That number is the driver that is running right now. modinfo -F version nvidia prints the version of the module file the kernel would load next time. Two different numbers confirm the mismatch, and the higher number is almost always the one on disk, because the upgrade moved forward while memory stayed behind.
Check the userspace half as well, since that is the side NVML actually loads:
ls -l /usr/lib/x86_64-linux-gnu/libnvidia-ml.so*
dpkg -l | grep -E 'nvidia-(driver|dkms|utils)|libnvidia-compute'libnvidia-ml.so.1 is a symlink to a versioned file, and that version should equal what modinfo -F version nvidia printed. If the symlink is missing, or points at a file that no longer exists, the upgrade itself did not finish. Repair that first with sudo dpkg --configure -a and sudo apt -f install, because no amount of module reloading fixes a half-installed package set.
Then find out when the change happened:
grep -B1 -A3 -i nvidia /var/log/apt/history.log
sudo grep -i nvidia /var/log/unattended-upgrades/unattended-upgrades.logAn Upgrade: line naming libnvidia-compute-560 or nvidia-dkms-560 with today's timestamp is the cause in writing: apt changed the driver under a running module. If that line appears only in the unattended-upgrades log, nobody on your team did it by hand, and the section on unattended-upgrades below is the part you want.
Find what is holding the nvidia modules open
You cannot unload a module that something is using, so list the holders before you try.
lsmod | grep nvidiaThe third column is a use count and the fourth is a list of module names. nvidia_uvm, nvidia_drm and nvidia_modeset all sit on top of nvidia, so nvidia almost always shows a non-zero count. nvidia_uvm (unified virtual memory) is loaded by CUDA work. nvidia_modeset and nvidia_drm come from the display path.
Processes holding the device files are the other half of the answer:
sudo fuser -v /dev/nvidia*
sudo lsof /dev/nvidia*On a server the usual holders are nvidia-persistenced, a CUDA job such as a training script or a local model server, an ffmpeg process doing NVENC transcoding, a container started with --gpus all, and a display manager on a box that has a desktop installed. Stop them by name. Killing processes at random on a GPU host tends to leave half a job behind.
sudo systemctl stop nvidia-persistenced
docker psAny container in that list that was given the GPU counts as a holder. Stop it with docker stop <name> before you touch the modules.
Unload and reload the nvidia modules
Unload the children first, then the base module. modprobe -r works out the dependency order for you, which is why it is easier than rmmod.
sudo modprobe -r nvidia_uvm nvidia_drm nvidia_modeset nvidia
lsmod | grep nvidiaNo output from lsmod means memory is clear. Now load the current module and read the versions again:
sudo modprobe nvidia
cat /proc/driver/nvidia/version
nvidia-sminvidia-smi also loads the module on demand through nvidia-modprobe, and the first CUDA program loads nvidia_uvm the same way, so you do not have to load every module by hand. The driver version in the nvidia-smi header should now match modinfo -F version nvidia. That agreement is the whole fix.
Two failures are common at this step. The first:
rmmod: ERROR: Module nvidia is in use by: nvidia_uvm nvidia_modesetA child module is still loaded, so unload that child before the base module. The second:
rmmod: ERROR: Module nvidia_drm is in useSomething is using the DRM (direct rendering manager) interface. On a machine booted with nvidia_drm.modeset=1 the console framebuffer holds it, and on a box with a desktop the display manager does. Stop the display manager with sudo systemctl stop gdm3, or drop to a text target with sudo systemctl isolate multi-user.target, then unload again.
When a reboot is simply the answer
Reboot when the console framebuffer holds nvidia_drm, when the workload on the card cannot be stopped cleanly right now, or when you do not know what is holding the module and the machine is not serving traffic. A reboot loads every module fresh from disk, so it ends the mismatch every time. Unloading by hand is worth the effort only when uptime matters more than the ten minutes you may spend on it.
After the reboot, run nvidia-smi once and confirm the header version matches modinfo -F version nvidia. If the mismatch comes back after a clean reboot, memory was never the problem: the module files on disk are inconsistent, which is the next section.
The same error after a kernel upgrade, and what DKMS did
DKMS (dynamic kernel module support) rebuilds nvidia.ko for every kernel you install. A new kernel arrives, DKMS builds the driver against it, and the next boot loads that build.
uname -r
dkms status
modinfo -n nvidiadkms status should list your driver against the running kernel, for example nvidia/560.35, 6.8.0-79-generic, x86_64: installed. A kernel with no entry has no module at all, and then nvidia-smi fails with a message about not being able to communicate with the driver instead of a version mismatch. That is a different problem: see the NVIDIA driver failing to load on Ubuntu Server.
A mismatch that survives a reboot usually means two copies of the module sit on the module path, one built by DKMS under updates/dkms/ and an older one from a package that was never removed.
find /lib/modules/$(uname -r) -name 'nvidia*.ko*'
modinfo -n nvidiamodinfo -n nvidia prints the exact file the kernel will load. If that path is the stale copy, remove the stale file, rebuild the module index with sudo depmod -a, and reload. If DKMS never built for the running kernel, build it by hand using the version string from dkms status:
sudo dkms install nvidia/560.35 -k $(uname -r)Fewer kernel trees on disk means fewer chances for a stale module to win, so clearing out old kernels is worth doing on a GPU host. One more trap lives here. A freshly built module has to be signed on a machine with Secure Boot enabled, and an unsigned one is refused at load time with Key was rejected by service, which is the Secure Boot signature failure rather than a version problem.
Containers see the mismatch after the host upgrades
The NVIDIA Container Toolkit mounts the host driver libraries into the container when the container starts. A container that is already running keeps the files it was given. Upgrade the driver on the host and nvidia-smi inside that container reports a mismatch against the host module, even after the host itself is healthy again.
Restart the container so the toolkit mounts the current host libraries, substituting your own container and service names:
docker restart jellyfin
docker compose up -d --force-recreate jellyfinIf a plain restart does not clear it, recreate the container, which rebuilds the whole mount set from scratch. The difference between those two operations is worth knowing before you need it: restart and rebuild do not replace the same things. Media servers show this in a specific way, where playback still works but every GPU transcode fails right after a host driver upgrade, so check the container before you re-tune Jellyfin hardware transcoding on an NVIDIA card.
Stop unattended-upgrades from doing this under a running job
The driver is one package set spanning kernel and userspace, so an automatic upgrade of it always creates this mismatch until the next reload. On a machine that runs GPU jobs unattended, take the driver out of the automatic path and upgrade it yourself in a window where a reboot is fine. Add the two package prefixes to /etc/apt/apt.conf.d/50unattended-upgrades:
Unattended-Upgrade::Package-Blacklist {
"nvidia-";
"libnvidia-";
};Those entries are regular expressions matched from the start of the package name, so the two prefixes cover the driver metapackage, the DKMS package and the compute libraries. Confirm the result before you trust it:
sudo unattended-upgrade --dry-run --debug 2>&1 | grep -i -e blacklist -e nvidiaA simpler option is an apt hold, which blocks a manual upgrade too:
sudo apt-mark hold nvidia-driver-560 nvidia-dkms-560
apt-mark showholdEither way you now own those updates. The driver runs inside the kernel, so leaving it pinned for a year is a security decision rather than a tidy one: put it in the routine you already use for checking your server for known CVEs, and release the hold with sudo apt-mark unhold when you upgrade on purpose.
Check this before you blame the card
The mismatch is decided by comparing two version strings, one from a library file and one from a loaded module. The GPU is never asked, so this message can never be a symptom of failing hardware. Real hardware trouble looks different. lspci | grep -i nvidia printing nothing means the device is not attached, and on a rented VPS with a GPU that points at the passthrough device rather than at your driver. sudo dmesg | grep -iE 'nvrm|xid' prints Xid lines when the GPU itself reports a fault. A different message, NVIDIA-SMI has failed because it couldn't communicate with the NVIDIA driver, means no module is loaded at all.
Fix it in order
- Read both versions:
cat /proc/driver/nvidia/versionandmodinfo -F version nvidia. - Confirm the cause in
/var/log/apt/history.logor the unattended-upgrades log. - List holders with
lsmod | grep nvidiaandsudo fuser -v /dev/nvidia*. - Stop those holders, then run
sudo modprobe -r nvidia_uvm nvidia_drm nvidia_modeset nvidia. - Run
nvidia-smi. If the modules refuse to unload, reboot instead of fighting them. - If the error survives a reboot, check
dkms statusandmodinfo -n nvidiafor a stale module file.
FAQ
Can I fix Failed to initialize NVML: Driver/library version mismatch without rebooting?
Yes, when you can get every user off the GPU. Stop the processes holding /dev/nvidia*, stop nvidia-persistenced, stop any container that was given the card, then run sudo modprobe -r nvidia_uvm nvidia_drm nvidia_modeset nvidia. lsmod | grep nvidia returning nothing means the old driver is out of memory, and the next nvidia-smi loads the new module and prints a matching version. If nvidia_drm will not unload because the console framebuffer is using it, a reboot is the shorter path.
Why does rmmod say the nvidia module is in use when nothing is running?
rmmod: ERROR: Module nvidia is in use by: nvidia_uvm nvidia_modeset is not about your processes. It means other kernel modules are stacked on top of nvidia and have to go first. Unload nvidia_uvm, nvidia_drm and nvidia_modeset, or list them all in one modprobe -r command so the order is handled for you. nvidia-persistenced also keeps the driver loaded on purpose, and it is a service, so scanning ps output for your own jobs will not reveal it.
Does this error mean my GPU is broken?
No. NVML compares the version string of its own library file with the version reported by the loaded kernel module, and it fails before it touches the hardware. To rule out a real device problem, run lspci | grep -i nvidia to confirm the card is attached and sudo dmesg | grep -iE 'nvrm|xid' to look for Xid faults reported by the GPU itself.
How do I keep unattended-upgrades from upgrading the NVIDIA driver?
Add "nvidia-"; and "libnvidia-"; to the Unattended-Upgrade::Package-Blacklist block in /etc/apt/apt.conf.d/50unattended-upgrades, then verify with sudo unattended-upgrade --dry-run --debug. An apt-mark hold on the driver and DKMS packages works too, and it also stops a manual apt upgrade from surprising you. Both choices move driver patching onto your schedule, so pair them with a reminder to upgrade and reboot on purpose.