Fix NVIDIA DKMS build failures after upgrade
A kernel upgrade made the NVIDIA DKMS build fail. The log DKMS leaves under /var/lib/dkms names the real cause. Here is how to read it and fix it.
Why the NVIDIA DKMS build fails after a kernel upgrade
An NVIDIA DKMS build fails after a kernel upgrade because the driver source, the kernel it is being compiled for, and the tools needed to compile it have stopped matching each other. DKMS (dynamic kernel module support) rebuilds the NVIDIA kernel module automatically every time apt installs a new kernel. When that rebuild fails, the package transaction ends with a DKMS error naming the module and the kernel version, and you are left with a new kernel that has no NVIDIA module on it.
None of this needs guessing, because DKMS writes a full compiler log for every build it attempts. Work through it in a fixed order: which kernel is running, which kernels DKMS has built for, what the first error in the log says, and then the four causes that produce nearly all of these failures on a server.
What DKMS does during a kernel upgrade
The NVIDIA driver ships as source code, not as a ready-made binary module. The nvidia-dkms-* package unpacks that source into /usr/src/nvidia-<driver-version>/ and registers it with DKMS. A hook in /etc/kernel/postinst.d/ then runs dkms autoinstall as soon as a new kernel package is unpacked. That hook compiles the module against the new kernel's headers, signs it when a machine owner key is configured, and installs it under /lib/modules/<kernel-version>/updates/dkms/.
Two consequences matter for debugging. The build runs while the new kernel is still not booted, so it depends on the headers package for that kernel and not on the kernel you are running. And the build is an ordinary compile, which means a real compiler error message is sitting on your disk right now. Most people never open it and start reinstalling drivers instead, which fixes nothing when the cause is a missing headers package.
Step 1: which kernel is running, and which kernels has DKMS built for
uname -r
ls -1 /lib/modules
dkms statusuname -r reports the running kernel. /lib/modules holds one directory per kernel whose modules are installed. dkms status lists every module version DKMS knows about, with the kernels it has built that module for and the state of each pair.
Answer these from your own output before you change anything:
- Is there a
dkms statusline for the exact kernel thatuname -rprinted? - Is there a line for the newest entry in
/lib/modules, which may not be the kernel you booted? - Do all the lines name the same driver version, or is an old driver version still registered beside a newer one?
- Which kernel has no NVIDIA line at all? That is the one whose build failed.
Step 2: find the build log DKMS wrote
DKMS keeps the working tree of the last attempt and the archived logs of previous ones. The live log for a build that just failed is here:
sudo ls -l /var/lib/dkms/nvidia/*/build/make.logThe per-kernel copy, written once a build completes, is here:
sudo ls -l /var/lib/dkms/nvidia/*/*/x86_64/log/make.logIf neither glob matches, search for whichever log is newest:
sudo find /var/lib/dkms -name make.log -printf '%T@ %p\n' | sort -n | tail -n 3Check the driver version in the path against the version in dkms status. Are they the same? A log under an old driver version is a log from a build you ran weeks ago, and reading it will send you after a cause that no longer exists.
Step 3: read the first compiler error, not the summary at the bottom
The last lines of the log are make reporting that one of its sub-commands returned non-zero. They name the stage that failed, not the reason. The reason is the first line that contains error:, several hundred lines above.
sudo grep -n -m 5 'error:' /var/lib/dkms/nvidia/<driver-version>/build/make.log
sudo less /var/lib/dkms/nvidia/<driver-version>/build/make.logRead that first match and decide which of these it is. Does it name a header file that the compiler could not open? Does it name a kernel function, a struct, or a struct field that the driver's code is using in a way the kernel no longer accepts? Does it mention space, a write failure, or a truncated file? Each of those points at a different cause, and the sections below take them in the order they occur on a server.
Cause 1: the headers for that kernel are missing or mismatched
A module is compiled against the header files of one specific kernel build. Ubuntu ships them in linux-headers-<kernel-version>, and /lib/modules/<kernel-version>/build is a symlink into that package. If the symlink points at a directory that is not there, the compiler cannot open a single kernel header, so the build stops on the first #include it meets.
readlink -f /lib/modules/$(uname -r)/build
test -d /lib/modules/$(uname -r)/build && echo present || echo missing
dpkg -l | grep -E 'linux-(image|headers)'Does the symlink resolve to a directory that exists? Does the dpkg -l list show a linux-headers package whose version string matches the linux-image package for the failing kernel? If either answer is no, install the matching headers and the compiler:
sudo apt install linux-headers-$(uname -r)
sudo apt install build-essential dkmsThen fix the reason it went missing. Headers arrive automatically only when a headers metapackage is installed alongside the image metapackage. If linux-image-generic is installed but linux-headers-generic is not, every future kernel will arrive without headers and every future build will fail the same way. Install the headers metapackage whose name mirrors the image metapackage you actually have. That pairing is also where hardware enablement kernels catch people out, because the generic and the HWE series are separate metapackages that move at different speeds, which is worth understanding in full through the difference between the HWE and GA kernels on Ubuntu Server.
Cause 2: the driver is too old for the new kernel
The Linux kernel makes no stability promise about its internal interfaces. Functions get renamed, arguments get added, and struct fields move between releases. A driver released before a kernel cannot know about any of that, so the compile stops at the first call the new kernel no longer offers in the old shape. This is the cause when the first error: in the log names a kernel symbol rather than a missing file: an unknown type, an implicit declaration, or a call with the wrong number of arguments.
No amount of reinstalling the same driver will change this. The driver branch has to move forward.
ubuntu-drivers devices
ubuntu-drivers list
apt-cache search '^nvidia-driver-[0-9]'ubuntu-drivers devices names the branch Ubuntu recommends for the GPU in this machine. On a headless server the -server variants are usually the right ones, since they track the branches NVIDIA supports for datacenter cards. Install the branch you chose, then check that DKMS registered it:
sudo apt install nvidia-driver-<branch>-server
dkms status -m nvidiaIf the old driver version is still listed after the new one installs, remove it so future kernel upgrades stop trying to build it. The --all flag removes that version for every kernel, so read the version string twice before you press enter:
sudo dkms remove -m nvidia -v <old-driver-version> --allCause 3: the new kernel is installed but you have not booted it
The build runs at package install time, against a kernel you are not running yet. So a failure here can be reported for a kernel that uname -r does not mention, and a reboot alone sometimes clears it: when a transaction was interrupted part way, the headers can land after the build hook has already run.
uname -r
ls -1 /boot/vmlinuz-*
sudo dpkg --configure -a
sudo apt --fix-broken installIs the newest vmlinuz in /boot the same version that uname -r printed? If it is not, you are running an older kernel than the one that failed. Finish the transaction with the two commands above, reboot, and then let DKMS try again for the kernel you are now on:
sudo dkms autoinstall -k $(uname -r)If the box has to keep working while you sort the driver out, the old kernel is still installed and still has its module. Boot that one deliberately rather than hoping the menu picks it, which is what choosing which kernel a VPS boots by default covers.
Cause 4: the disk was too full to finish the build
Compiling the NVIDIA module writes thousands of object files under /var/lib/dkms, and the final install step writes a module and a new initramfs into /boot. When a write fails, the compiler stops and the first error in the log is about space or a broken write rather than about any line of code.
df -h /var /boot /usr /tmp
df -i /varIs there at least a gigabyte free on the filesystem holding /var? Does /boot have room for another initramfs, which is often over 100 MB? Has /var run out of inodes even though it reports free space? A /boot partition filled by old kernels is the most common version of this on a long-lived server, and clearing it is a routine job described in removing old kernels safely on Ubuntu. If the free space figure looks wrong compared with what you can find by hand, the reason is usually a deleted file still held open by a process, which the gap between df and du on a full disk explains.
Rebuild for one kernel without reinstalling the driver
Once you believe the cause is fixed, do not reinstall the driver package to test it. Ask DKMS to build for exactly one kernel. The build step writes a fresh log and reproduces the failure in seconds, with no apt transaction in the way.
sudo dkms build -m nvidia -v "<driver-version>" -k "<kernel-version>"
sudo dkms install -m nvidia -v "<driver-version>" -k "<kernel-version>"
dkms status -m nvidiaDKMS 3.x also accepts the shorter nvidia/<driver-version> form in place of the -m and -v pair. To rebuild everything DKMS knows about for one kernel, use sudo dkms autoinstall -k <kernel-version>. After a successful install, reboot into that kernel and confirm the module is loaded with lsmod | grep nvidia and nvidia-smi.
Hold the kernel or the driver until the pair matches
Sometimes the working pair is the pair you already had, and the newest kernel in the archive is simply ahead of the newest driver your distribution packages. You can stop apt from moving one half of the pair:
apt-mark showhold
sudo apt-mark hold linux-image-generic linux-headers-generic
sudo apt-mark unhold linux-image-generic linux-headers-genericHold the kernel metapackages when you are already on the newest available driver. Hold the driver package instead when you have a working combination and want kernel security updates to keep arriving. A held kernel stops receiving security fixes, so treat that hold as temporary, write down the date you set it, and check for a newer driver branch every few weeks. Holding the kernel and forgetting about it is a worse outcome than a broken driver.
A build failure is not a load failure
These three problems look similar from the console and have nothing in common underneath, so confirm which one you have before you start fixing it.
If dkms status shows the module installed for your running kernel and sudo modprobe nvidia still refuses, the build succeeded. On a machine with Secure Boot enabled, the kernel checks the module's signature and refuses an unsigned or wrongly signed one, which is a key enrolment problem covered in why the kernel rejects the NVIDIA module key under Secure Boot.
If the module loads and nvidia-smi complains that the driver and library versions differ, the module built correctly but the user-space libraries came from a different driver version, which is the case handled in fixing the NVML driver and library version mismatch. For the wider set of symptoms where a GPU simply does not appear after a boot, work through why the NVIDIA driver fails to load on Ubuntu Server, and when you need the kernel's own account of what happened at boot, reading the kernel log on a VPS shows where those messages are kept.
FAQ
Where does DKMS put the build log for a failed NVIDIA build?
The live log of the last attempt is at /var/lib/dkms/nvidia/<driver-version>/build/make.log, and the archived per-kernel copy is at /var/lib/dkms/nvidia/<driver-version>/<kernel-version>/x86_64/log/make.log. DKMS names the path in the error it reports during the failing apt transaction. If you have already lost that output, find the newest one with sudo find /var/lib/dkms -name make.log -printf '%T@ %p\n' | sort -n | tail -n 3. Read the first line in it containing error:, not the last lines, which only report that the make step returned non-zero.
Do I need to reinstall the NVIDIA driver after every kernel update?
No. DKMS exists so that you do not have to: the kernel package's post-install hook rebuilds the module for each new kernel on its own. Reinstalling the same driver version after a failed build changes nothing, because the source it compiles is identical. Install a newer driver only when the first compiler error in the log names a kernel function or struct, which is the signal that the driver predates the kernel's internal interface.
Which linux-headers package do I need on Ubuntu Server?
For the kernel you are running right now, sudo apt install linux-headers-$(uname -r). For every kernel you install in future, you also need the headers metapackage matching your image metapackage, so that headers arrive in the same transaction as the kernel. Run dpkg -l | grep -E 'linux-(image|headers)' and compare the two lists: if a linux-image metapackage has no linux-headers counterpart with the same suffix, that is why the build failed.
How do I stop the next kernel update from breaking the driver again?
After any kernel upgrade, run dkms status before you reboot and confirm there is an NVIDIA line for the newly installed kernel version. If there is not, you can fix it while the working kernel is still running. When your driver branch is behind the kernel series your system tracks, hold the kernel metapackages with sudo apt-mark hold until a newer driver is available, and record the date, because held packages stop receiving security updates.