Updated: 28 September 2026 · Applies to: NVIDIA driver on Ubuntu 24.04 / 26.04 LTS
The message NVIDIA-SMI has failed because it couldn't communicate with the NVIDIA driver. Make sure that the latest NVIDIA driver is installed and running. means the driver is not loaded into the running kernel. The usual causes are a kernel update without a restart, a driver that was not built for the new kernel, Secure Boot blocking an unsigned driver, or the open-source nouveau driver taking the card. Work through the steps in this order.
1. Make sure the server sees the GPU
lspci | grep -i nvidia
If nothing is listed, the problem is not the driver: open a support ticket with this output.
2. Look at what is loaded and what the kernel says
lsmod | grep -E 'nvidia|nouveau' sudo dmesg | grep -iE 'nvrm|nvidia' | tail -20 uname -r
Note the kernel version from uname -r and any error line from dmesg; you need them in the steps below and for a support ticket.
3. Restart after a kernel update
If apt upgrade installed a new kernel and the server was not restarted, or the reverse, the driver may belong to another kernel. Restart once:
sudo reboot
Then run nvidia-smi again.
4. Rebuild the driver for the running kernel (DKMS)
A driver installed with NVIDIA's -dkms-open packages or nvidia-open is compiled for each kernel and needs the matching kernel files:
sudo apt install -y linux-headers-$(uname -r) dkms status sudo dkms autoinstall sudo reboot
dkms status should list the NVIDIA module as installed for the kernel shown by uname -r.
5. Check Secure Boot
mokutil --sb-state
If it says SecureBoot enabled and the driver was compiled locally, the kernel refuses to load it because it is not signed. Reinstall with Ubuntu's signed packages, which work with Secure Boot:
sudo ubuntu-drivers install --gpgpu sudo reboot
Ubuntu's documentation notes that a driver built through DKMS needs a Machine Owner Key (MOK) to be enrolled during the installation; the system asks for it at the next restart.
6. Disable the nouveau driver
If lsmod shows nouveau and no nvidia module, nouveau has taken the card. Block it and rebuild the boot image:
echo -e "blacklist nouveau\noptions nouveau modeset=0" | sudo tee /etc/modprobe.d/blacklist-nouveau.conf sudo update-initramfs -u sudo reboot
7. Reinstall the driver
If the steps above did not help, install the driver again as described in How to install the NVIDIA driver on Ubuntu 24.04 / 26.04 for a GPU server? and restart.
When it is a hardware problem
If dmesg contains GPU has fallen off the bus or Xid 79, the card lost its connection to the server. That points to power, cooling or the hardware itself and cannot be fixed with software. Open a support ticket and include the dmesg lines.
Frequently asked questions
It worked yesterday and today it fails. Why?
The most common reason is an automatic or manual kernel update. The new kernel needs its own build of the driver. Restart the server, then follow step 4.
Can I fix this from inside a container?
No. The driver lives on the host. Fix it on the server itself, then start the container again.
Need a GPU server, or a hand with the setup?
- GPU dedicated servers: NVIDIA GPU servers for AI training and inference; our engineers can install the driver, CUDA, PyTorch or a private LLM and hand it over ready to use.
- Private LLM installation: we install Ollama, the GPU driver and a chat interface on your own server.
Prefer a hand with the setup? Our engineers can do it for you: Hire an Expert, or use our on-demand server management.
