Updated: 28 September 2026 · Applies to: NVIDIA GPU servers on Ubuntu 24.04 / 26.04 LTS
While a model trains or answers requests, you want to know whether the GPU is busy, how much of its memory is used and how hot it runs. The nvidia-smi tool, installed with the NVIDIA driver, shows all of this. nvtop gives a live, htop-like view.
See the current state
nvidia-smi
The table shows, for each GPU: temperature, power use, memory used and total, GPU utilisation in percent, and (in the lower table) the processes that use the GPU.
Refresh the view every second
watch -n 1 nvidia-smi
Press Ctrl+C to stop.
Show only the numbers you care about
nvidia-smi --query-gpu=index,name,utilization.gpu,memory.used,memory.total,temperature.gpu,power.draw --format=csv -l 5
This prints one line per GPU every 5 seconds. To write the values to a file for later, add -f gpu-log.csv and run it in the background with nohup or in a screen session.
Live view with nvtop
sudo apt install -y nvtop nvtop
Press F10 or q to leave it.
Find out which program uses the GPU
- Look at the process table at the bottom of
nvidia-smi; note the PID. - Identify it:
ps -fp PID. - Stop it if it is stuck:
kill PID(usekill -9 PIDonly if it does not react).
For processes started inside Docker, the process table may be empty. Use docker ps to see your containers instead.
What to watch for
- Memory used close to memory total: the next request may fail with an out-of-memory error. Use a smaller model, a smaller batch size or fewer parallel requests.
- Utilisation stays near 0 % while your program runs: the program is probably running on the CPU. See How to install PyTorch with GPU (CUDA) support and test it? or How to check that Ollama is using the GPU (and fix slow CPU-only answers)?.
- Temperature that keeps climbing under load, or clocks that drop: open a support ticket with the output of
nvidia-smi -q -d TEMPERATURE,PERFORMANCE.
Frequently asked questions
Is it normal that memory stays used after my program ends?
No. Something still runs. Check the process table, or look for a container that is still up with docker ps.
Can I monitor the GPU from my own dashboard?
Yes. NVIDIA provides the DCGM exporter for Prometheus and Grafana. Our engineers can set that up for you.
Need a GPU server, or a hand with the setup?
- GPU dedicated servers: NVIDIA GPU servers for AI training and inference; our engineers can install the driver, CUDA, PyTorch or a private LLM and hand it over ready to use.
- Private LLM installation: we install Ollama, the GPU driver and a chat interface on your own server.
Prefer a hand with the setup? Our engineers can do it for you: Hire an Expert, or use our on-demand server management.
