Updated: 28 September 2026 · Applies to: NVIDIA GPU servers on Ubuntu 24.04 / 26.04 LTS
Servers with several GPUs (for example four A100 or seven A40 cards) run every program on all visible GPUs unless you limit it. Limiting is useful to run several jobs side by side, or to keep a card free for another user. The number of a GPU is the same across most tools once you fix the ordering.
1. List your GPUs
nvidia-smi -L
The output looks like GPU 0: NVIDIA A100 ... (UUID: GPU-xxxxxxxx-...). Note the number and the UUID of the card you want.
2. Make CUDA use the same numbers as nvidia-smi
By default CUDA can number the cards by speed, which may differ from the order in nvidia-smi. If your cards are of different models, set this once so that the numbers match:
export CUDA_DEVICE_ORDER=PCI_BUS_ID
Add the line to ~/.bashrc to keep it for future logins. With identical cards you can skip this step.
3. Limit a program to some GPUs
CUDA_VISIBLE_DEVICES=0,1 python3 train.py
The program only sees GPU 0 and GPU 1, and inside the program they are numbered 0 and 1. You can also use UUIDs: CUDA_VISIBLE_DEVICES=GPU-xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx.
4. Limit a Docker container
docker run --rm --gpus '"device=0,1"' your-image
This needs the NVIDIA Container Toolkit (How to use your NVIDIA GPU in Docker with the NVIDIA Container Toolkit?).
5. Limit Ollama
Ollama reads the same variable. Open its service settings and add the line, then restart it:
sudo systemctl edit ollama
In the editor add:
[Service] Environment="CUDA_VISIBLE_DEVICES=0,1"
sudo systemctl restart ollama
Ollama's documentation recommends UUIDs, because numeric IDs can change order. To force Ollama to use only the CPU, set an invalid ID such as -1.
Check the link speed between GPUs
nvidia-smi topo -m
The matrix shows how each pair of GPUs is connected (for example NV for NVLink or PHB and SYS for connections through the motherboard). Training jobs that split one model over several GPUs run best on cards with fast links.
Frequently asked questions
Do I need to reboot after changing these settings?
No. They are read when a program starts. For a service such as Ollama, restart the service.
The wrong GPU is used although I set CUDA_VISIBLE_DEVICES.
Set CUDA_DEVICE_ORDER=PCI_BUS_ID (step 2) or use the UUID instead of the number.
Need a GPU server, or a hand with the setup?
- GPU dedicated servers: NVIDIA GPU servers for AI training and inference; our engineers can install the driver, CUDA, PyTorch or a private LLM and hand it over ready to use.
- Private LLM installation: we install Ollama, the GPU driver and a chat interface on your own server.
Prefer a hand with the setup? Our engineers can do it for you: Hire an Expert, or use our on-demand server management.
