Updated: 28 September 2026 · Applies to: NVIDIA GPU servers on Ubuntu 24.04 / 26.04 LTS

Servers with several GPUs (for example four A100 or seven A40 cards) run every program on all visible GPUs unless you limit it. Limiting is useful to run several jobs side by side, or to keep a card free for another user. The number of a GPU is the same across most tools once you fix the ordering.

1. List your GPUs

nvidia-smi -L

The output looks like GPU 0: NVIDIA A100 ... (UUID: GPU-xxxxxxxx-...). Note the number and the UUID of the card you want.

2. Make CUDA use the same numbers as nvidia-smi

By default CUDA can number the cards by speed, which may differ from the order in nvidia-smi. If your cards are of different models, set this once so that the numbers match:

export CUDA_DEVICE_ORDER=PCI_BUS_ID

Add the line to ~/.bashrc to keep it for future logins. With identical cards you can skip this step.

3. Limit a program to some GPUs

CUDA_VISIBLE_DEVICES=0,1 python3 train.py

The program only sees GPU 0 and GPU 1, and inside the program they are numbered 0 and 1. You can also use UUIDs: CUDA_VISIBLE_DEVICES=GPU-xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx.

4. Limit a Docker container

docker run --rm --gpus '"device=0,1"' your-image

This needs the NVIDIA Container Toolkit (How to use your NVIDIA GPU in Docker with the NVIDIA Container Toolkit?).

5. Limit Ollama

Ollama reads the same variable. Open its service settings and add the line, then restart it:

sudo systemctl edit ollama

In the editor add:

[Service]
Environment="CUDA_VISIBLE_DEVICES=0,1"
sudo systemctl restart ollama

Ollama's documentation recommends UUIDs, because numeric IDs can change order. To force Ollama to use only the CPU, set an invalid ID such as -1.

Check the link speed between GPUs

nvidia-smi topo -m

The matrix shows how each pair of GPUs is connected (for example NV for NVLink or PHB and SYS for connections through the motherboard). Training jobs that split one model over several GPUs run best on cards with fast links.

Frequently asked questions

Do I need to reboot after changing these settings?
No. They are read when a program starts. For a service such as Ollama, restart the service.

The wrong GPU is used although I set CUDA_VISIBLE_DEVICES.
Set CUDA_DEVICE_ORDER=PCI_BUS_ID (step 2) or use the UUID instead of the number.

Need a GPU server, or a hand with the setup?

  • GPU dedicated servers: NVIDIA GPU servers for AI training and inference; our engineers can install the driver, CUDA, PyTorch or a private LLM and hand it over ready to use.
  • Private LLM installation: we install Ollama, the GPU driver and a chat interface on your own server.

Prefer a hand with the setup? Our engineers can do it for you: Hire an Expert, or use our on-demand server management.

Was this answer helpful? 0 Users Found This Useful (0 Votes)