Updated: 28 September 2026 · Applies to: Ollama 0.34 on Ubuntu 24.04 / 26.04 LTS with NVIDIA GPUs

If Ollama answers slowly on a GPU server, it is probably running the model on the CPU. You can see this in one command and, in most cases, fix it in a few minutes.

1. See where the model runs

  1. Load a model by asking it something: ollama run llama3.2 "Say hello"
  2. Right afterwards run:
    ollama ps
  3. Read the PROCESSOR column: 100% GPU means the model sits fully in GPU memory (fast), 100% CPU means it runs from system memory (slow), and a mix such as 40%/60% CPU/GPU means it only partly fits into the GPU.

2. Check that the driver works

nvidia-smi

If this fails, fix the driver first: How to fix the "NVIDIA-SMI has failed" error (cannot communicate with the NVIDIA driver)?. Ollama supports NVIDIA GPUs with driver 550 or newer (570 or newer for older Pascal-generation cards).

3. Restart Ollama after driver changes

Ollama detects the GPU when it starts. After installing or repairing the driver, restart it:

sudo systemctl restart ollama

On some systems Ollama loses the GPU after a suspend and resume cycle. Reload the NVIDIA memory driver to bring it back:

sudo rmmod nvidia_uvm && sudo modprobe nvidia_uvm

4. Read what Ollama detected

journalctl -e -u ollama

Look for lines about the GPU, or for a message that no compatible GPU was found. For more detail, turn on debug logging:

sudo systemctl edit ollama

In the editor add these lines, save, and restart Ollama:

[Service]
Environment="OLLAMA_DEBUG=1"
sudo systemctl restart ollama

5. The model is bigger than the GPU memory

If ollama ps shows a split between CPU and GPU, the model plus its working memory does not fit into the VRAM. Choose a smaller model or a smaller size of the same model (for example a lower parameter count), or give Ollama a smaller context window with the OLLAMA_CONTEXT_LENGTH setting. On a server with several GPUs, Ollama spreads a large model over the cards it may use; see How to choose which GPU an application uses on a multi-GPU server? for selecting cards.

6. Ollama in Docker

If you run Ollama in a container, start it with --gpus all and install the NVIDIA Container Toolkit first: How to use your NVIDIA GPU in Docker with the NVIDIA Container Toolkit?.

Keep a model loaded

Ollama unloads a model after 5 minutes without requests, so the next request has to load it again. To keep it loaded longer, set OLLAMA_KEEP_ALIVE in the service settings (see step 4 for how to edit them), for example Environment="OLLAMA_KEEP_ALIVE=1h"; a value of -1 keeps models loaded until you stop them.

Frequently asked questions

Ollama works on the GPU in a terminal, but not from my application.
Both use the same service. Check that your application calls the server you think it does (ollama ps shows what is loaded), and that no second Ollama, for example in a container, answers instead.

Does an AMD GPU work?
Ollama supports selected AMD cards with the ROCm driver. Follow Ollama's hardware-support page for the current list.

Official documentation: Ollama documentation: hardware support and FAQ.

Want your own private AI without the setup work?

Prefer a hand with the setup? Our engineers can do it for you: Hire an Expert, or use our on-demand server management.

Was this answer helpful? 0 Users Found This Useful (0 Votes)