Updated: 28 September 2026 · Applies to: Ollama 0.34 on Ubuntu 24.04 / 26.04 LTS with NVIDIA GPUs
If Ollama answers slowly on a GPU server, it is probably running the model on the CPU. You can see this in one command and, in most cases, fix it in a few minutes.
1. See where the model runs
- Load a model by asking it something:
ollama run llama3.2 "Say hello" - Right afterwards run:
ollama ps
- Read the PROCESSOR column:
100% GPUmeans the model sits fully in GPU memory (fast),100% CPUmeans it runs from system memory (slow), and a mix such as40%/60% CPU/GPUmeans it only partly fits into the GPU.
2. Check that the driver works
nvidia-smi
If this fails, fix the driver first: How to fix the "NVIDIA-SMI has failed" error (cannot communicate with the NVIDIA driver)?. Ollama supports NVIDIA GPUs with driver 550 or newer (570 or newer for older Pascal-generation cards).
3. Restart Ollama after driver changes
Ollama detects the GPU when it starts. After installing or repairing the driver, restart it:
sudo systemctl restart ollama
On some systems Ollama loses the GPU after a suspend and resume cycle. Reload the NVIDIA memory driver to bring it back:
sudo rmmod nvidia_uvm && sudo modprobe nvidia_uvm
4. Read what Ollama detected
journalctl -e -u ollama
Look for lines about the GPU, or for a message that no compatible GPU was found. For more detail, turn on debug logging:
sudo systemctl edit ollama
In the editor add these lines, save, and restart Ollama:
[Service] Environment="OLLAMA_DEBUG=1"
sudo systemctl restart ollama
5. The model is bigger than the GPU memory
If ollama ps shows a split between CPU and GPU, the model plus its working memory does not fit into the VRAM. Choose a smaller model or a smaller size of the same model (for example a lower parameter count), or give Ollama a smaller context window with the OLLAMA_CONTEXT_LENGTH setting. On a server with several GPUs, Ollama spreads a large model over the cards it may use; see How to choose which GPU an application uses on a multi-GPU server? for selecting cards.
6. Ollama in Docker
If you run Ollama in a container, start it with --gpus all and install the NVIDIA Container Toolkit first: How to use your NVIDIA GPU in Docker with the NVIDIA Container Toolkit?.
Keep a model loaded
Ollama unloads a model after 5 minutes without requests, so the next request has to load it again. To keep it loaded longer, set OLLAMA_KEEP_ALIVE in the service settings (see step 4 for how to edit them), for example Environment="OLLAMA_KEEP_ALIVE=1h"; a value of -1 keeps models loaded until you stop them.
Frequently asked questions
Ollama works on the GPU in a terminal, but not from my application.
Both use the same service. Check that your application calls the server you think it does (ollama ps shows what is loaded), and that no second Ollama, for example in a container, answers instead.
Does an AMD GPU work?
Ollama supports selected AMD cards with the ROCm driver. Follow Ollama's hardware-support page for the current list.
Official documentation: Ollama documentation: hardware support and FAQ.
Want your own private AI without the setup work?
- Private LLM installation: we install and configure Ollama, the GPU driver and a chat interface on your own server.
- GPU dedicated servers: NVIDIA GPU servers for AI inference and training.
- AI implementation services: private LLMs, RAG, n8n automation and API integrations built on your own servers.
Prefer a hand with the setup? Our engineers can do it for you: Hire an Expert, or use our on-demand server management.
