Updated: 28 September 2026 · Applies to: Ollama 0.34 on Ubuntu 24.04 / 26.04 LTS

Ollama runs open large language models (Llama, Mistral, Gemma, Qwen and many others) on your own server. Your prompts and documents stay on the server; nothing is sent to an outside AI service. This guide installs Ollama on Ubuntu and runs a first model. It works on a GPU server and, more slowly, on a plain VPS or dedicated server with only a CPU.

Before you start

  1. Ubuntu 24.04 LTS or 26.04 LTS with curl installed (sudo apt install -y curl).
  2. GPU server: the NVIDIA driver must work, so nvidia-smi prints a table (How to install the NVIDIA driver on Ubuntu 24.04 / 26.04 for a GPU server?). Ollama needs driver 550 or newer. Without a GPU Ollama uses the CPU, which is much slower.
  3. Free disk space: models are large, from about 2 GB for a small model to tens of GB for a big one. Check with df -h.

Install Ollama

  1. Run the official install script as a user with sudo rights:
    curl -fsSL https://ollama.com/install.sh | sh
  2. Check the version and the service:
    ollama -v
    systemctl status ollama
    The service should be active (running). Ollama starts automatically after every reboot.
  3. Check that the API answers on the server itself:
    curl http://127.0.0.1:11434
    The answer is Ollama is running.

Download and run a first model

  1. Start a small model. The first start downloads it:
    ollama run llama3.2
  2. Type a question at the >>> prompt and press Enter. Type /bye to leave.
  3. Useful commands:
    ollama list        # models on this server and their size
    ollama ps          # models loaded in memory right now
    ollama pull NAME   # download a model without starting it
    ollama rm NAME     # delete a model to free disk space

Browse the available models in the Ollama model library. The size of a model must fit into the GPU memory (VRAM) for full speed. If it is bigger, Ollama runs part of it on the CPU and answers become slow. ollama ps shows how a loaded model is split, see How to check that Ollama is using the GPU (and fix slow CPU-only answers)?.

What is running where

Add a chat web page

For a ChatGPT-style web interface on top of Ollama, install Open WebUI: How to install Open WebUI with Docker and connect it to Ollama?.

Frequently asked questions

Does Ollama send my data to the internet?
Running local models does not send your prompts anywhere. The server only needs internet access to download models.

Which model should I choose?
Start with a small model such as llama3.2 to check that everything works, then try larger ones. Choose a model whose file size (see ollama list) is comfortably below the memory of your GPU.

Can I run Ollama on a VPS without a GPU?
Yes, small models run on the CPU, but expect slow answers. A GPU server is recommended for anything beyond testing.

Official documentation: Ollama documentation: Linux (Ollama 0.34).

Want your own private AI without the setup work?

Prefer a hand with the setup? Our engineers can do it for you: Hire an Expert, or use our on-demand server management.

Was this answer helpful? 0 Users Found This Useful (0 Votes)