Updated: 28 September 2026 · Applies to: Ollama 0.34 on Ubuntu 24.04 / 26.04 LTS
Ollama runs open large language models (Llama, Mistral, Gemma, Qwen and many others) on your own server. Your prompts and documents stay on the server; nothing is sent to an outside AI service. This guide installs Ollama on Ubuntu and runs a first model. It works on a GPU server and, more slowly, on a plain VPS or dedicated server with only a CPU.
Before you start
- Ubuntu 24.04 LTS or 26.04 LTS with
curlinstalled (sudo apt install -y curl). - GPU server: the NVIDIA driver must work, so
nvidia-smiprints a table (How to install the NVIDIA driver on Ubuntu 24.04 / 26.04 for a GPU server?). Ollama needs driver 550 or newer. Without a GPU Ollama uses the CPU, which is much slower. - Free disk space: models are large, from about 2 GB for a small model to tens of GB for a big one. Check with
df -h.
Install Ollama
- Run the official install script as a user with
sudorights:curl -fsSL https://ollama.com/install.sh | sh
- Check the version and the service:
ollama -v systemctl status ollama
The service should beactive (running). Ollama starts automatically after every reboot. - Check that the API answers on the server itself:
curl http://127.0.0.1:11434
The answer isOllama is running.
Download and run a first model
- Start a small model. The first start downloads it:
ollama run llama3.2
- Type a question at the
>>>prompt and press Enter. Type/byeto leave. - Useful commands:
ollama list # models on this server and their size ollama ps # models loaded in memory right now ollama pull NAME # download a model without starting it ollama rm NAME # delete a model to free disk space
Browse the available models in the Ollama model library. The size of a model must fit into the GPU memory (VRAM) for full speed. If it is bigger, Ollama runs part of it on the CPU and answers become slow. ollama ps shows how a loaded model is split, see How to check that Ollama is using the GPU (and fix slow CPU-only answers)?.
What is running where
- The API listens on
127.0.0.1:11434, so only the server itself can reach it. To use it from other machines safely, read How to make Ollama reachable from other machines (OLLAMA_HOST) safely?. - Models are stored in
/usr/share/ollama/.ollama/models. To move them to another disk, see How to move Ollama models to another disk (OLLAMA_MODELS)?. - Logs:
journalctl -e -u ollama
Add a chat web page
For a ChatGPT-style web interface on top of Ollama, install Open WebUI: How to install Open WebUI with Docker and connect it to Ollama?.
Frequently asked questions
Does Ollama send my data to the internet?
Running local models does not send your prompts anywhere. The server only needs internet access to download models.
Which model should I choose?
Start with a small model such as llama3.2 to check that everything works, then try larger ones. Choose a model whose file size (see ollama list) is comfortably below the memory of your GPU.
Can I run Ollama on a VPS without a GPU?
Yes, small models run on the CPU, but expect slow answers. A GPU server is recommended for anything beyond testing.
Official documentation: Ollama documentation: Linux (Ollama 0.34).
Want your own private AI without the setup work?
- Private LLM installation: we install and configure Ollama, the GPU driver and a chat interface on your own server.
- GPU dedicated servers: NVIDIA GPU servers for AI inference and training.
- AI implementation services: private LLMs, RAG, n8n automation and API integrations built on your own servers.
Prefer a hand with the setup? Our engineers can do it for you: Hire an Expert, or use our on-demand server management.
