GPU Dedicated Servers: How to Choose by Workload and GPU Memory

How to choose a GPU dedicated server for AI inference, training, rendering or video transcoding: GPU memory maths, NVENC support and what else to check.

Short answer: a GPU dedicated server is a physical server with one or more graphics cards that are yours alone. Choose it by workload. Running an AI model needs enough GPU memory to hold that model. Training needs several times more. Rendering needs a card your renderer supports. Video transcoding needs hardware video encoders, which some data-centre GPUs do not have. Then check the CPU, RAM, storage, bandwidth and software support around the card.

Ucartz GPU dedicated servers page
Our GPU dedicated servers page.

What a GPU dedicated server is

A GPU does thousands of small calculations in parallel, which suits neural networks and 3D rendering, and many NVIDIA cards also have a separate hardware video encoder. In a dedicated server the whole machine, GPUs included, belongs to you: full root access, no virtualisation layer and no other customers on the same cards. That makes performance predictable and lets you install any driver, CUDA version or framework you need.

Not every AI project needs a GPU on day one. Small or quantised models can run on CPUs at a lower cost, as we explain in CPU vs GPU VPS for Hugging Face models.

Match the GPU to the workload

AI inference: GPU memory decides what fits

To answer requests, a model’s weights must sit in GPU memory. The weights alone take the number of parameters times the bytes per parameter:

  • 16-bit precision: 2 bytes per parameter, so a 7-billion-parameter model needs about 14 GB.
  • 8-bit: 1 byte per parameter, about 7 GB for the same model.
  • 4-bit: half a byte per parameter, about 3.5 GB.

On top of the weights you need room for the context of each conversation and for the runtime itself, and that grows with longer prompts and more users at once. As rough starting points for 4-bit models, we size 7B to 8B models at 8 GB of GPU memory or more, 13B to 14B at 12 to 16 GB, and 70B-class models at 40 GB or more. A 70B model at 16-bit precision needs about 140 GB for the weights alone, which is more than one 80 GB card.

Training and fine-tuning: several times more memory

Training keeps much more than the weights in memory. According to Hugging Face’s documentation on training memory, mixed-precision training with the Adam optimiser uses about 6 bytes per parameter for the weights, 8 bytes for the optimiser states and 4 bytes for the gradients: 18 bytes per parameter before activations, which grow with batch size and sequence length. For a 1-billion-parameter model that is already about 18 GB; for 7 billion it is about 126 GB. Full training of larger models is therefore a multi-GPU job.

With several GPUs, how they talk to each other matters. NVIDIA quotes 600 GB/s for an NVLink bridge between two A100 cards, against 64 GB/s for PCIe Gen4. Ask whether the cards in a multi-GPU server are bridged or communicate over PCIe only.

3D rendering

Check that your renderer supports the card. The Blender manual, for example, lists CUDA and OptiX for NVIDIA cards with compute capability 5.0 or higher, and notes that OptiX uses the ray-tracing hardware in RTX cards. Professional cards such as the RTX A5000 and the data-centre A40 also have ECC (error-correcting) memory, according to NVIDIA’s specifications.

Video transcoding: look for NVENC

GPU transcoding uses the card’s dedicated video encoder, called NVENC on NVIDIA cards. Not every GPU has one. NVIDIA’s encode and decode support matrix shows that the A100 has no NVENC at all, so a card built for AI can be the wrong choice for a streaming pipeline.

GPUGPU memoryNVENC encodersGood to know
A100 80GB80 GB HBM2e, about 2 TB/sNoneCan be split into up to seven isolated GPU instances (MIG)
A4048 GB GDDR6 with ECC1Two cards can be linked with NVLink
RTX A500024 GB GDDR6 with ECC1Workstation card for rendering and mid-size models
GeForce RTX 409024 GB GDDR6X2AV1 encoding; up to 12 encode sessions at once
GeForce RTX 509032 GB GDDR73AV1 encoding; up to 12 encode sessions at once
Figures from NVIDIA’s product specifications and its video encode and decode support matrix.

FFmpeg uses NVENC through its h264_nvenc, hevc_nvenc and av1_nvenc encoders. To see whether your FFmpeg build includes them, run:

ffmpeg -hide_banner -encoders | grep nvenc

With FFmpeg 9.0 we got:

 V....D av1_nvenc            NVIDIA NVENC av1 encoder (codec av1)
 V....D h264_nvenc           NVIDIA NVENC H.264 encoder (codec h264)
 V....D hevc_nvenc           NVIDIA NVENC hevc encoder (codec hevc)

The encoders appearing in the list only means FFmpeg was built with support; they work once the server has an NVENC-capable card and the NVIDIA driver. The FFmpeg documentation covers the encoder options.

What to check besides the GPU

  • GPU count and memory per card. Two 24 GB cards are not one 48 GB card: a model that needs 40 GB must be split across them, which not all software supports.
  • CPU cores and system RAM. Data loading, pre-processing and tokenising run on the CPU and use system RAM. A strong GPU waiting on a weak CPU is wasted money.
  • NVMe storage. Model files and datasets are large, and they load faster from NVMe than from SATA disks.
  • Bandwidth and port speed. Downloading model weights or moving datasets can mean hundreds of gigabytes. Check the monthly transfer and the port speed.
  • Software setup. Who installs the NVIDIA driver, CUDA, PyTorch or your inference server, and what it costs.
  • Delivery time and billing. GPU servers are real hardware. Check how long delivery takes and whether billing is monthly or longer.
  • Hardware failure. Ask how quickly a failed GPU, disk or memory module is replaced.

First checks when your GPU server is ready

Once you can log in over SSH, confirm that the operating system and the driver can see every card:

lspci | grep -i nvidia
nvidia-smi
nvidia-smi -L

Then work through the setup in order:

  1. Check that the server can see its NVIDIA GPU and what the results mean.
  2. Install the NVIDIA driver on Ubuntu if nvidia-smi is missing.
  3. Install PyTorch with CUDA support and test it, or use the NVIDIA Container Toolkit with Docker.
  4. Monitor GPU use, memory and temperature while your first real job runs.

GPU dedicated servers at Ucartz

Our GPU dedicated servers are bare-metal machines with NVIDIA A100, A40, RTX A5000, RTX 4090 and RTX 5090 cards, from a single GPU up to seven cards in one server. Every plan has NVMe storage, full root access, free setup and Free Basic Managed Support. The page lists the delivery time for each plan, typically 24 working hours. If you want the driver, CUDA, PyTorch or a private LLM installed and handed over ready to use, our engineers can do it as a paid task, and we are happy to help you size the server before you order.

Ashily Shaji
Ashily Shaji