Skip to content

Self-host vLLM

updated Aug 2026prices checked · Jul 2026
We earn commissions when you shop through the links below. Full disclosure →

High-throughput LLM inference and serving engine with PagedAttention and continuous batching — an OpenAI-compatible API server built for GPU-backed, multi-user production workloads.

Key facts

LicenseApache-2.0
StackPython, CUDA
Min RAM16384 MB
Official imageyes
Difficulty
Our recommendation

Reach for vLLM when a GPU server needs to serve many people at once — continuous batching and PagedAttention keep one card ahead of a whole team's chat traffic, behind the same OpenAI-compatible API your clients already speak. If it's just you, or your VPS has no GPU, Ollama does the job with a tenth of the setup.

What you need

  • Any VPS with at least 16384 MB of RAM
  • A domain you control — most self-hosted setups need HTTPS in front of them
  • About an afternoon — budget time for troubleshooting
Where to host itaffiliate disclosure
Hetzner Cloudrun it on
From $23.59/mo · 2 vCPU / 4 GB / 80 GB · EU + US
Get Hetzner Cloud (opens in new tab)
Kamaterafree trial
From $4/mo · 1 vCPU / 1 GB / 20 GB · US + EU + Asia
Start free on Kamatera → (opens in new tab)
DigitalOceanalso works on
From $6/mo · 1 vCPU / 1 GB / 25 GB · US + EU + Asia
Deploy on DigitalOcean → (opens in new tab)

Paid link — we earn a commission if you shop through it.

Install

Run these commands on your server:

# vLLM — official image needs an NVIDIA GPU (CPU builds exist but are a separate image)
docker run -d --runtime nvidia --gpus all -v ~/.cache/huggingface:/root/.cache/huggingface \
  -p 8000:8000 --ipc=host vllm/vllm-openai:latest --model Qwen/Qwen3-0.6B

What you take on

vLLM is production-grade serving infrastructure, and it asks you to operate like it:

non-negotiableIt effectively requires a GPU, NVIDIA first. The official vllm/vllm-openai image assumes CUDA (AMD gets a separate vllm/vllm-openai-rocm image) — you'll install the drivers and the NVIDIA container toolkit before anything serves, and the CPU builds exist mostly for development, not the throughput vLLM is chosen for.
non-negotiableVRAM math comes first. vLLM loads full Hugging Face weights, so the model claims most of your card before the clever cache management helps — size the GPU to the model you actually want to serve, or startup will fail with an out-of-memory error, not a warning.
non-negotiableIt moves fast. Releases land frequently and flags shift between them — pin an exact image tag rather than latest, and re-read the server flags on every upgrade instead of assuming yours still mean the same thing.

An alternative to

Head-to-head

More in AI Chat Interfaces

We use analytics cookies (Google Analytics, PostHog) to see which guides are useful. No ad networks, no cross-site tracking. See our privacy policy.