Skip to content

Run Claude Code with Ollama on Your Own VPS

Updated Aug 2026

We earn commissions when you shop through the links below. Full disclosure →

Point Claude Code at a self-hosted Ollama server — run open-source coding models like Qwen3-Coder with no API bill, using Ollama's Anthropic-compatible API.

Before you start
  • Ollama 0.15 or newer, installed locally or on a VPS (see Deploy Ollama on a VPS)
  • Claude Code installed on your workstation
  • Enough RAM/VRAM for a tool-calling coding model (~9 GB for the smallest good option)
Need a box for this guide? Kamatera's free tier lets you spin one up now.Start free on Kamatera → (opens in new tab)

Why run Claude Code against your own models?

Claude Code is a terminal coding agent, and by default it talks to Anthropic's API. Since early 2026, Ollama ships an Anthropic-compatible /v1/messages endpoint — which means Claude Code can drive open-source models running on hardware you control. No per-token bill, no code leaving your infrastructure.

The trade-off is honest: local models are noticeably weaker than the hosted Claude models at multi-step agentic work. For sensitive codebases, air-gapped environments, or high-volume grunt work (test generation, refactors, docs), a self-hosted backend is a real option. For hard architectural work, expect a quality gap.

Pick a model that can call tools

Claude Code leans hard on tool calling — the model must support it, and it needs context headroom. Ollama's own integration docs recommend at least 32K context, ideally 64K or more.

Models we'd actually point Claude Code at, all tool-calling capable:

  • glm-4.7-flash — the lightest option that behaves well in agent loops; runs in roughly 9 GB of memory. Start here on modest hardware.
  • qwen3-coder:30b — the sweet spot for dedicated coding work. It's a ~19 GB download and wants 24 GB of VRAM (or unified memory) to run comfortably.
  • deepseek-r1:32b — strong reasoning; needs about 20 GB of VRAM.

On a CPU-only VPS these run, but slowly — agent loops make many model calls, so tokens-per-second matters more here than in chat. See the sizing section of our Ollama deployment guide for RAM math.

The one-command path: ollama launch claude

Recent Ollama releases include a launcher that wires everything for you:

ollama launch claude

It configures Claude Code to talk to your local Ollama and starts it. Pin a model explicitly with:

ollama launch claude --model qwen3-coder:30b

That's the whole setup when Ollama runs on the same machine as Claude Code.

Where to host itaffiliate disclosure
Hetzner Cloudrun it on
2 vCPU · 4 GB RAM · 80 GB SSD · $23.59/mo
Get Hetzner Cloud (opens in new tab)
Kamaterafree trial
1 vCPU · 1 GB RAM · 20 GB SSD · $4.00/mo
Start free on Kamatera → (opens in new tab)
DigitalOceanalso works on
1 vCPU · 1 GB RAM · 25 GB SSD · $6.00/mo
Deploy on DigitalOcean → (opens in new tab)

Paid link — we earn a commission if you shop through it.

The manual path: point Claude Code at a VPS

If Ollama runs on a VPS (the setup this site cares about), configure Claude Code by environment variable:

export ANTHROPIC_AUTH_TOKEN=ollama   # required by the CLI, ignored by Ollama
export ANTHROPIC_BASE_URL=http://<your-vps-tunnel>:11434
claude --model qwen3-coder:30b

ANTHROPIC_AUTH_TOKEN can be any non-empty value — Ollama doesn't check it. ANTHROPIC_BASE_URL is the Ollama server's address; no /v1 suffix needed. There's no /login flow in this mode.

Do not expose port 11434 to the public internet. Ollama has no built-in authentication — anyone who can reach the port can run your models. Reach the VPS over a private channel instead:

# SSH tunnel: your laptop's localhost:11434 → the VPS's Ollama
ssh -N -L 11434:localhost:11434 you@your-vps

Then use ANTHROPIC_BASE_URL=http://localhost:11434 as if it were local. A mesh VPN like Headscale or NetBird does the same job permanently, and a reverse proxy with basic auth in front of Ollama works if you must serve multiple users.

What won't work like the hosted API

Ollama's Anthropic-compatibility layer is deliberately scoped. The documented gaps that matter for Claude Code:

  • No prompt caching. Every request re-processes the system prompt and conversation history, so long sessions get slower and burn more compute than the hosted API would.
  • No forced tool choice. The tool_choice parameter isn't supported; occasionally a model will answer in prose when Claude Code expected a tool call. Better models misfire less.
  • Approximate token counts. The token-counting endpoint isn't implemented, so context-window bookkeeping is an estimate.
  • Images must be base64. URL-referenced images aren't fetched.

None of these break the core loop — edit, run, test, commit works. They're the reason the experience trails the hosted models even before model quality enters the picture.

Cost math

A VPS that runs glm-4.7-flash around the clock costs a flat monthly price — compare that against a metered API bill that scales with every agent loop. If you're running batch agentic workloads daily, the crossover comes fast; if you use Claude Code an hour a week, the hosted API is cheaper and better. Our Ollama VPS guide has current provider picks sized for each model tier.

Next steps

How to self-host Ollama →More self-hosted ai chat interfaces tools →Best VPS for Ollama →Automatic HTTPS with Caddy →Deploy Coolify on a VPS →How to Deploy Actual Budget on a VPS →How to Deploy AnythingLLM on a VPS →How to Deploy Appwrite on a VPS →How to Deploy Audiobookshelf on a VPS →How to Deploy Authelia on a VPS →How to Deploy authentik on a VPS →How to Deploy Baserow on a VPS →How to Deploy Beszel on a VPS →How to Deploy Bitwarden on a VPS →How to Deploy BookStack on a VPS →How to Deploy CapRover on a VPS →How to Deploy Checkmate on a VPS →How to Deploy Directus on a VPS →How to Deploy docker-mailserver on a VPS →How to Deploy Docmost on a VPS →How to Deploy Dokku on a VPS →How to Deploy Dokploy on a VPS →How to Deploy Firefly III on a VPS →How to Deploy Forgejo on a VPS →How to Deploy Gatus on a VPS →How to Deploy Ghostfolio on a VPS →How to Deploy Gitea on a VPS →How to Deploy GitLab on a VPS →How to Deploy GlitchTip on a VPS →How to Deploy Grafana on a VPS →How to Deploy Graylog on a VPS →How to Deploy Headscale on a VPS →How to Deploy Healthchecks on a VPS →How to Deploy Home Assistant on a VPS →How to Deploy Immich on a VPS →How to Deploy Jan on a VPS →How to Deploy Jellyfin on a VPS →How to Deploy Karakeep on a VPS →How to Deploy Keycloak on a VPS →How to Deploy Leantime on a VPS →How to Deploy LibreChat on a VPS →How to Deploy Linkwarden on a VPS →How to Deploy LocalAI on a VPS →How to Deploy Mailcow on a VPS →How to Deploy Mailu on a VPS →How to Deploy Matomo on a VPS →How to Deploy Mattermost on a VPS →How to Deploy Meilisearch on a VPS →How to Deploy Memos on a VPS →How to Deploy n8n on a VPS →How to Deploy Navidrome on a VPS →How to Deploy NetBird on a VPS →How to Deploy Netdata on a VPS →How to Deploy Nextcloud on a VPS →How to Deploy Next.js to a VPS →How to Deploy Nginx Proxy Manager on a VPS →How to Deploy NocoDB on a VPS →How to Deploy ntfy on a VPS →How to Deploy Ollama on a VPS →How to Deploy Open WebUI on a VPS →How to Deploy OpenHands on a VPS →How to Deploy OpenObserve on a VPS →How to Deploy OpenProject on a VPS →How to Deploy Outline on a VPS →How to Deploy Pangolin on a VPS →How to Deploy Paperless-ngx on a VPS →How to Deploy Passbolt on a VPS →How to Deploy Plane on a VPS →How to Deploy Plausible Analytics on a VPS →How to Deploy Pocket ID on a VPS →How to Deploy PocketBase on a VPS →How to Deploy Prometheus on a VPS →How to Deploy Psono on a VPS →How to Deploy Radarr on a VPS →How to Deploy Rocket.Chat on a VPS →How to Deploy SigNoz on a VPS →How to Deploy Sonarr on a VPS →How to Deploy Stalwart on a VPS →How to Deploy Stirling-PDF on a VPS →How to Deploy Supabase on a VPS →How to Deploy Synapse on a VPS →How to Deploy Taiga on a VPS →How to Deploy TeamPass on a VPS →How to Deploy Tinyauth on a VPS →How to Deploy Traefik on a VPS →How to Deploy Trilium on a VPS →How to Deploy Twenty CRM on a VPS →How to Deploy Umami on a VPS →How to Deploy Uptime Kuma on a VPS →How to Deploy Vaultwarden on a VPS →How to Deploy Vikunja on a VPS →How to Deploy wg-easy on a VPS →How to Deploy Wiki.js on a VPS →How to Deploy Zabbix on a VPS →How to Deploy Zitadel on a VPS →How to Deploy Zulip on a VPS →Docker & Compose on Ubuntu 26.04 →Building AI Workflows with n8n →Install Open WebUI with Ollama →Adding AI-Powered Insights to Plausible Analytics →Building AI-Powered Apps with Supabase and pgvector →

We use analytics cookies (Google Analytics, PostHog) to see which guides are useful. No ad networks, no cross-site tracking. See our privacy policy.