You Fit Self-Host Claude? The Honest Answer
Claude weights no dey public, so your server no fit run am. See wetin you fit self-host instead: open models, API gateway, or Claude Code.
You fit self-host Claude? No, and na why
You no fit self-host Claude. Anthropic no publish the model weights, so no file dey to download, no container dey to run, and no licence dey wey go allow you serve am from your own hardware. Every Claude request dey go Anthropic's API or hosted partner like Amazon Bedrock, Google Vertex AI, or Microsoft Foundry. To run am for machine wey you own no be configuration problem. The artefact simply no dey outside Anthropic.
Na the short answer be that. The longer answer be say most people wey dey ask this question no really want the weights. Dem want one of three things wey all dey possible for server wey you control: capable model wey dey run locally, gateway wey hold their API keys and limit their spending, or coding agent wey dey live for their own box instead of their laptop. This guide cover all three, with the commands.
Wetin “self-hosted Claude” usually mean
Search traffic for “self hosted Claude” dey split into some different wants, and each one need different answer.
Some people want privacy. Dem no want prompts commot from their network. Na only local open weight model fit solve that, because any Claude request by definition na request go Anthropic.
Some people want cost control. Dem dey worry say runaway agent fit finish credits. Gateway fit solve that, and e work with Claude, so you still keep the model quality.
Some people want independence from laptop. Dem want agent wey go continue work while dem close the lid. VPS fit solve that, and Claude Code dey run for am without wahala.
Some people want the phrase “self hosted OpenRouter”. That one still na gateway, and the usual answer na LiteLLM.
Find out which one be your own, because the correct build different for each case.
Self-host open model with Ollama
If you need make sure say no prompt dey leave your server, run open weight model. The model families wey fit work well for rented server today na Llama, Qwen, Mistral, Gemma, and DeepSeek. All of dem publish weights wey you fit download and run.
Ollama na the fastest way to start. The install script na one line, and e go set up systemd service for Ubuntu.
curl -fsSL https://ollama.com/install.sh | sh
systemctl status ollamasystemctl status ollama suppose print active (running). Then pull model and talk to am.
ollama pull qwen3:8b
ollama run qwen3:8b "Summarise what a reverse proxy does in two sentences."The first pull go download several gigabytes, so the model must fit inside RAM or GPU memory before e fit answer anything. For quantised models, rough rule be say: 8 billion parameter model need about 6 GB free, 14 billion parameter model need about 10 GB, while 70 billion parameter model need more memory than most general purpose VPS plans get. If the machine memory no reach, kernel go kill the process and you go see Error: llama runner process has terminated, with out of memory line inside dmesg. Check free -h before you blame the model. The same memory budget also decide how much long prompt the model actually read, because Ollama quietly truncate anything wey pass modest default window. So increase num_ctx and size the KV cache na the first thing to check when long documents return half summarised.
Ollama also serve HTTP API on 127.0.0.1:11434. Na this one make e useful to other software, instead of only being chat toy.
curl http://127.0.0.1:11434/api/generate -d '{"model":"qwen3:8b","prompt":"ping","stream":false}'If the first request after quiet period take thirty seconds, while the next one return immediately, nothing spoil: Ollama unload the model after five minutes of idling. Keep am resident with keep_alive go remove that reload delay.
Leave that port bound to localhost. Ollama port wey open for public IP na free GPU for anybody wey find am. The full setup, including systemd unit, GPU detection, and putting reverse proxy in front, dey covered for guide for running Ollama on VPS. If you dey serve more than one user at the same time, read comparison between Ollama and vLLM first, because Ollama single stream design go become bottleneck well before the hardware reach limit.
Be honest about the gap. Good open model for mid sized VPS fit genuinely help with summarising, classifying, drafting, and simple extraction. For long multi step reasoning, large codebases, and agentic tool use, e no dey close to frontier hosted model, and no amount of prompt tuning fit close that gap. Choose local model for the work wey e good at, and pay for hosted one where the difficulty dey real.
LiteLLM wey go be your own gateway
Na this be the "self hosted OpenRouter" wey people dey search for. Gateway dey between your applications and every model provider. Your apps go hold one key, wey point to your server. Na only that server go keep the real provider keys. You fit set spending limit for each key, route different apps go different models, and log every request for one place.
LiteLLM na the common choice because e dey speak OpenAI compatible API and e dey proxy to Anthropic, Ollama, and most other providers through the same endpoint. Run am for Docker with config file.
model_list:
- model_name: claude
litellm_params:
model: anthropic/claude-sonnet-5
api_key: os.environ/ANTHROPIC_API_KEY
- model_name: local
litellm_params:
model: ollama/qwen3:8b
api_base: http://127.0.0.1:11434Save am as litellm_config.yaml and start the proxy. E dey listen on port 4000.
docker run -v $(pwd)/litellm_config.yaml:/app/config.yaml \
-e ANTHROPIC_API_KEY=$ANTHROPIC_API_KEY \
-e LITELLM_MASTER_KEY=sk-1234 \
-p 4000:4000 docker.litellm.ai/berriai/litellm:latest \
--config /app/config.yamlLITELLM_MASTER_KEY na the admin credential, so treat am like root password and no use the example value for production. Call the proxy exactly as you go call hosted API.
curl http://localhost:4000/v1/chat/completions \
-H 'Authorization: Bearer sk-1234' \
-H 'Content-Type: application/json' \
-d '{"model": "claude","messages": [{"role": "user","content": "Say hello in five words."}]}'Healthy response na normal JSON wey get choices array. 401 mean say Authorization header no match your master key. 400 wey name the model mean say model for your request no match any model_name for the config file.
The reason to build this instead of calling Anthropic directly na the spending limit. Issue one separate virtual key for each application, and give each key e own budget.
curl 'http://0.0.0.0:4000/key/generate' \
--header 'Authorization: Bearer sk-1234' \
--header 'Content-Type: application/json' \
--data-raw '{"models": ["claude"], "max_budget": 100}'That key fit spend one hundred dollars and reach one model, and nothing else. If agent misbehave for three in the morning, the blast radius na one key instead of your whole account. That pattern, together with the monitoring around am, na the subject of how to control agent costs for VPS. If you still dey decide whether to pay per token at all, API versus subscription cost comparison go help you work through the arithmetic.
Notice wetin the gateway no dey do. E no make Claude local, and e no hide your prompts from Anthropic. Requests still dey leave your server go the provider. Wetin you gain na control over keys, spending, routing, and logs.
Run Claude Code for your own VPS
The third wish na di easiest one to grant. Claude Code na client. E dey run anywhere wey you install Node.js, and e dey talk to the API through HTTPS. If you put am for server wey you own, the agent go continue to work after you shut your laptop. E also mean say the agent's blast radius na box wey you fit rebuild, instead of your main machine.
npm install -g @anthropic-ai/claude-code
claude --versionRun am inside tmux so SSH connection wey drop no go kill long job. The setup, including how to manage the session, dey explained for how to run Claude Code for VPS with tmux. Give the agent im own unprivileged user. Before you give am write access to anything wey matter to you, read the safety rules for running Claude Code for server.
This na self-hosting the agent, no be the model. E good make we clear about this, because na here people dey mix things up. You own the process, the filesystem, the network egress, and the logs. Anthropic still own the inference.
Wetin each option really go cost you
Prices dey change, so see these figures as general guide, no be fixed quote. As of July 2026, Claude Sonnet 5 dey list for $3 per million input tokens and $15 per million output tokens, while Claude Opus 5 na $5 and $25. Local model no cost anything per token. Instead, e go cost whatever the server costs per month, whether you use am or not.
The break-even point dey lower than many people expect. VPS wey get enough memory to run useful open model costs real money every month, and e dey idle most of the time. If your usage dey come in bursts, hosted API usually cheaper. If your usage constant, or your data no fit leave your network, local model win for both reasons.
The honest mixed answer na wetin most teams finally choose. Run open model locally for high-volume, low-difficulty work. Route hard requests go hosted frontier model. Put gateway in front of both, so applications no need know which one dem dey use. This also let you move the boundary between dem without touching application code. That architecture na the practical version of "self hosted Claude". Unlike the literal version, e actually dey exist. If you also want run the whole agent stack by yourself, this roundup of self-hosted AI agents cover wetin dey available.
FAQ
I fit download Claude model weights and run dem locally?
No. Anthropic never release weights for any Claude model, and no licence dey allow self-hosting. Anything wey dem advertise online as downloadable "Claude model" either na different model wey get misleading name, or wrapper wey dey call the API. If e need API key, e no dey run locally.
Which open model dey closest to Claude?
No exact match dey, and the leading models dey change every few months. The open-weight families wey worth testing na Llama, Qwen, Mistral, Gemma, and DeepSeek. For summarising, classification, and simple code edits, good open model with 8 to 14 billion parameters fit really help. For long multi-step reasoning and agentic tool use, the gap between am and hosted frontier model still big. Test am with your own prompts instead of trusting leaderboard.
LiteLLM na self-hosted OpenRouter?
Functionally yes, for routing and key management. LiteLLM dey run for your server, e dey provide one OpenAI-compatible endpoint, and e dey proxy requests to Anthropic, Ollama, and most other providers. You get per-key spend caps, model routing, and one place to read logs. But e no give you local inference: requests to Claude still dey travel go Anthropic.
If I run Claude Code for my own server, my code go remain private?
No. Claude Code dey send the file contents wey e read to Anthropic API, no matter where the process dey run. Wetin VPS give you na isolation for the agent, not privacy for the content. Give am dedicated unprivileged user, keep am away from credentials and unrelated repositories, and treat everything wey e fit read as content wey go leave the server.