Run Ollama cloud models from a small VPS
Ollama on a small no-GPU VPS with cloud models doing the inference: signin without a browser, a private API on 11434, quotas, and a local fallback.
What you are building
You can run Ollama cloud models from a small VPS with no GPU, because the box never loads the weights. Ollama runs on the server as an ordinary service on port 11434, and any model name ending in :cloud sends the actual work to Ollama's hosted hardware. The VPS holds the API endpoint and the account credential. Nothing heavy lands on it.
The result is one private endpoint at http://127.0.0.1:11434 that your apps, your coding agent and Open WebUI can all talk to, backed by a model far larger than that plan could ever hold in memory. Every command, model name and limit below was checked on 30 September 2026 against Ollama v0.35.0, released 28 September 2026, on Ubuntu 24.04.
Install Ollama and run it as a service
curl -fsSL https://ollama.com/install.sh | sh
systemctl is-active ollama
curl -s http://127.0.0.1:11434/api/versionsystemctl is-active ollama should print active, and /api/version should return a JSON object holding the version string. The install script creates a system user named ollama whose home is /usr/share/ollama, and a systemd unit that starts on boot. The full walkthrough, including firewall setup and what each unit file setting does, is in the guide to self-hosting an LLM with Ollama on a VPS, so this post only covers what changes when the inference is not local.
One setting matters here more than anywhere else. The daemon's default bind address is 127.0.0.1:11434, loopback only. Leave it alone. The reason is in the security section below.
Sign the server in to ollama.com without a browser
Ask for a cloud model before the server has an account attached and Ollama tells you so directly:
You need to be signed in to Ollama to run Cloud models.The fix is ollama signin. On a desktop it opens a browser tab. A VPS has no browser to open, so the command prints the URL and waits:
ollama signinIf your browser did not open, navigate to:
<a one time URL on ollama.com>Copy that URL into the browser on your own machine, sign in to your ollama.com account, and approve the request. The ollama signin process on the server finishes on its own once you approve. Run it a second time and it answers You are already signed in as user '<your account>', which is the check that it worked.
Which credential ends up on the server
The credential is an Ed25519 key pair, not a password and not an API key. Ollama generates it the first time the daemon starts. Ollama's own FAQ gives the public half on Linux as /usr/share/ollama/.ollama/id_ed25519.pub, and the private half sits beside it as /usr/share/ollama/.ollama/id_ed25519. Both belong to the ollama service user, which is why the path is under that user's home and not under yours.
sudo ls -l /usr/share/ollama/.ollama/Signing in pairs that public key with your ollama.com account. The private key stays on the VPS forever and is what proves the machine to Ollama's cloud, so anyone who can read that file can spend your account's usage credits. Keep it out of your backups' plain-text tier, and do not copy an image of this server around as a template.
If the browser approval cannot happen at all, there is a manual route. Print the public key and add it to your account under ollama.com/settings/keys:
sudo cat /usr/share/ollama/.ollama/id_ed25519.pubOLLAMA_API_KEY is a separate mechanism, and confusing the two costs people an afternoon. An API key authenticates requests you send straight to https://ollama.com/api/chat with an Authorization: Bearer header, from a script that does not run Ollama at all. The local daemon never reads it: OLLAMA_API_KEY is not one of the variables Ollama's environment configuration knows about, so setting it in the unit file with systemctl edit ollama.service has no effect on :cloud models. Sign the daemon in instead.
To take the access away later, run ollama signout, which answers You have signed out of ollama.com, then revoke the key in your account settings so a restored backup of the disk cannot use it again.
Run a cloud model, and confirm nothing large was downloaded
ollama run gemma4:cloud "Reply with one short sentence about disk latency."
ollama ls
du -sh /usr/share/ollama/.ollama/modelsTokens stream back within a second or two on a box that could not hold the model. The Ollama cloud docs state it plainly: cloud models do not need to be downloaded. ollama ls shows the evidence, because a cloud entry prints a single - in the SIZE column instead of a byte count. There is no blob on disk to measure, so there is no size to report, and du on the models directory barely moves. If you want the full picture of when Ollama does write gigabytes to your disk and where, the difference between pull and run and where model files actually live covers the local case this one avoids.
One naming detail trips up scripts. In the CLI and the local API you use the cloud tag, gemma4:cloud. For requests you send directly to ollama.com you use the name the registry returns, such as gemma4:31b. Same model, two names, chosen by which endpoint you are calling.
Which cloud models can you run today?
curl -s https://ollama.com/api/tagsThat is the documented way to list what the cloud currently serves, and the same set is browsable at Ollama's cloud model list. On 30 September 2026 it held sixteen entries, including gemma4, glm-5.3, glm-5.3-flash, deepseek-v4.1-flash, deepseek-v4-pro, kimi-k3, kimi-k2.7-code, minimax-m3, nemotron-3-ultra, mistral-large-3 and gpt-oss. Treat that as a snapshot of the day, not a fixed menu, and read the list from the API in anything automated.
Each model page also shows a usage level, which is how fast that model spends your allowance. gemma4:cloud is labelled Low Usage. kimi-k2.7-code:cloud is labelled High Usage. Picking a lighter model for bulk work and a heavy one for the hard questions is the whole of cost control here. Deciding which of these you would rather host yourself is a different question, and which models a server you own can realistically hold answers it from the hardware side.
Keep port 11434 private, because it now spends your account
The local Ollama API has no authentication. Any client that can reach the socket is trusted, and in OpenAI-compatible requests Ollama accepts and ignores whatever API key you send. On a normal local-model server that is merely bad. Here it is worse, because every request that reaches port 11434 is now billable capacity on your ollama.com account. Publish it and a scanner spends your credits.
So keep the default bind. Do not set OLLAMA_HOST=0.0.0.0:11434 on this box, keep ufw default deny incoming in place, and reach the API through an SSH tunnel from your own machine:
ssh -N -L 11434:127.0.0.1:11434 you@your-vps
curl -s http://127.0.0.1:11434/api/tagsWith the tunnel up, http://127.0.0.1:11434 on your laptop is the server's loopback port, and nothing is listening on a public interface. A private network interface between your own servers works the same way, with the bind address set to that private address rather than to 0.0.0.0. The full list of ways this endpoint gets left open and what port 11434 actually exposes are worth reading before you decide a reverse proxy is simpler.
Point a coding agent and Open WebUI at it
With the tunnel open, any OpenAI-compatible client works against http://127.0.0.1:11434/v1 with a placeholder key, since the key is ignored locally:
curl -s http://127.0.0.1:11434/v1/chat/completions \
-H 'Content-Type: application/json' \
-H 'Authorization: Bearer ollama' \
-d '{"model":"kimi-k2.7-code:cloud","messages":[{"role":"user","content":"Write a bash one liner that counts open TCP sockets."}]}'A coding agent needs the same three values: the base URL, a placeholder key, and the model name with its :cloud suffix. Wiring Ollama into a coding agent covers the per-tool configuration and the tool-calling support you need to check per model.
Open WebUI runs on the same box in Docker and reads the Ollama address from OLLAMA_BASE_URL:
docker run -d --network=host \
-e OLLAMA_BASE_URL=http://127.0.0.1:11434 \
-e WEBUI_SECRET_KEY="$(openssl rand -hex 32)" \
-v open-webui:/app/backend/data \
--name open-webui --restart always \
ghcr.io/open-webui/open-webui:main--network=host is doing real work in that command. A container on the default bridge network cannot reach a daemon bound to the host's loopback address, so the usual -p 3000:8080 form would fail with a connection refused in the Open WebUI logs. Sharing the host network stack means 127.0.0.1:11434 inside the container is the same socket as outside it.
The trade is that Open WebUI then listens on port 8080 of the host itself. That is still safe here, and the reason is a Docker detail worth knowing: a port published with -p is reached through Docker's own nat rules, which run before ufw's input rules and therefore ignore them, while a host-network container's port goes through the normal input chain and ufw does block it. So tunnel that one too:
ssh -N -L 8080:127.0.0.1:8080 you@your-vpsOpen http://127.0.0.1:8080, create the first account, and the cloud models appear in the model picker like local ones. If Open WebUI is heavier than you want on a 1 GB box, the lighter chat front ends for a VPS all speak the same API.
How small can the VPS be?
Size the plan from what the box is actually doing, which is almost nothing. It accepts a TCP connection, parses JSON, opens a TLS connection to ollama.com, and copies a token stream back to the client. That is I/O and a little JSON handling. No weights are read from disk, no KV cache is allocated, no GPU is involved, and the usual sizing rule of roughly one gigabyte of RAM per billion parameters does not apply because no parameters are here. This is the whole reason a small plan is enough, and it is also the whole trade: what you give up by letting Ollama's cloud do the inference is the other half of the decision.
One vCPU and 1 GB of RAM covers the relay itself. Two things push that up. The first is concurrency, since each open stream is a connection the daemon holds while it waits. The second is any local model you keep alongside, which is real RAM and real disk. If you pull the fallback model described below, plan on 2 GB of RAM and a few gigabytes of free disk.
OLLAMA_KEEP_ALIVE is irrelevant on this server. It controls how long a model stays loaded in your memory, and a cloud model is never loaded into your memory at all, so there is nothing to keep alive. Keeping a model resident between requests only matters once you are running the weights yourself. CPU architecture barely matters either, which makes this one of the few AI workloads where picking an ARM VPS over x86 costs you nothing.
The honest limits, as of September 2026
Four limits apply, and each one has a page you should read rather than trust a blog about.
Usage is metered per plan. Every plan carries a monthly allowance of usage credits spent per token, plus a cap on how many requests can run at once. The concurrency cap is the one that shapes your server design, because it decides whether two people can use your endpoint at the same moment.
The data behind this chart
[
{
"plan": "Free",
"concurrent_requests": 1
},
{
"plan": "Pro",
"concurrent_requests": 3
},
{
"plan": "Max",
"concurrent_requests": 10
},
{
"plan": "Team",
"concurrent_requests": 10
}
]Those are the published figures for the 4 plans listed on the Ollama plans page on 30 September 2026. A free account gets 1 concurrent request, which means your coding agent and your Open WebUI tab will queue behind each other. The top plans allow 10. Read the credit allowances off that page rather than from any article, including this one, and watch what you have spent under ollama.com/settings/usage.
Your prompts leave the server. This is the part people skip. The text your app sends, and the text the model returns, travel to Ollama's infrastructure and are processed there. Ollama's cloud documentation states that it processes cloud prompts and responses to answer your requests, and that it does not use them to train models, with the details in its privacy policy. That is a reasonable position and it is still a different position from a model running on hardware you rent. If the data cannot leave your control, this design is the wrong one.
Context length differs from the same model run locally. Ollama's context length documentation says cloud models are set to their maximum context length by default, while a local model defaults by available VRAM, which is 4k of context under 24 GiB. A GPU-less VPS is firmly in that bracket. So gemma4:cloud starts with the model's full window, and gemma4:31b pulled onto a small local box would start at 4k until you raise OLLAMA_CONTEXT_LENGTH. Published windows for cloud models are large: kimi-k2.7-code:cloud advertises 256K and glm-5.3-flash:cloud advertises 1M as of 30 September 2026. How Ollama picks a context length and how to change it explains why a prompt gets silently truncated when you get this wrong. Long generations bring a second problem, since a client that waits a long time for the first token can give up before Ollama answers, which is the case the context deadline exceeded error describes.
Cloud models are retired on published dates. Ollama's usage settings show upcoming retirements for models you have recently used, and the instruction is to switch models before the retirement date. Models you downloaded locally are not affected. In practice that means a model name hard-coded in your app is a dated dependency. Keep the name in one config value, not scattered through your code, and check that settings page when you review the bill.
Keep one small local model as a fallback
When the allowance runs out, when the key is revoked, or when the VPS cannot reach ollama.com, a cloud request returns an error instead of text. Your app sees an HTTP failure, not a slow answer. The cheapest insurance is one small model that lives on the disk you are paying for:
ollama pull gemma3:1b
ollama run gemma3:1b "Reply with the single word ok."
ollama lsgemma3:1b is listed at 815MB on ollama.com, so it fits the same small plan and runs on CPU. It will not match a frontier cloud model on anything hard. That is not the job. The job is that your service degrades to a worse answer instead of to a stack trace, and that you keep a working endpoint while you sort out billing or a retired model name. In ollama ls you will now see both kinds side by side: the local model with a real size, the cloud model with a -. That one column is the clearest summary of this whole setup.
FAQ
Do I need a GPU to run Ollama cloud models on a VPS?
No. The model weights never load on your server, so there is nothing for a GPU to accelerate. Your VPS opens a TLS connection to ollama.com and relays a token stream, which is I/O work. One vCPU and 1 GB of RAM is enough for the relay, and you only need more if you also keep a local model pulled as a fallback.
Where does the ollama.com credential live on my server?
In an Ed25519 key pair that the daemon generates on first start. Ollama's FAQ documents the public half on Linux as /usr/share/ollama/.ollama/id_ed25519.pub, with the private half beside it, both owned by the ollama service user. ollama signin pairs that public key with your account, so the private key on the server is what can spend your usage credits. OLLAMA_API_KEY is a separate credential for calling https://ollama.com/api/chat directly, and the local daemon does not read it.
Does a cloud model take up disk space on my VPS?
No, and you can prove it. Ollama's cloud documentation says cloud models do not need to be downloaded, and ollama ls prints a single - in the SIZE column for a cloud entry because there is no blob on disk to measure. Check du -sh /usr/share/ollama/.ollama/models before and after your first cloud run and the number will not meaningfully change.
Can I expose port 11434 so my other machines can use it?
Not to the internet. The local Ollama API has no authentication and ignores any API key an OpenAI-compatible client sends, so whoever reaches the port gets to spend your ollama.com allowance. Keep the default 127.0.0.1:11434 bind and reach it over ssh -N -L 11434:127.0.0.1:11434 you@your-vps, or bind it to a private network address that only your own servers can route to.
Why is a cloud model's context length different from the same model run locally?
Because the defaults are set from different facts. Ollama's documentation says cloud models are set to their maximum context length by default, while a local model's default comes from available VRAM and is 4k under 24 GiB, which covers every GPU-less VPS. So the cloud copy accepts a long prompt that the local copy of the same model would truncate until you raise OLLAMA_CONTEXT_LENGTH and give the box the memory to back it.