SSD Nodes Learn Hosting plans →
Guides Matt ConnorBy Matt Connor

Is Ollama free? Local use vs Ollama Cloud

Ollama is MIT licensed and free to run on hardware you own. What costs money is Ollama Cloud and the machine with enough RAM to hold the model.

Is Ollama free?

Yes, Ollama is free: the software is MIT licensed, and running models on your own hardware is unmetered. What costs money is Ollama Cloud, the hosted service that runs models on Ollama's GPUs, and the machine you run Ollama on yourself.

So the real question is which bill you want to pay. Local use turns the cost into hardware: enough RAM to hold the model, and a processor fast enough to make it useful. Cloud use turns the cost into a subscription plus per-token credits. This guide puts numbers on the cloud side and names the costs on the local side that no pricing page shows you.

What does Ollama do?

Ollama is a local model runner. It downloads open-weight models (models whose trained weights are published, so anyone can run them), loads them into memory, and serves them over an HTTP API on port 11434. The ollama command line talks to that same API, and so do editors and coding agents. Ollama does not train models and does not make its own. It packages models from other labs, mostly in the GGUF file format, and runs them with an inference engine that uses the ggml library from the llama.cpp project. If you want to know what the wrapper adds on top of that engine, Ollama versus llama.cpp covers the trade.

On Linux the install is one line from the official README:

curl -fsSL https://ollama.com/install.sh | sh
ollama --version

ollama --version should print a version number. A warning that it could not connect to a running instance means the ollama systemd service did not start, and journalctl -u ollama -n 50 shows why.

Ollama vs ChatGPT

ChatGPT is a hosted assistant. It is a product with a chat interface, accounts, web search and a proprietary model that runs on OpenAI's servers. Ollama is an engine plus an API. Its desktop app has a simple chat window, but the model is whichever open model you pull, and the quality you get depends on that model and on the hardware that runs it. A 4-billion-parameter model on a laptop CPU will not match a frontier hosted model on hard reasoning tasks. What you get in exchange is control: your prompts stay on your machine, and the model does not change under you unless you pull a new version. If you are weighing a monthly subscription against this, replacing a paid Claude plan with an open model works through where the quality gap is real and where it is not.

Ollama Cloud pricing as of October 2026

Ollama Cloud runs larger models on Ollama's own hardware. You reach them through the same CLI and API as a local model, after you sign in with ollama signin. Cloud models carry a cloud tag, such as gpt-oss:120b-cloud, and otherwise behave like any other model name.

ChartOllama Cloud plans, read from ollama.com/pricing on 2026-10-02
The data behind this chart
[
  {
    "plan": "Free",
    "price_usd_month": 0,
    "included_credits_usd": 0,
    "concurrent_requests": 1
  },
  {
    "plan": "Pro",
    "price_usd_month": 20,
    "included_credits_usd": 60,
    "concurrent_requests": 3
  },
  {
    "plan": "Max",
    "price_usd_month": 100,
    "included_credits_usd": 300,
    "concurrent_requests": 10
  },
  {
    "plan": "Team (early access)",
    "price_usd_month": 500,
    "included_credits_usd": "1,000",
    "concurrent_requests": 10
  }
]

These figures were read from the Ollama pricing page on 2026-10-02. Pricing pages change often, so check the page again before you rely on any number here.

The Free plan costs $0 and allows 1 concurrent request. Its credit is a small starter allowance on a smaller set of starter models. The pricing page gives no dollar figure for that allowance, so the chart records 0 dollars of fixed credit. You can add paid credits to unlock the other models without taking a subscription.

Pro costs $20 a month, or $200 a year billed annually, and includes $60 of usage each month with 3 concurrent requests. Max costs $100 a month for $300 of usage and 10 concurrent requests. Team is in early access at $500 a month, with $1,000 of usage shared across unlimited users. An Enterprise plan with custom pricing also exists, which is why it is not in the chart.

The credit model is new. Ollama moved Pro, Max and Team to per-token pricing on 2026-08-31. Before that date, cloud usage was limited by time windows (a 5-hour session limit and a weekly limit) rather than by a dollar pool. Now each plan's included usage resets monthly on the day your subscription started. Unused credit does not roll over. When the pool runs out, you keep going at the published per-token rate for each model, so a heavy month costs more than the sticker price.

On privacy, Ollama states that its cloud keeps zero data retention, does not log prompts, and does not train on your data. That is a policy you trust, not a property you can check. A local model is the only setup where you can verify that prompts never leave the box. For a longer cost comparison, read Ollama Cloud against running your own server, and for a middle path, running Ollama Cloud models from a small VPS keeps the API on your server while the heavy model runs remotely.

The costs of local Ollama that no pricing page shows

Local Ollama is free to run, but the hardware is not. These are the costs that decide whether a local setup is cheaper than a cloud plan.

How much RAM does the model need?

The whole model must sit in memory while it runs. A model with 8 billion parameters at the common 4-bit quantization (weights stored in about 4 to 5 bits each instead of 16) is close to 5 GB on disk, and it needs that much free memory plus room for the context. A 70-billion-parameter model at the same quantization needs about 40 GB. This one number decides your hardware budget, so work it out before you buy anything: how big an LLM fits in your RAM gives the arithmetic, and Q4 versus Q8 versus FP16 quantization shows what each step down in size costs in quality.

Check what is loaded and where it runs:

ollama ps

The PROCESSOR column should read 100% GPU on a GPU machine or 100% CPU on a CPU-only server. A split such as 48%/52% CPU/GPU means the model did not fit in video memory, so part of it runs on the slower CPU and every token waits for that part.

How much disk does Ollama use?

Every model you pull stays on disk until you remove it with ollama rm. On a Linux install that runs as a service, models live under /usr/share/ollama/.ollama/models.

ollama list
sudo du -sh /usr/share/ollama/.ollama/models

A few experiments with large models can fill a small VPS disk quickly. When the disk is full, a pull fails partway and the model will not load. How ollama pull and ollama run store models explains where the space goes and how to move the model directory.

A VPS or a box at home?

A machine at home costs money once, plus electricity. It can hold a consumer GPU, which is the biggest speed gain available. It usually sits behind your router's NAT (network address translation), with no fixed public IP address and a slow upload link, so reaching it from outside takes extra work.

A VPS (virtual private server) costs a monthly fee, with a public IP address and a data-centre network from day one. Most VPS plans have no GPU, so models run on the CPU, and RAM becomes the line item that drives the price. A plan with 16 GB of RAM runs 7 to 8 billion parameter models well enough for a personal API. Running Ollama on a VPS walks through the full setup.

Model weights carry their own licences

Ollama's MIT licence covers Ollama's code. It does not cover the models. Each model ships under its own terms. Llama models use the Llama Community License, which includes an acceptable-use policy and a separate clause for products above 700 million monthly active users. Gemma models use the Gemma Terms of Use. Many Qwen and Mistral models use Apache 2.0, which is permissive. Read the licence of the exact model you plan to ship with:

ollama show llama3.2 --license

This prints the licence text bundled with that model. If you build a product on a model, that text is the one that binds you, not Ollama's.

The real drawbacks of running Ollama yourself

The API has no authentication

Ollama binds to 127.0.0.1:11434 by default, so only the machine itself can reach it. Set OLLAMA_HOST=0.0.0.0 to reach it from your laptop, and anyone on the internet who finds the port can use it too, because the server has no API key or password. A stranger can run prompts on your CPU, pull large models to fill your disk, or delete the models you have. Check what the port is bound to:

sudo ss -ltnp | grep 11434

127.0.0.1:11434 means local only. 0.0.0.0:11434 or *:11434 means every network interface, which on a VPS means the public internet. Put a reverse proxy with authentication or a private network in front of it, as described in securing your Ollama API endpoint.

CPU-only inference is slow

To generate each token, the model reads its full set of weights from memory. So generation speed is limited by memory bandwidth divided by model size, and a CPU server has far less memory bandwidth than a GPU. Reading the prompt is a separate step that is limited by raw compute, so a long prompt on a CPU means a long wait before the first word appears. Measure your own machine instead of trusting a published figure:

ollama run llama3.2 --verbose "Explain DNS in two sentences."

After the answer, --verbose prints timing lines. prompt eval rate is how fast the prompt was read, and eval rate is how fast new tokens were generated, both in tokens per second.

The default context window is small

The context window is how many tokens (word pieces) the model can see at once, counting your prompt and its answer. As of October 2026, Ollama's documentation sets the default by video memory: under 24 GiB gets 4k tokens, 24 to 48 GiB gets 32k, and 48 GiB or more gets 256k. A CPU-only VPS has no video memory, so it falls in the 4k bucket. A prompt longer than the window is cut to fit, so the model silently ignores the start of a long document. The server log shows a truncating input prompt warning when this happens:

journalctl -u ollama -n 100 | grep -i truncat

Raise the default for the service with a systemd override:

sudo systemctl edit ollama
[Service]
Environment="OLLAMA_CONTEXT_LENGTH=16384"
sudo systemctl restart ollama

A larger context needs more memory, because the model keeps a cache entry for every token in the window. Raise it in steps and watch ollama ps and free -h.

Concurrency is limited

As of October 2026, OLLAMA_NUM_PARALLEL defaults to 1 in the documentation. Each loaded model answers one request at a time, and other requests wait in a queue of up to 512. Past that, the server answers with HTTP 503, which means it is overloaded. Raising OLLAMA_NUM_PARALLEL lets a model serve several requests at once, but the context memory is multiplied by the number of parallel slots, so four slots at 16k tokens need the memory of a 64k context. That is fine for one person and a coding agent. For many users at once, a serving engine built for batching fits better, and Ollama versus vLLM compares the two.

When local is cheaper and when the cloud is

Local Ollama is cheaper when your model fits in hardware you already own or a VPS you already pay for. Then each extra prompt costs nothing. It is also the only choice when prompts must stay on your own machine.

Ollama Cloud is cheaper when you want models far larger than your RAM allows, and your use is light enough that the included credits cover most months. Buying a GPU machine to run a 120-billion-parameter model a few times a week is hard to justify against a $20 plan. The two also mix well, because the same API serves both local and cloud models. Run small models locally for routine work, and send the hard prompts to a cloud model.

FAQ

Is Ollama free for commercial use?

Ollama itself is MIT licensed, so you can use it in commercial work at no cost. The model you run is a separate question, because each model has its own licence. Run ollama show <model> --license and read the terms. Apache 2.0 models are permissive, while the Llama and Gemma licences add their own conditions.

Do I need an Ollama account to run models locally?

No. Installing Ollama, pulling open models and running them on your own hardware needs no account and no payment. You only need to sign in with ollama signin to use Ollama Cloud models, including the starter models on the Free plan.

What does the Ollama Cloud Free plan include?

As of October 2026, the Free plan costs $0, allows 1 concurrent request, and includes a small monthly starter allowance on a limited set of starter models. You can add paid credits to unlock more models. Ollama moved to this per-token credit model on 2026-08-31, so check the pricing page for current figures.

Why is Ollama so slow on my VPS?

Most VPS plans have no GPU, so the model runs on the CPU. Each generated token reads the full model from memory, so speed is limited by memory bandwidth and model size. Run ollama run <model> --verbose to see the eval rate, and try a smaller or more heavily quantized model if it is too slow.

Is the Ollama API safe to expose to the internet?

Not on its own. The Ollama server has no authentication, so anyone who reaches port 11434 can run prompts, pull models and delete models. Keep the default bind of 127.0.0.1:11434, and put an authenticating reverse proxy or a private network in front of it for remote access.