SSD Nodes Learn Hosting plans →
How to do am Matt ConnorBy Matt Connor · Updated 2026-08-29

How to Use Ollama with Your Coding Agent

Connect your coding agent to Ollama with one base URL and any dummy key. See the 11434 endpoints, context-length gotcha, keep-alive, and best local-model jobs.

Wetin you dey connect

You fit use Ollama with your coding agent, and the connection small pass wetin people dey expect. You go change one base URL and choose one model name. The API key field still need value, but the local server dey ignore am, so any string go work.

Ollama dey listen on port 11434 and e dey serve two request shapes at the same time. /v1/chat/completions na the OpenAI-compatible shape, and Ollama documentation describe the key for there as required but ignored. /v1/messages na the Anthropic-compatible shape, wey Claude Code dey speak. Your agent already dey speak one of the two, so nothing else about am go change.

That part fit take five minutes. Whether the result go usable depend on two settings wey almost nobody dey change, the context length and the keep-alive, plus whether you give the model the kind work wey e good at. Each one get im own section, and the honest limits dey for the end.

Which coding agents dey accept local base URL

The test na one question: does the tool expose a base URL setting? If e do, e fit talk to your server.

Ollama publishes integration pages for Claude Code, OpenCode, Codex, Cline, Roo Code, Zed, JetBrains IDEs and VS Code. Aider documents its own Ollama support separately. This one cover most of wetin people mean by coding agent for August 2026. Dem no all dey speak the same format, and na this difference dey make setup fail.

  • Most agents want OpenAI-compatible endpoint. Give dem the base URL http://localhost:11434/v1 and any non-empty API key string.
  • Claude Code no accept OpenAI base URL at all. E dey speak Anthropic Messages API, so you need set ANTHROPIC_BASE_URL to http://localhost:11434, where Ollama dey serve /v1/messages.
  • Codex dey use OpenAI Responses API. Ollama dey serve /v1/responses too, and dem add am for version 0.13.3.
  • Agent wey no get base URL setting no fit redirect, because endpoint dey built into the client. Put translation layer for front instead, like self-hosted LiteLLM gateway, then expose your model again for the format wey the client require.

Ollama fit write these configs for you. ollama launch opencode starts OpenCode with inline config for the model wey you choose, ollama launch claude do the same thing for Claude Code, and ollama launch droid --config writes the config without launching the tool.

Install Ollama and pull model wey fit call tools

curl -fsSL https://ollama.com/install.sh | sh
systemctl status ollama --no-pager
ollama pull qwen3-coder:30b
ollama ls

The installer dey add systemd unit and start am, so systemctl status ollama suppose print active (running). If e no do so, journalctl -e -u ollama go print the reason.

The model suppose support tool calling, because na tool calling agent dey use work. E go read file, write patch, run test, then read the failure and try again. Model wey no fit emit tool call go describe the edit for prose instead of making am, and agent go loop or stop. Look for tools label for the model page on ollama.com before you pull am. qwen3-coder:30b get this label, and as of August 2026 that tag na 19 GB download with 256K context window. If your box na CPU-only or RAM no plenty, the memory arithmetic for the Qwen 27B tag on a VPS show wetin fit enter 8 to 64 GB before you spend the download. Once you pull am, those gigabytes go enter root disk of the server. Na the part of VPS wey get the least free space. So where Ollama dey keep its model files and how to move dem elsewhere worth reading before disk full.

Now confirm which names the server actually dey serve:

curl http://localhost:11434/v1/models

The strings for that response na wetin your agent config must contain, character for character. Check am first to settle most model-not-found errors. If Ollama never install, the longer walkthrough dey for how to self-host an LLM with Ollama on a VPS.

Point OpenCode go Ollama

Edit ~/.config/opencode/opencode.json:

{
  "$schema": "https://opencode.ai/config.json",
  "provider": {
    "ollama": {
      "npm": "@ai-sdk/openai-compatible",
      "name": "Ollama",
      "options": {
        "baseURL": "http://localhost:11434/v1"
      },
      "models": {
        "qwen3-coder:30b": {
          "name": "qwen3-coder 30b"
        }
      }
    }
  }
}

The key under models na the model name wey OpenCode send go Ollama, so e gats match ollama ls exactly. The name field na only the label wey dey model picker. Start opencode, switch go Ollama provider, and monitor journalctl -e -u ollama to confirm say request reach your server, no be somewhere else. Setup for the agent itself dey covered for running OpenCode for VPS.

Point Claude Code go Ollama

export ANTHROPIC_AUTH_TOKEN=ollama
export ANTHROPIC_API_KEY=""
export ANTHROPIC_BASE_URL=http://localhost:11434
claude --model qwen3-coder:30b

ANTHROPIC_API_KEY dey set to empty string on purpose. If real key dey inside environment, e go send your requests go hosted API instead. This one fit cause bill, and you no go get local inference. ollama launch claude dey set all these things for you.

Know wetin compatibility layer no include. E no implement tool_choice or prompt caching. E never get token counting endpoint too, so the token numbers wey you see na only approximations from the model own tokenizer. Claude Code still ships with large system prompt and large tool set, so e need more context than chat client. The bigger question of wetin carry over and wetin no carry over dey covered for whether you fit self-host Claude.

Point Aider go Ollama

export OLLAMA_API_BASE=http://127.0.0.1:11434
aider --model ollama_chat/qwen3-coder:30b

Aider documentation recommend ollama_chat/ prefix pass ollama/. E still let you pin context window for each model inside .aider.model.settings.yml, wey useful when one model need different window from the server default:

- name: ollama_chat/qwen3-coder:30b
  extra_params:
    num_ctx: 65536

Why setup wey dey work still dey produce nonsense

Na this section matter pass. Ollama dey choose default context length from the VRAM (video memory for GPU) wey e fit see, and dem don publish those defaults:

ChartOllama default context length by available VRAM, documented August 2026
The data behind this chart
[
  {
    "label": "Under 24 GiB VRAM",
    "default_context_tokens": "4,096"
  },
  {
    "label": "24 to 48 GiB VRAM",
    "default_context_tokens": "32,768"
  },
  {
    "label": "48 GiB VRAM or more",
    "default_context_tokens": "262,144"
  }
]

Most VPS plans, and every CPU-only server, go land for first row: 4,096 tokens. Na only big GPU go get the 262,144 tokens for last row.

Agent dey pass 4096 tokens before e do any work. System prompt, tool definitions, repository listing, and first file wey e open don already pass that size. Wetin happen next na the main problem: nothing go show error. Aider documentation talk say Ollama dey silently discard context wey pass the window. The oldest tokens go comot, so model go answer with confidence about file wey e no fit see again, or forget instruction wey you give am two steps earlier. Na this mechanism dey cause most reports say local model too dull to write code. Choosing the number itself na another decision, and wetin num_ctx dey cost for KV cache memory at each size worth reading before you settle for one.

Ollama documentation talk say tasks like agents and coding tools suppose use at least 64000 tokens. Set am for server:

sudo systemctl edit ollama.service

Add these lines for override file:

[Service]
Environment="OLLAMA_CONTEXT_LENGTH=64000"

Then reload and restart:

sudo systemctl daemon-reload
sudo systemctl restart ollama
ollama ps

ollama ps na the check. E go print CONTEXT column, and that number na wetin model actually receive. Your ID and SIZE go different:

NAME               ID              SIZE     PROCESSOR    CONTEXT    UNTIL
qwen3-coder:30b    a1b2c3d4e5f6    24 GB    100% GPU     64000      4 minutes from now

Set am for server instead of agent for two reasons. OpenAI chat completions schema no get field for context length, so OpenAI-compatible client no fit ask for one. And the setting dey apply per server, so every agent wey you point to the box go inherit am. Output side get its own ceiling, and unlike context length, e dey travel through compatibility endpoint. So num_predict and the max_tokens field wey map to am na wetin you suppose use when reply stop halfway through patch. If one model need different window, bake am into copy with a Modelfile:

FROM qwen3-coder:30b
PARAMETER num_ctx 65536
ollama create qwen3-coder-64k -f Modelfile

Context no free. Longer window dey cost more memory, so monitor PROCESSOR column. 100% GPU na wetin you want. Once part of model spill enter CPU, token rate go fall reach level wey make agent loop unusable. Measuring tokens per second for local LLM na how you go find the real ceiling for your box. How much RAM and CPU coding agent VPS need cover how to size machine before you buy am.

Keep model loaded between requests

By default, Ollama dey unload model 5 minutes after the last request. This one good for chat box but e no good for agent work. You fit pause to read diff, timer go finish, then the next request go reload tens of gigabytes of weights from disk before the first token show. E go look like say system hang.

OLLAMA_KEEP_ALIVE dey take duration string like 10m or 24h, plain number of seconds, -1 to keep model loaded forever, or 0 to unload am immediately. Set am beside the context length:

[Service]
Environment="OLLAMA_CONTEXT_LENGTH=64000"
Environment="OLLAMA_KEEP_ALIVE=-1"

The keep_alive request field dey only for Ollama native /api/generate and /api/chat endpoints. E no dey for compatibility endpoints, so agent no fit set am for each request. The environment variable na the only control wey you get. When you need the memory back, ollama stop qwen3-coder:30b go unload the model without stopping the server. If you want the setting to remain after reboot, or you want compare keeping the weights for memory all day with getting that memory back, keeping an Ollama model loaded in memory fit work through both.

Ollama run for separate server

Ollama dey bind to localhost. To reach am from another machine, set OLLAMA_HOST=0.0.0.0:11434 for the same systemd override, then restart the service.

Do this only for private network. Ollama documentation talk say local API no need authentication, so if port 11434 open to internet, anybody fit use your hardware and read anything wey your agent send. Two safe options dey. Keep the bind for localhost and forward the port through SSH from your laptop:

ssh -N -L 11434:localhost:11434 you@your-vps

Your agent still dey point to http://localhost:11434/v1 and e no go know any difference. The other option na VPN, with Ollama bound to the VPN address instead of 0.0.0.0. If many people or many agents go share one box, Ollama scheduler no build for that kind load, and the comparison between Ollama and vLLM show where the throughput difference start to cause problem.

Where local coding model dey win, and where e no dey

Agent wey model you host dey drive no fit replace frontier API for every task. E dey win clearly for four kinds of work.

  • Bulk mechanical edits, where each change small and you fit check am. Rename things across repository, add type hints, write docstrings, translate comments. Model fit run for hours and bill no go change.
  • Work wey must not leave your hardware. Client code under confidentiality agreement, or internal repository wey you no get permission to send to third party.
  • Offline and air-gapped machines, where hosted API no dey available to call at all.
  • Predictable cost. Once you don pay for server, agent wey dey burn tokens inside loop no cost anything extra. This na opposite of metered API. Where GPU VPS dey break even against API tokens get the arithmetic.

E dey lose for long multi-step tasks. "Find why this test dey fail, fix the cause, update the callers" need many correct tool calls one after another, while the complete history still dey for context. Model for 8B to 14B range on modest server fit produce malformed tool call, or lose the plan after few turns. Then you go spend more time directing am pass the time wey the task suppose take. This no be prompt problem wey you fit solve just by writing better prompt. Na capacity problem.

E still dey lose anytime mistake cost plenty and you no go read every line. Give local model narrow jobs wey you go verify the output, and keep hosted model for work wey you no go check step by step.

Failure modes, and the strings you will see

curl: (7) Failed to connect to localhost port 11434 after 0 ms: Connection refused. Server no dey run, or agent dey point to another host. Run systemctl status ollama, then journalctl -e -u ollama.

Agent reports say model no exist. The name for your config no match any name wey server dey serve. Compare am with curl http://localhost:11434/v1/models and copy the string from there. Tag na part of the name, so config wey name tag wey you never pull go fail even though similar model dey installed.

Agent answers with prose and e never edit file. Either model no get tool support, or request plus the tool definitions don already fill context window. Check tools label for model page, then check CONTEXT column for ollama ps.

Long silence before first token, then normal speed. Keep-alive don expire and system dey read weights from disk again. Set OLLAMA_KEEP_ALIVE.

Model contradicts file wey e just read. Na context truncation. ollama ps usually dey show CONTEXT value wey small pass wetin you think say you set, because environment variable go your shell instead of the systemd unit.

Everything dey work, but e slow, and PROCESSOR no be 100% GPU. Model plus its context no fit inside VRAM. Reduce context length, or move to smaller model or smaller quantisation. Before you re-pull, wetin q4_K_M, q8_0 and fp16 each cost for memory, and where quality really dey drop go tell you how much space stepping down go free and wetin you go give up.

FAQ

I fit point Claude Code go Ollama?

Yes, but e no work with OpenAI-compatible URL. Claude Code dey speak Anthropic Messages API, and Ollama dey serve that format for /v1/messages on the same port 11434. Export ANTHROPIC_BASE_URL=http://localhost:11434, ANTHROPIC_AUTH_TOKEN=ollama and empty ANTHROPIC_API_KEY, then start am with claude --model qwen3-coder:30b. ollama launch claude go write the same settings for you. This compatibility layer no implement tool_choice or prompt caching, and e no get token counting endpoint, so token counts wey e report na approximation.

Why local model dey answer about code wey e no fit see?

Na because request don pass the context window, and the oldest part don drop without any error. Ollama dey set default context from the VRAM wey e find, and below 24 GiB that default na 4,096 tokens. Agent system prompt and tool definitions fit exceed that by themselves. Set OLLAMA_CONTEXT_LENGTH=64000 inside the systemd unit, restart Ollama, then confirm say the CONTEXT column for ollama ps dey show the new value.

Which model make I run for coding agent on VPS?

Choose the biggest model wey get tools label and still fit inside memory with 64k context window. Prefer model wey dem tune for code. qwen3-coder:30b na the common choice for GPU server wey get enough VRAM. If that tag too big for your server, the RAM figures and CPU-only speeds for Nemotron 3.5 Lightning fit help you compare before you download am. Below about 14B parameters, model fit still answer code questions well but fail for multi-step edits. Agent work dey expose small formatting mistakes for tool calls. Test am with one real task from your own repository instead of sample prompt.

I need GPU to run coding agent on my own model?

For practical use, yes. CPU-only inference dey work and e okay for single questions, but agent dey send many requests for each task and each request dey read long history again. Slow token rate fit turn two-minute task into one hour. Check the PROCESSOR column for ollama ps: any value wey no be 100% GPU mean say part of the model dey run for CPU, and token rate go drop sharply.

#ollama#coding-agent#openai-compatible#local-llm#self-hosted-ai