SSD Nodes Learn Hosting plans →
Guides Matt ConnorBy Matt Connor

Can Ollama Search the Web? Three Ways to Do It

Ollama does not browse on its own. Give a local model search through Ollama's web search API, Open WebUI with SearXNG, or a short Python tool loop you control.

Can Ollama search the web?

Can Ollama search the web? Not by itself. The Ollama runtime loads a model and generates text. It never opens a browser and never sends a query anywhere. A model can only search when some program gives it a search tool, runs the search when the model asks for one, and feeds the results back into the conversation.

That program can be one of three things. It can be Ollama's own hosted search API, a chat front end such as Open WebUI, or a short script you write yourself. This guide sets up each one. For each route it also says who sees your search queries and what the route costs, because those two answers differ a lot between the three.

How does a local model search at all?

Search through Ollama uses tool calling. Your program sends the model a chat request plus a list of tools. Each tool has a name, a description and its arguments. A model trained for tool use can reply with a tool call instead of text, for example searxng_search(query="ollama release notes"). Ollama passes that call back to your program in message.tool_calls. Your program runs the search, adds the results as a message with the role tool, and sends the whole conversation back. The model then writes its answer from those results.

Two facts follow from this. First, the model never touches the network. Your code does. Second, the model must support tool calling. Check that on your own machine:

ollama show qwen3:4b

The Capabilities section of the output should list tools. If it lists only completion, the model was not trained for tool calls. A chat request that includes tools then fails with an error that ends in does not support tools. Check each model you plan to use with ollama show, or on that model's own page in the Ollama library. Do not rely on a list of "tool-capable models" in a blog post, because new models and new versions of old models change the answer.

If Ollama is not running on your server yet, start with running Ollama on your own VPS and come back here.

Search results are long, which matters for one setting. Every result the program sends back takes space in the context window, which is the amount of text the model can see at once. The Ollama documentation recommends a context length of about 32,000 tokens for search agents. Ollama's default context is often smaller than that, so in a long search session the original question can fall out of the window. The model then answers a question it can no longer see. Setting num_ctx for a larger context window explains how to raise it and how much memory that costs.

Does Ollama work offline?

Local inference works offline. Once a model is pulled, ollama run needs no network, and your prompt never leaves the machine. Any search needs network. Every route below sends your query to some server on the internet. The only real choice is which server, and who runs it. A machine with no network can answer only from what the model learned in training, so it knows nothing newer than its training data.

Route one: Ollama's web search API

Ollama runs a hosted search service with two endpoints. web_search returns results for a query. web_fetch returns the text of one web page. Both work with local models: the model runs on your server, and only the search and fetch requests go to ollama.com.

You need a free Ollama account and an API key. Create the key at ollama.com/settings/keys, then test it from the server:

export OLLAMA_API_KEY='paste-your-key-here'
curl https://ollama.com/api/web_search \
  --header "Authorization: Bearer $OLLAMA_API_KEY" \
  -d '{"query":"what is ollama?"}'

A working key returns a JSON object with a results list. Each result has a title, a url and a content field. The request also accepts max_results, which defaults to 5 and allows up to 10. A missing or wrong key gets an authorization error instead of results. The usual cause is that the variable was exported in a different shell. Run echo ${OLLAMA_API_KEY:+set} in the shell you are using. It prints set when the variable exists, and an empty line when it does not.

Now install the Python library into a virtual environment and pull a model that supports tools:

sudo apt install -y python3-venv
python3 -m venv ~/ollama-search
~/ollama-search/bin/pip install ollama requests
ollama pull qwen3:4b

The virtual environment is needed because Ubuntu 24.04 refuses pip install into the system Python with an externally-managed-environment error.

The Python library ships web_search and web_fetch as ready-made functions, so you pass them straight in as tools. This loop is the example from Ollama's documentation, shortened a little and with a larger context window:

from ollama import chat, web_fetch, web_search

available_tools = {"web_search": web_search, "web_fetch": web_fetch}
messages = [{"role": "user", "content": "What is new in the latest Ollama release?"}]

while True:
    response = chat(
        model="qwen3:4b",
        messages=messages,
        tools=[web_search, web_fetch],
        think=True,
        options={"num_ctx": 32768},
    )
    if response.message.content:
        print("Content:", response.message.content)
    messages.append(response.message)
    if not response.message.tool_calls:
        break
    for tool_call in response.message.tool_calls:
        print("Tool call:", tool_call.function.name, tool_call.function.arguments)
        function_to_call = available_tools.get(tool_call.function.name)
        if function_to_call:
            result = function_to_call(**tool_call.function.arguments)
            content = str(result)[:8000]
        else:
            content = f"Tool {tool_call.function.name} not found"
        messages.append({"role": "tool", "content": content, "tool_name": tool_call.function.name})

Save it as search.py and run it with ~/ollama-search/bin/python search.py. You should see one or more Tool call: lines and then a Content: line with the answer. If the model answers at once with no tool call, it decided it did not need to search. Ask about something from this week, which the model cannot know from training, and it should call the tool. The [:8000] cut keeps one long page from filling the whole context window.

Who sees the queries: Ollama does. Every search term and every fetched URL goes to ollama.com under your account. The prompt and the model's answer stay on your server. The search terms usually reveal what the conversation is about, though, so treat them as part of the conversation.

What it costs: the API works on a free account. Ollama sets the usage limits for each plan and can change them, so check the current numbers on ollama.com before you build anything that depends on them. For the wider question of where Ollama's free use ends, see whether Ollama is free, local versus cloud.

Route two: a front end that searches for the model

Open WebUI is a chat interface for Ollama, and it can run web searches itself. With web search turned on for a chat, Open WebUI uses the model to write a search query, runs that query against a search backend, and puts the results into the prompt. In this mode it does not depend on the model's own tool calling, so it also works with small models that lack tool support.

The backend this guide recommends is SearXNG, a self-hosted metasearch engine. A metasearch engine sends your query to several public search engines and merges their results. In Open WebUI, open the Admin Panel, then Settings, then Web Search. Turn on web search, choose searxng as the engine, and set the query URL. If both run in the same Docker network, the URL is http://searxng:8080/search?q=<query>. The /search?q=<query> part is required.

Open WebUI asks SearXNG for JSON, and SearXNG serves only HTML until you allow JSON in its settings.yml:

search:
  formats:
    - html
    - json

Restart SearXNG after the edit, then test it from the server:

curl -s -o /dev/null -w '%{http_code}\n' 'http://localhost:8080/search?q=test&format=json'

200 means JSON is on. 403 means SearXNG refused the JSON request, because json is missing from formats or SearXNG was not restarted after the change. In Open WebUI this shows up as a search that returns nothing. The complete setup, including the Docker Compose file, is in the SearXNG JSON API guide for Open WebUI.

Who sees the queries: the upstream engines that SearXNG forwards to, such as Google or DuckDuckGo. They see the query arrive from your server's IP address, with no browser cookies and no account attached. Nobody else sees it, as long as your SearXNG instance is not open to the public.

What it costs: nothing beyond the server. There is one real catch. Public engines rate-limit traffic from datacenter IP addresses, so an engine can start answering with a CAPTCHA instead of results. Fixing SearXNG engine CAPTCHA errors covers what to change when that happens.

Route three: your own tool loop against SearXNG

The third route uses the same SearXNG instance with no chat interface at all. You write a small Python function that calls the SearXNG JSON API and give it to the model as a tool. This gives you full control: you decide how many results the model sees, what each result contains, and whether every query is logged.

Run this on the server where both Ollama and SearXNG run. It uses the virtual environment from route one. Change SEARXNG_URL if your instance listens somewhere else.

import requests
from ollama import chat

SEARXNG_URL = "http://localhost:8080/search"

def searxng_search(query: str) -> str:
    """Search the web through a private SearXNG instance.

    Args:
        query: The search query

    Returns:
        The top results, each with title, URL and snippet
    """
    r = requests.get(SEARXNG_URL, params={"q": query, "format": "json"}, timeout=20)
    r.raise_for_status()
    lines = []
    for item in r.json().get("results", [])[:5]:
        lines.append(item.get("title", "") + "\n" + item.get("url", "") + "\n" + item.get("content", ""))
    return "\n\n".join(lines) or "No results."

messages = [{"role": "user", "content": "What changed in the latest SearXNG release?"}]

for _ in range(5):
    response = chat(
        model="qwen3:4b",
        messages=messages,
        tools=[searxng_search],
        options={"num_ctx": 32768},
    )
    messages.append(response.message)
    if not response.message.tool_calls:
        print(response.message.content)
        break
    for call in response.message.tool_calls:
        print("Searching:", call.function.arguments)
        if call.function.name == "searxng_search":
            output = searxng_search(**call.function.arguments)
        else:
            output = "Unknown tool: " + call.function.name
        messages.append({"role": "tool", "content": output, "tool_name": call.function.name})

The docstring is part of the program. The Ollama library reads the function's signature and its Args: section to build the tool description that the model sees. A vague description leads to vague tool calls, so describe the tool the way you would describe it to a colleague.

Save the file as searx_agent.py and run ~/ollama-search/bin/python searx_agent.py. You should see a Searching: line with the query the model chose, then the answer. Two failures are common, and both show up clearly:

  • requests.exceptions.HTTPError: 403 Client Error: Forbidden means SearXNG still has JSON turned off. Fix formats as in route two.
  • requests.exceptions.ConnectionError means nothing listens at SEARXNG_URL. Check the port with curl -I http://localhost:8080.

The loop stops after five rounds. That cap is a guard: a model that keeps asking for more searches would otherwise keep the script running with no end. The five-result cut and the short snippets keep each round small, so the context window lasts longer.

Who sees the queries: the same upstream engines as in route two, and no one else. What it costs: only the server. The same SearXNG endpoint can serve other agents too, which a SearXNG search skill for AI agents builds on.

Which route should you pick?

Pick route one if you want search working in five minutes and you are comfortable with ollama.com seeing your search terms. It also gives you web_fetch, which returns the full text of a page. With SearXNG you get snippets only, unless you write your own fetch tool.

Pick route two if people will chat with the model in a browser. Open WebUI handles the search step for them, and it works with models that cannot call tools.

Pick route three if search is part of a program, a scheduled job or an agent. It is the only route where you control every request, and it needs nothing beyond what you already run. If your agent is a coding assistant, using Ollama with your coding agent shows how to connect the model side.

Keep port 11434 and SearXNG private

All three routes talk to Ollama on port 11434. Ollama has no authentication on that port, so anyone who can reach it can run your models and use your server's memory. By default Ollama listens only on 127.0.0.1, which is safe. Do not change OLLAMA_HOST to 0.0.0.0 to make a remote front end work. Put a proxy with authentication in front instead, as described in securing your Ollama API endpoint. The Ollama API and port 11434 explains what that port exposes.

The same rule applies to SearXNG. An open instance with JSON enabled is a free search API for anyone who finds it, and the upstream engines will block your server's IP address because of their traffic. If you must expose it, follow hardening a public SearXNG instance first.

FAQ

Can Ollama search the web without an API key?

Yes. Ollama's hosted web_search API needs a free account and an API key, but the other two routes do not. Run your own SearXNG instance with JSON output enabled, then either connect it to Open WebUI or call it from a short Python tool loop. In both cases the queries go from your server to public search engines, and no account is involved.

Does Ollama work offline?

Local inference works offline. After ollama pull, a model runs with no network connection, and your prompts stay on the machine. Web search always needs network, whichever route you use. An offline model can only answer from its training data, so it cannot know about events after its training cutoff.

Why does my model never call the search tool?

First run ollama show <model> and check that Capabilities lists tools. A model without tool support fails with an error ending in does not support tools. If the model supports tools but still answers directly, it decided no search was needed. Ask about something recent, and make the tool description in the function's docstring clear about when to use it.

Does Ollama's web search send my data to ollama.com?

The hosted web_search and web_fetch APIs send every search query and every fetched URL to ollama.com under your account. Your full prompt and the model's answer stay on your server when the model runs locally. If you want no third party other than the search engines themselves to see your queries, use a self-hosted SearXNG instance instead.

#ollama#web-search#searxng#tool-calling#local-llm