SSD Nodes Learn Hosting plans →
Guides Matt ConnorBy Matt Connor

Self-host Hindsight agent memory on a VPS

Run Hindsight 0.10.2 on your own VPS with Docker: retain, recall and reflect behind loopback, an external Postgres, and the LLM choice that sets cost.

What Hindsight is

Hindsight is a memory server for AI agents. It runs as a long-lived service, keeps long-term memory in named memory banks, and exposes three operations: retain writes something down, recall searches it, reflect reasons over what the bank holds. Every bank also gets its own MCP (model context protocol) endpoint, so a coding agent can use the bank as a tool with no custom code. Put it on a VPS you control and Claude Code on your laptop, Codex on your desktop and your own agents on the same server can all read and write one memory.

This guide follows Hindsight v0.10.2, released 29 September 2026, and the documentation at hindsight.vectorize.io as it read on 30 September 2026. The project is built by Vectorize. Vectorize says Hindsight reaches state-of-the-art scores on long-term memory benchmarks. That is the project measuring itself, so read it as a reason to test the thing, not as a result you can quote.

Two details separate Hindsight from a vector database with an API in front of it. First, retain does not store your text as written: it extracts facts, resolves entities and indexes them. Second, retain and reflect both call an LLM. That means the server has a running cost per write, which is the part people miss when they compare memory backends on disk usage alone. The difference between the memory types an agent keeps and what each one costs to maintain is worth reading before you decide how much your agents should write down.

The documented quickstart, and why it is wrong on a VPS

The installation docs give this command:

export OPENAI_API_KEY=sk-xxx
docker run -it --pull always --name hindsight --restart unless-stopped --shm-size=1g \
  -p 8888:8888 -p 9999:9999 \
  -e HINDSIGHT_API_LLM_API_KEY=$OPENAI_API_KEY \
  -v hindsight-data:/home/hindsight/.pg0 \
  ghcr.io/vectorize-io/hindsight:latest

Port 8888 is the REST API and the MCP endpoint. Port 9999 is the Control Plane, the web UI where you browse banks, run recall queries and look at the knowledge graph. The volume holds an embedded PostgreSQL database at /home/hindsight/.pg0, which is why the command sets --shm-size=1g: Postgres wants shared memory, and Docker's 64 MB default is too small for it.

On a laptop that command is fine. On a VPS it publishes an unauthenticated API and an unauthenticated web UI on every interface, including your public IP.

A container port published with -p 8888:8888 is inserted into Docker's own iptables chain, which the kernel consults before the chain ufw writes. The port answers from the internet while sudo ufw status still reports the port as denied.

This is the single most important thing to get right here, and it is not specific to Hindsight: see why a Docker published port ignores your ufw rules for the mechanism and the fix. The short version is that you bind the published port to the loopback address yourself.

Bind both ports to 127.0.0.1

docker run -d --pull always --name hindsight --restart unless-stopped --shm-size=1g \
  -p 127.0.0.1:8888:8888 -p 127.0.0.1:9999:9999 \
  -e HINDSIGHT_API_LLM_API_KEY=$OPENAI_API_KEY \
  -v hindsight-data:/home/hindsight/.pg0 \
  ghcr.io/vectorize-io/hindsight:0.10.2

Two changes from the documented line. The ports carry an explicit 127.0.0.1: prefix, so Docker's rule only accepts traffic from the box itself. The image is pinned to 0.10.2 instead of latest, so a --pull always restart cannot swap the version under you.

Check it came up, from a shell on the VPS:

docker logs -f hindsight
curl -s http://127.0.0.1:8888/v1/default/banks

GET /v1/default/banks lists the banks in the default tenant. A JSON response means the API is serving. Connection refused means the container is still starting or has exited, and docker logs will say which. First start takes a while on a small VPS, because the full image downloads its local embedding model on first run.

Size the box for that. The docs put the full image at about 1.5 GB of RAM minimum and 2 GB recommended, because it loads local embedding and reranking models, and they put the slim image (ghcr.io/vectorize-io/hindsight:0.10.2-slim) at 512 MB minimum and 1 GB recommended, because slim sends embedding and reranking to external services instead. PostgreSQL wants another 512 MB on top. A 1 GB VPS running the full image will be killed by the kernel out-of-memory handler partway through the first retain.

How to reach the API and the UI once they are on loopback

An SSH tunnel is the version that needs nothing installed on the server:

ssh -N -L 8888:127.0.0.1:8888 -L 9999:127.0.0.1:9999 you@your-vps

While that runs, http://localhost:9999 in your browser is the Control Plane on the VPS, and http://localhost:8888 is the API. Anything on your laptop that can be pointed at localhost:8888 now works unchanged, which includes every MCP client below.

A tunnel per machine gets tiring once you have three machines and a phone. The other answer is a private network that all of them join, so the VPS has a stable address only your devices can route to. Running Headscale as your own Tailscale control server gives you that without a third party holding the coordination, and then you bind Hindsight to the tailnet address instead of 127.0.0.1 and skip the tunnel.

Does the Hindsight API ask for a password?

No, not unless you configure it. As of 30 September 2026 the MCP server documentation states that the MCP endpoint is open, with no authentication required by default. The docs describe API key authentication as something you turn on, using a tenant extension:

export HINDSIGHT_API_TENANT_EXTENSION=hindsight_api.extensions.builtin.tenant:ApiKeyTenantExtension
export HINDSIGHT_API_TENANT_API_KEY=your-secret-key

Clients then send the key in the Authorization header. Set both as -e flags on the container if you want them.

Set the key and keep the loopback binding. The key protects you from someone who reaches the port. The loopback binding is what stops them reaching it, and it also covers the Control Plane on 9999, which is a separate surface from the API.

The Compose variant with an external PostgreSQL

The single container carries its own embedded Postgres, which is simple and awkward to back up, because the database files sit inside a Docker volume that is being written to. If you already run Postgres, or you want the database managed and dumped on its own schedule, the repository ships a Compose file for exactly that.

curl -fsSLO https://raw.githubusercontent.com/vectorize-io/hindsight/v0.10.2/docker/docker-compose/external-pg/docker-compose.yaml

That file defines two services. db runs pgvector/pgvector:pg${HINDSIGHT_DB_VERSION:-18} with a named volume, and it does not publish a port, so the database is reachable only on the Compose network. hindsight runs ghcr.io/vectorize-io/hindsight:${HINDSIGHT_VERSION:-latest} and publishes 8888 and 9999. The file requires you to set HINDSIGHT_DB_PASSWORD and HINDSIGHT_API_LLM_API_KEY before it will start.

cat > .env <<'EOF'
HINDSIGHT_VERSION=0.10.2
HINDSIGHT_DB_PASSWORD=a-long-random-string
HINDSIGHT_API_LLM_API_KEY=sk-xxx
EOF
chmod 600 .env
docker compose up -d
docker compose logs -f hindsight

Edit the ports: entries in the file to "127.0.0.1:8888:8888" and "127.0.0.1:9999:9999" before that first up. The published-port rule above applies to Compose exactly as it applies to docker run, and a Compose file that shipped with 8888:8888 will expose the API the moment the stack starts. If the Compose syntax here is new, the Compose basics for running services on a VPS covers the file layout and the .env handling.

Pointing Hindsight at a Postgres you already run instead means setting HINDSIGHT_API_DATABASE_URL to a connection string. The database needs pgvector, which is why the bundled file uses the pgvector/pgvector image rather than stock postgres.

Choosing the LLM, which is what sets your running cost

Hindsight calls a model every time it retains and every time it reflects. The provider is one environment variable, HINDSIGHT_API_LLM_PROVIDER, and the docs list more than twenty values. Three shapes matter for a self-hosted box.

A hosted API key. HINDSIGHT_API_LLM_PROVIDER=openai or anthropic, plus HINDSIGHT_API_LLM_API_KEY, plus HINDSIGHT_API_LLM_MODEL if you do not want the documented default for that provider. Simplest to run, and the meter runs on writes rather than on reads, which surprises people who assumed memory is cheap because storage is cheap.

A local model. HINDSIGHT_API_LLM_PROVIDER=ollama with HINDSIGHT_API_LLM_BASE_URL pointed at your Ollama server, no API key needed. The documented default model for that provider is gemma3:12b. One trap: inside the container, localhost is the container, so a base URL of http://localhost:11434/v1 reaches nothing. Point it at the host, http://172.17.0.1:11434/v1 on a default Docker bridge, or put Ollama in the same Compose network and use the service name. Sizing is the real question, and running Ollama on a VPS to self-host the model covers what a 12B model actually needs in RAM. Adding that to a box that is already running Hindsight plus Postgres usually means a bigger VPS, not a free lunch.

A subscription you already pay for. The docs list claude-code, which uses Claude Pro or Max credentials managed by Claude Code and needs no API key, and openai-codex, which reads ~/.codex/auth.json as created by codex auth login. Both read a credential from the filesystem of the machine running the Hindsight API. On a server that means getting that credential into the container and keeping it there, which is a stored secret on a box you also expose to agents. The rules for keeping credentials out of reach of your agents apply here more than usual, because this credential is on the same host as the memory the agents read.

Embeddings are configured separately. The default is a local embedding model in the full image, which is why that image is large and why the slim image needs an external embedding service.

Connecting a coding agent

Hindsight ships an installer that wires its plugin into coding agents. Run it on the machine where the agent runs, not on the VPS:

npx @vectorize-io/hindsight-coding-agents install claude-code \
  --server self-hosted \
  --api-url http://localhost:8888

With the SSH tunnel open, http://localhost:8888 is your VPS. The same installer takes codex and cursor-cli in place of claude-code. It also reads HINDSIGHT_API_URL, HINDSIGHT_API_TOKEN and HINDSIGHT_SERVER_MODE from the environment, and a config file at ~/.hindsight/coding-agent.json. The plugin hooks the agent's session start, prompt and session end, so it recalls before work and retains after.

If you would rather add the memory as plain tools and keep control of when they are called, use the MCP endpoint directly:

claude mcp add --transport http hindsight http://localhost:8888/mcp \
  --header "Authorization: Bearer your-secret-key" \
  --header "X-Bank-Id: my-bank"

Each bank is also addressable as its own endpoint at http://localhost:8888/mcp/{bank_id}/. The server exposes retain, recall and reflect as tools, along with tools for listing memories and documents and managing banks. Banks do not need creating first: the docs say a bank is created with default settings on the first write to it, and a read of a bank that does not exist returns 404.

Write the first memory by hand so you can see the shape of it:

curl -X POST http://127.0.0.1:8888/v1/default/banks/my-bank/memory/retain \
  -H "Content-Type: application/json" \
  -d '{"items": [{"content": "The staging database runs on port 5433, not 5432.", "context": "infrastructure"}]}'

Then open the Control Plane on 9999 through the tunnel and look at the bank. You should see the fact extracted and entities resolved, rather than your sentence stored as a blob. If the bank is empty, check docker logs hindsight for an LLM error: a wrong API key fails here, at retain time, and not at startup.

What a shared memory server adds to your risk

Banks are isolated from each other, and that is the main control you have. Memories in one bank are not visible to another, so splitting by project or by person is not organisation, it is the boundary.

Inside a bank there is no such boundary. Every agent with access reads what any other agent wrote, which means one bad write propagates to all of them and survives the session that created it. That is the whole problem of a poisoned memory that keeps being recalled long after the attack, and a shared server widens the blast radius from one agent to your whole fleet. Treat any content an agent ingests from the web or from a pull request as untrusted input to retain, not as a fact.

Hindsight has a feature named Memory Defense, and it is worth knowing what it does and does not cover. The docs describe it as a scan of memories before storage using a 45-pattern regex set, which redacts secrets and personal data such as API keys, connection strings and card numbers. It is disabled by default and is configured per bank through PATCH /v1/{tenant}/banks/{bank_id}/config:

{
  "memory_defense": {
    "enabled": true,
    "rules": [
      { "on": "sensitive_data", "action": "redact" }
    ]
  }
}

That protects you from an agent writing a leaked credential into long-term memory. It is not a defence against poisoned content, because a planted instruction contains no secret to match. Those are separate problems with separate answers.

The other slow problem is that a memory bank fills up with things that used to be true. A fact retained in March about a service you have since replaced still scores well on a recall query, so deciding what to prune and when belongs on the schedule from the start, not after the first wrong answer.

Deleting what it holds about a person

Hindsight's documented design is append-only, and it prefers invalidating a fact to deleting it: an invalidated fact drops out of recall, consolidation and the knowledge graph, but is kept for audit. That is good for debugging and it is not the same thing as erasure, which matters when someone asks you to remove their data rather than to stop using it.

What the docs do give you is coarse. DELETE /v1/default/banks/{bank}/memories/{id}/observations clears a memory's derived observations. The MCP surface includes delete_document, clear_memories and delete_bank, and the banks API documents DELETE /v1/default/banks/{bank_id}, which removes the bank and its aliases. There is no documented endpoint for deleting every memory about a named person across a bank.

So build the handle in before you need it. Give each person or customer their own bank, or retain their material under a document_id you can name later, because the deletion operations that exist work on a bank or on a document. Retro-fitting that after a year of mixed writes is the expensive version. What it actually takes to erase one person's data from an agent's memory goes through the derived copies, the embeddings and the extracted facts that a single delete usually misses.

Backups

Where the data lives depends on which deployment you chose, and it is the one path you should never guess. With the single container, it is the Docker volume mounted at /home/hindsight/.pg0. Stop the container before you archive it, because it is a live Postgres data directory and a copy taken mid-write is a copy of a torn database:

docker stop hindsight
docker run --rm -v hindsight-data:/data -v "$PWD":/backup alpine \
  tar czf /backup/hindsight-data.tgz -C /data .
docker start hindsight

With the external-pg Compose stack there is nothing to stop. Use pg_dump against the db service on its schedule, keep the dump off the box, and test a restore into a scratch database once, so you find out now rather than later whether the pgvector extension is present in the target.

Either way the LLM configuration is not in the backup. Keep a note of HINDSIGHT_API_LLM_PROVIDER and the model you set, because a restored database with a different embedding model behind it will not recall the way it used to.

If all of this is more server than you want for one laptop's worth of memory, the single-machine option is worth a look first: a local agent memory that runs beside the agent has no Postgres, no ports to bind and no LLM bill, at the cost of the sharing that is the entire reason to put Hindsight on a VPS.

FAQ

Does the Hindsight API require authentication by default?

No. As of 30 September 2026 the documentation states that the MCP endpoint is open with no authentication required, and API key auth is opt-in through HINDSIGHT_API_TENANT_EXTENSION set to the built-in ApiKeyTenantExtension plus HINDSIGHT_API_TENANT_API_KEY. Because the quickstart also publishes ports 8888 and 9999 on every interface, and a Docker published port is not filtered by ufw, the API and the Control Plane are reachable from the internet with no password until you change the port binding. Bind both to 127.0.0.1 and reach them over an SSH tunnel or a private network.

Can I run Hindsight without paying an LLM provider?

Yes, with a local model. Set HINDSIGHT_API_LLM_PROVIDER=ollama and HINDSIGHT_API_LLM_BASE_URL to your Ollama server, which needs no API key. The docs also list subscription providers, claude-code for Claude Pro or Max and openai-codex which reads ~/.codex/auth.json, and both read credentials from the machine running the API. The cost does not disappear in the local case, it moves to RAM on your VPS, since Hindsight plus PostgreSQL plus a model on one box is a much larger server than Hindsight alone.

Where does Hindsight store its data on a self-hosted install?

In the single-container deployment it is the embedded PostgreSQL database in the volume mounted at /home/hindsight/.pg0, which the documented docker run command names hindsight-data. In the external-pg Compose deployment it is the PostgreSQL service, which uses a pgvector image and keeps its files in its own named volume. Back up the volume with the container stopped, or run pg_dump against the database service.

Can I delete everything Hindsight knows about one person?

Not with one call. The design is append-only, and the documented preference is to invalidate a fact, which removes it from recall, consolidation and the knowledge graph while keeping it for audit. Deletion operates at coarser levels: clearing a memory's derived observations, deleting a document, clearing memories, or DELETE /v1/default/banks/{bank_id} for a whole bank. Plan for this by giving each person their own bank or a stable document_id, so a later erasure request has something to target.