SSD Nodes Learn Hosting plans →
Guides Matt ConnorBy Matt Connor

Hindsight vs mem0 for self-hosted agent memory

Hindsight 0.10.2 against mem0 2.2.1 for agent memory on a VPS: memory model, services you run, LLM cost, MCP, licence and what self-hosting drops.

Hindsight vs mem0: which one should you self-host?

Hindsight and mem0 are the two self-hosted agent memory servers most people shortlist, and they split on one question: how much of the system runs on your own box. Hindsight ships a single container that carries its own PostgreSQL, its own embedding model, its own reranker, and a built-in MCP (model context protocol) endpoint, so a call to an LLM (large language model) API is the only network hop it needs. mem0 ships a smaller server that reaches out to hosted services by default, and several of the features on its own benchmark page are in its managed platform only.

This compares Hindsight v0.10.2, released 29 September 2026, against the mem0ai Python package 2.2.1, the current release on PyPI as of 30 September 2026. Every claim below comes from each project's own README and documentation, read on 30 September 2026. Where a project states something only about its paid cloud, this says so.

What does each one store, and what can you ask it to do?

mem0 stores memories extracted from messages you hand it, and two operations carry almost all the traffic. add sends your messages to an LLM, which extracts facts and writes them. search turns a question into an embedding and returns matching memories.

from mem0 import Memory

m = Memory()
m.add(messages, user_id="alex")
m.search("What do you know about me?", filters={"user_id": "alex"})

The open source edition scopes memories by user_id, agent_id and run_id. The app_id scope, organisations and projects are platform features. mem0's April 2026 algorithm changed extraction to what the README calls "Single-pass ADD-only extraction", which removes the UPDATE and DELETE passes the older algorithm ran on every write.

On graphs, read the docs carefully before you plan around them. mem0's graph memory is now native: the docs say the platform "builds a native graph linking people, places, and concepts across your memories, with no external graph database to provision", and that "There is no Neo4j, Memgraph, or other graph store to deploy". The earlier Neo4j, Memgraph, Kuzu, Apache AGE and Neptune integrations are gone. The platform comparison page lists Graph Memory among the v3 ranking features, so a self-hoster on the open source edition does not get it.

Hindsight organises everything into memory banks, one per user or agent, each in strict isolation. It exposes three operations. Retain extracts facts, entities, temporal data and relationships from what you store. Recall retrieves without an LLM. Reflect reasons over retrieved memories using an LLM, guided by the bank's mission and directives. Inside a bank, memories are typed: world facts, experiences, observations (deduplicated beliefs that carry their evidence), mental models (standing answers to a question), and knowledge pages (living wiki-style documents).

curl -fsSL https://hindsight.vectorize.io/get-cli | bash
hindsight memory retain my-bank "Alice works at Google as a software engineer"
hindsight memory recall my-bank "What does Alice do?"
hindsight memory reflect my-bank "Tell me about Alice"

Recall runs four strategies at once: semantic vector matching, BM25 keyword search, graph traversal over entity and temporal links, and temporal filtering. The four ranked lists are merged with reciprocal rank fusion (RRF), then a cross-encoder rescores the top candidates, capped at 300 by default. That cap exists because a cross-encoder costs one model call per candidate, so it has to run last on a shortlist. If you care how a memory store's shape constrains the agent on top of it, a portable open format for agent memory is a useful contrast, and Hindsight's strict bank isolation is the same idea as keeping each project's memory in its own compartment.

What do you actually run on the box?

Hindsight's quickstart is one container, and the image carries an embedded PostgreSQL so nothing else is required to start.

export OPENAI_API_KEY=sk-xxx
docker run -it --pull always --name hindsight --restart unless-stopped \
  --shm-size=1g -p 8888:8888 -p 9999:9999 \
  -e HINDSIGHT_API_LLM_API_KEY=$OPENAI_API_KEY \
  -v hindsight-data:/home/hindsight/.pg0 \
  ghcr.io/vectorize-io/hindsight:latest

Port 8888 is the API and port 9999 is the control plane, a web interface for browsing banks and testing queries. A real deployment splits this up: hindsight-api is stateless and holds all state in PostgreSQL, hindsight-worker runs background tasks off the same queue, and PostgreSQL doubles as the task broker. Point at your own database with HINDSIGHT_API_DATABASE_URL. The stated prerequisite is PostgreSQL 14 or newer with a vector extension, and the docs name pgvector, pgvectorscale, vchord and scann. There is also a Helm chart and a bare metal path.

helm install hindsight oci://ghcr.io/vectorize-io/charts/hindsight \
  --set api.llm.provider=groq \
  --set api.llm.apiKey=gsk_xxxxxxxxxxxx \
  --set postgresql.enabled=true

mem0 has two shapes. As a library (pip install mem0ai) it defaults to a local Qdrant at /tmp/qdrant, a SQLite history file at ~/.mem0/history.db, an OpenAI model for extraction, and OpenAI text-embedding-3-small for embeddings. As a self-hosted server it is a docker compose bundle: a REST API on port 8888, a dashboard on port 3000, and PostgreSQL with pgvector underneath. It wants two environment variables before it will start properly.

export OPENAI_API_KEY=sk-xxx
export JWT_SECRET=$(openssl rand -hex 32)

JWT_SECRET signs access and refresh tokens, so a short or reused value is a real authentication problem. There is an AUTH_DISABLED switch the docs mark for local development only. The dashboard gives you a live audit log of every API call, a memory browser, an entity list and per-user API keys you can revoke.

One practical collision: both servers default their REST API to port 8888. If you want both on one VPS while you evaluate them, remap one of them at the docker level before you start.

How much RAM does each need on a VPS?

Hindsight publishes a component-by-component floor in its installation documentation. These are the vendor's own stated minimums, not measurements taken here.

ChartHindsight published RAM floor per component, in MB
The data behind this chart
[
  {
    "label": "API, full image",
    "min_ram_mb": 1536,
    "recommended_ram_mb": 2048
  },
  {
    "label": "API, slim image",
    "min_ram_mb": 512,
    "recommended_ram_mb": 1024
  },
  {
    "label": "Control plane UI",
    "min_ram_mb": 128,
    "recommended_ram_mb": 256
  },
  {
    "label": "PostgreSQL",
    "min_ram_mb": 512,
    "recommended_ram_mb": 1024
  }
]

Add those up and the full image with its own database wants roughly 2 GB minimum and closer to 3 GB to be comfortable, which puts a 4 GB VPS at the sensible end. The full API image asks for 1536 MB on its own while the slim image asks for 512 MB, and PostgreSQL wants 1024 MB or more.

That gap between the two images has a cause worth understanding, because it is the same cause that decides your bill. The full image runs BAAI/bge-small-en-v1.5 for embeddings and cross-encoder/ms-marco-MiniLM-L-6-v2 for reranking in process, downloaded from HuggingFace on first run. Two transformer models resident in memory is what the extra gigabyte buys. The slim image drops them, so it needs less RAM and needs an external embedding service instead.

mem0's server is lighter for the opposite reason: by default it sends every embedding to OpenAI, so there is no model in your process. Your floor is PostgreSQL plus a small Python API, which fits a 2 GB VPS. You pay for that in a network call on every write and every search, and in the fact that the text leaves your machine. Point mem0's embedder at a local Ollama instead and the RAM comes back, plus whatever Ollama's own model needs. There is no free version of this trade. Either the model lives on your box or the text goes to someone else's.

Which operations call an LLM, and what does that cost you?

Both systems bill you on writes and give you cheap reads, which is the single most useful thing to know before you wire either one into a chat loop.

Hindsight calls an LLM on retain, because extracting facts, entities and temporal data out of raw text is a language task. It calls an LLM on reflect, because reflect synthesises an answer. It does not call an LLM on recall: vector search, BM25, graph traversal and temporal filtering are database and index work, and the cross-encoder rerank is a local model. So a read-heavy agent that only ever calls recall adds no token cost at all.

mem0 calls an LLM on add, for the same extraction reason. Its documentation states that search "converts your natural language question into a vector embedding, then finds memories with similar embeddings in your database", with optional keyword and entity signals and an optional reranker. No LLM in the base search path. The difference from Hindsight is the embedder: mem0's default embedder is a paid API call on both add and search, while Hindsight's default embedder is local.

The practical consequence is that your memory bill tracks how chatty your agent is, not how much it remembers. An agent that retains every turn of every conversation runs an extraction call per turn. Work out that number before you deploy, because it is the line item that surprises people. The trade-offs by memory class are laid out in what each kind of agent memory costs to keep, and the same arithmetic decides when it is cheaper to prune a memory than to keep re-reading it.

Can you run either one entirely on local models?

Yes, and both document it, though Hindsight goes further out of the box.

Hindsight lists OpenAI, Anthropic, Gemini, Groq, Ollama, AWS Bedrock and DeepSeek among its providers, plus over 100 more through LiteLLM, plus any OpenAI-compatible endpoint. Ollama is three environment variables.

export HINDSIGHT_API_LLM_PROVIDER=ollama
export HINDSIGHT_API_LLM_BASE_URL=http://localhost:11434/v1
export HINDSIGHT_API_LLM_MODEL=llama3

It also documents a built-in llama.cpp server: set HINDSIGHT_API_LLM_PROVIDER=llamacpp and there is no external dependency at all. Combined with the in-process embedder and reranker, that is a deployment with no outbound calls.

mem0's open source configuration takes an Ollama provider for both the LLM and the embedder.

config = {
    "llm": {
        "provider": "ollama",
        "config": {"model": "mixtral:8x7b", "temperature": 0.1, "max_tokens": 2000},
    },
}

You will need to supply the Ollama host yourself and size the VPS for the model, which is the part people underestimate. If a fully local stack is the requirement rather than a preference, compare both against a memory store designed to run entirely on the machine the agent runs on.

Does either ship MCP, and which coding agents work?

This is where the two differ most sharply for a self-hoster.

Hindsight's server ships MCP built in. The README states: "Every server ships a built-in Model Context Protocol endpoint, one per bank, enabled by default: http://localhost:8888/mcp/{bank_id}/". One endpoint per bank means the bank boundary is also the tool boundary, so an agent connected to one bank cannot read another. The integrations page counts 53 official and 6 community integrations, with a coding agents section listing Aider, Claude Code Skills, Continue, Cursor, Devin Desktop, GitHub Copilot, Grok Build, OMO, OpenHands, Roo Code, ZCode and Zed, plus around 30 framework and SDK integrations including LangGraph, CrewAI, Pydantic AI and the Vercel AI SDK.

mem0's MCP server is hosted by mem0. Its documentation gives the URL as https://mcp.mem0.ai/mcp and states plainly: "Nothing runs on your machine: the server is hosted by Mem0, and your client connects to it over HTTPS." Clients listed include Claude Desktop, Claude Code, Codex, Cursor, Windsurf, VS Code and OpenCode, added with one command.

npx mcp-add \
  --name mem0-mcp \
  --type http \
  --url "https://mcp.mem0.ai/mcp" \
  --clients "claude code,cursor,windsurf,vscode,opencode"

Read that carefully if your reason for self-hosting is that data must not leave your infrastructure. Using mem0's MCP server means using mem0's cloud. Pairing a self-hosted mem0 server with MCP is work you do yourself, and it sits alongside whatever else you are already running as MCP servers on your VPS.

What licence is each one under?

Hindsight is released by Vectorize.io under the MIT licence, which is short, permissive, and imposes no obligation beyond keeping the copyright notice. mem0 is released under Apache 2.0, which is also permissive and additionally includes an explicit patent grant and a requirement to note any changes you make to the source.

What does the self-hosted edition lack against the vendor's cloud?

For mem0, the docs publish a direct comparison, and the open source column is missing more than infrastructure. There is no app_id scope, no organisation or project structure with member roles, and no project-wide event feed (you get per-memory history(memory_id) only). The v3 search-time ranking features are platform only, which covers Graph Memory, Memory Decay, Temporal Reasoning and the background consolidation mem0 calls Dream. Custom categories, webhooks, memory export, batch operations, feedback and summaries are also platform side. The self-hosted server does include its own dashboard, API key issuance and request audit log, so it is not bare.

For Hindsight, the Cloud documentation frames its additions as operational rather than algorithmic: managed infrastructure, team management with role-based access control, usage analytics, credit-based billing, single sign-on, enforced multi-factor authentication, audit logs, webhook and SIEM delivery, and Memory Defense capabilities beyond the regex-based redaction in the open edition. The retain, recall and reflect engine itself is the MIT-licensed code.

That asymmetry matters more than any benchmark number. mem0 holds back retrieval quality features; Hindsight holds back team and compliance features.

What do the published accuracy numbers actually say?

Both projects publish accuracy claims, and both published them about themselves.

ChartSelf-published accuracy claims, read 30 September 2026 (percent)
The data behind this chart
[
  {
    "label": "LongMemEval",
    "hindsight": 94.6,
    "mem0": 94.4
  },
  {
    "label": "LoCoMo",
    "hindsight": 92.0,
    "mem0": 92.5
  },
  {
    "label": "BEAM 1M",
    "hindsight": 73.9,
    "mem0": 64.1
  },
  {
    "label": "BEAM 10M",
    "hindsight": 64.1,
    "mem0": 48.6
  }
]

Hindsight's figures come from its own benchmarks hub, which presents them as results on the Agent Memory Benchmark and claims the top position on every dataset there. Its README adds a reproduction claim: "The benchmark performance data for Hindsight has been independently reproduced by research collaborators at the Virginia Tech Sanghani Center for Artificial Intelligence and Data Analytics and The Washington Post." Named third parties reproducing a result is a stronger position than a vendor page alone, and it is still a claim made by the vendor about who reproduced it.

mem0's figures come from its own memory evaluation page, written by a party to the comparison. To its credit that page is explicit about scope. It states the results "use a single-pass retrieval setup (one retrieval call, one answer, no agentic loops) at a top_200 retrieval budget", that scores carry roughly a one point confidence interval because LLM judges are inconsistent, and, most importantly for you: "Scores reflect Mem0's managed platform, which includes proprietary optimizations not available in the open-source SDK."

So the LongMemEval figure of 94.4 is not a figure you can expect from the software you install. Hindsight's 94.6 is measured on a system whose engine is the MIT-licensed code, though the harness is still the vendor's. Do not read the four pairs above as a head to head. Two vendors scoring their own systems on benchmarks with different harnesses and different judges produce numbers that sit next to each other without being comparable.

Why these numbers cannot settle the question

Different evaluation harnesses make different choices about retrieval budget, judge model, prompt format and whether agentic loops are allowed. mem0 documents a top_200 budget and a single retrieval call. Hindsight reports through the Agent Memory Benchmark. Neither ran the other's harness. The only number that decides your deployment is the one you measure on your own data, with your own questions, and both projects are installable in under an hour, so measuring it is cheaper than arguing about it.

The two risks both of these share

A poisoned memory reaches every agent that reads it. Memory is durable by design, so text an attacker gets stored once is text your agent re-reads on every future task. In Hindsight the blast radius is one bank, because banks are strictly isolated and the MCP endpoint is per bank. In mem0 the blast radius is one scope (user_id, agent_id or run_id), because that is what search filters on. Hindsight documents an optional Memory Defense layer that scans retain operations against a set of PII and secret patterns, which is a leak control rather than an injection control. Neither project documents a filter that stops a plausible-sounding false fact from being stored. The defence is architectural: narrow scopes, and a review path for anything written from untrusted input. Work through how a poisoned memory gets written and what it does downstream before you let user content reach a retain or add call.

Deleting what the server holds about one person has to be an operation you can actually run. mem0 documents this well for self-hosters. delete(memory_id) removes one memory and delete_all(user_id="alice") removes everything under one scope, both available in the open source edition. Batch delete is platform only, and the documented full wipe uses all four filters set to "*", including the app_id that open source does not have.

client.delete(memory_id="your_memory_id")
client.delete_all(user_id="alice")

Hindsight's self-hosted API documents curation rather than erasure. A PATCH on a memory edits it, invalidates it (a soft retire that drops it from recall and the knowledge graph while keeping audit history), or restores it. There are delete endpoints for mental models, knowledge-base nodes and directives. The two-step erasure flow with a preview token and an erasure receipt, where a repeat within 24 hours returns the original receipt rather than re-running, is documented on the Hindsight Cloud side. If you self-host and you owe someone a deletion, check the delete paths in your own version's spec at /openapi.json before you promise a timeline, and design one bank per person so an erasure has a clean boundary. The general shape of that obligation is covered in how to actually erase one user's data from an agent memory store.

Pick Hindsight if, pick mem0 if

Pick Hindsight if the point of self-hosting is that data stays on your box. It is the only one of the two whose default embedder, reranker and MCP endpoint all run locally, and with the llama.cpp provider it runs with no outbound calls at all. Pick it if you want graph traversal and BM25 in retrieval without paying for a cloud tier, if you want an MCP endpoint per bank that your coding agent can point at today, and if you can spare 3 GB of RAM.

Pick mem0 if you want the smallest thing that works, if you are happy for extraction and embeddings to go to a hosted API, and if you value the Python and JavaScript SDK surface and the large integration ecosystem more than local execution. Pick it also if you expect to move to the managed platform later, because the open source edition is the same core engine and the migration is a config change rather than a rewrite. If that is your choice, the full walkthrough for standing up a mem0 memory server on a VPS covers the compose file, the reverse proxy and the API key handling.

Whichever you choose, run both against your own questions for a week before you commit. The published numbers were produced by the vendors who sell the products, and your data is not their benchmark.

FAQ

Can I run mem0's MCP server on my own VPS?

No. mem0's documentation describes the MCP server as hosted by mem0 at https://mcp.mem0.ai/mcp and states "Nothing runs on your machine". Connecting a coding agent to mem0 over MCP therefore means sending memory traffic to mem0's cloud, even if your mem0 server runs on your own box. Hindsight takes the other approach: its self-hosted server exposes a built-in MCP endpoint per bank at http://localhost:8888/mcp/{bank_id}/, enabled by default, so no third party is in the path.

Does self-hosted mem0 include graph memory?

Not as of 30 September 2026. mem0's graph memory is now native to the managed platform, and the docs say there is "no external graph database to provision", replacing the earlier Neo4j, Memgraph, Kuzu, Apache AGE and Neptune integrations. The platform versus open source page lists Graph Memory among the v3 search-time ranking features, alongside Memory Decay, Temporal Reasoning and Dream, none of which are in the open source edition. Hindsight's self-hosted recall does include graph traversal over entity and temporal links as one of its four parallel retrieval strategies.

How much RAM does each one need, and why do they differ?

Hindsight publishes a floor of 1536 MB for its full API image, 512 MB for the slim image, 512 MB for PostgreSQL and 128 MB for the control plane, so a realistic single-box install wants 3 GB. The full image is heavier because it runs BAAI/bge-small-en-v1.5 for embeddings and a MiniLM cross-encoder for reranking inside the API process. mem0's server is lighter because it sends embeddings to OpenAI by default, so a 2 GB VPS is enough. Move mem0 to a local embedder and its footprint rises to match.

Which memory operations cost money in tokens?

Writes, in both systems. Hindsight calls an LLM on retain (to extract facts and entities) and on reflect (to synthesise an answer), and calls no LLM on recall. mem0 calls an LLM on add, and its search path uses the embedder rather than an LLM, with reranking optional. So the driver of your bill is how many turns your agent stores, not how many it reads back. An agent that retains every conversation turn runs one extraction call per turn.

Can I trust the published accuracy numbers?

Treat them as vendor claims with useful detail attached. Hindsight publishes its results on its own benchmarks hub and states in its README that the data has been independently reproduced by research collaborators at the Virginia Tech Sanghani Center and The Washington Post. mem0 publishes its results on its own evaluation page, and that page states clearly that the scores reflect the managed platform "which includes proprietary optimizations not available in the open-source SDK", under a single-pass retrieval setup at a top_200 budget with about a one point judge-driven confidence interval. The two sets of numbers were produced by different harnesses, so they do not form a head to head, and neither settles which system is better on your data.