Erase a user's data from agent memory
A deletion request touches more than one table. Find every copy of a user's data in a self-hosted agent memory stack, delete it, and prove it is gone.
What a deletion request really asks for
To erase a user's data from agent memory you have to delete the same fact in several places, because a memory stack stores one fact about a person four or five times over: as a row in a relational table, as a vector in an index, sometimes as a node in a graph, and as the raw conversation line it was extracted from. Deleting the row removes the agent's answer. It leaves the material the agent used to produce that answer, so the next extraction run writes the fact back.
The stack in this guide is the common self-hosted one: Postgres with pgvector or Qdrant holding the embeddings, a memory layer such as Mem0 on top, and the logs, caches and backups around both. Work in the order below. The backup step comes last because it is the one that decides whether your deletion claim is honest. This is engineering and not legal advice: it is the work that has to exist before anyone can promise a user that their data is gone.
The inventory: every place a fact about a person comes to rest
Write this list for your own stack before you run a single DELETE. Each line is a place where a fact can survive its deletion somewhere else.
- The memory rows themselves, usually one table of extracted facts keyed by a user id.
- The vector index, which holds the embedding of each fact and, in most setups, a copy of the text in a payload or a joined column.
- The graph store, if you enabled one, where a fact becomes nodes and relations that outlive the sentence they came from.
- The raw conversation log: the
messagestable, the JSONL transcripts on disk, the job payloads still sitting in Redis. - The memory layer's own audit log. Mem0's open-source
Memorykeeps one, and it stores the text of anything you delete. - Caches: an embedding cache keyed by a hash of the text, a Redis session, a materialised view, the agent's scratch files under
/opt. - Derived artefacts: nightly summaries, extracted profile traits, evaluation sets and test fixtures exported from real conversations, fine-tuning files.
- Logs outside the app: reverse proxy access logs carrying the user id in a query string, application logs that print whole prompts, the systemd journal.
- Backups and snapshots:
pg_dumpfiles, filesystem or volume snapshots, Qdrant collection snapshots, and the copy someone pulled to a laptop to debug an incident. - Processors you do not host: the model provider that saw the prompt, an error tracker holding breadcrumbs, product analytics.
The copies people forget are the last four. A vector index that has stopped returning a row in search results can still hold that row's vector on disk, and last night's backup is a complete working copy of everything you are about to delete.
Scope the subject by one stable key
Pick one identifier that every store carries, and scope every delete by that key alone. Matching on a name or an email substring is how you delete somebody else's data by accident, because two users share a surname and a free-text LIKE cannot tell them apart. Start by counting what you are about to remove.
SELECT count(*) FROM memories WHERE user_id = 'u_8412';
SELECT count(*) FROM messages WHERE user_id = 'u_8412';
SELECT count(*) FROM memories
WHERE user_id <> 'u_8412' AND content ILIKE '%ada lovelace%';The third query is the awkward one. It finds facts about your subject stored under a different person's key, because one user told the agent something about another person. Those sentences are still the subject's data, and deleting them edits a different user's record. Decide the rule now and write it down. The usual answer is to redact the identifying words inside the text, keep the row, and record that you did it.
If the person has two accounts, or changed email at some point, resolve that into a list of keys before you start and run every step below for each key in the list.
Step 1: delete the source records
Wrap the live deletes in one transaction, so a half-finished run cannot leave the memory row gone and the transcript in place. Record the request in a log that holds keys and timestamps only, never the deleted content, because an audit trail that quotes what you erased is a copy of what you erased.
BEGIN;
INSERT INTO erasure_log (subject_key, requested_at, ticket)
VALUES ('u_8412', now(), 'req-4471');
DELETE FROM memories WHERE user_id = 'u_8412';
DELETE FROM messages WHERE user_id = 'u_8412';
COMMIT;Stop the extraction and summarisation jobs before you open that transaction. A summariser that fires while you are working reads the transcript and writes a fresh memory row for a user you have just cleared, and nothing in your logs will look like a fault.
On Mem0's open-source library the same scope is one call. Read the memory list first and keep it, because those ids are the only handle you will have on the audit log afterwards.
from mem0 import Memory
m = Memory.from_config(config_dict)
before = m.get_all(filters={"user_id": "u_8412"}, top_k=1000)
m.delete_all(user_id="u_8412")delete_all walks the matching memories and removes each one from the vector store. Hold before as a file outside the stack for the length of this job, then delete that file too.
Why the deleted memory is still in history.db
Mem0's open-source Memory writes an event log to SQLite, at ~/.mem0/history.db unless MEM0_DIR or history_db_path moves it. As of September 2026 the table is created like this.
CREATE TABLE IF NOT EXISTS history (
id TEXT PRIMARY KEY,
memory_id TEXT,
old_memory TEXT,
new_memory TEXT,
event TEXT,
created_at DATETIME,
updated_at DATETIME,
is_deleted INTEGER,
actor_id TEXT,
role TEXT
)Deleting a memory adds a row whose event is DELETE and whose old_memory holds the memory's previous text in full, so the sentence you just removed is still on disk in plain text. Notice what the table does not have: a user_id column. That means you cannot scope this cleanup by subject after the fact, which is why step 1 saved the ids.
sqlite3 ~/.mem0/history.dbPRAGMA secure_delete = ON;
DELETE FROM history WHERE memory_id IN ('id-one', 'id-two');
VACUUM;A plain DELETE in SQLite unlinks the row and leaves its bytes in a free page inside the file, so a strings pass over the file can still show the text. secure_delete overwrites freed content as it goes, and VACUUM rebuilds the file from the live rows. Run both, in that order.
If you enabled a graph store, the person's nodes need the same treatment, and the edges attached to them go with the node rather than surviving as dangling relations. Confirm which property your memory layer writes the subject key into, then delete by that property.
MATCH (n {user_id: 'u_8412'}) DETACH DELETE n;Step 2: make the vector index drop the vectors
This is the step most deletion scripts miss, because the index looks clean from the outside long before it is clean on disk.
In Postgres, DELETE marks a row dead. The tuple stays in its heap page and its entry stays in the HNSW (hierarchical navigable small world) index until vacuum runs. Search results are correct immediately, because the dead index entry fails its visibility check and the row never reaches the reader. The vector itself is still in the file. pgvector's README gives the order to fix that, and the reason: vacuuming an HNSW index is slow, so reindex first.
REINDEX INDEX CONCURRENTLY memories_embedding_hnsw_idx;
VACUUM memories;REINDEX INDEX CONCURRENTLY builds a fresh index from live rows only, so the deleted vectors are never copied into it. Plain VACUUM then marks the dead heap space reusable without rewriting the file, which means the old tuples stay readable at the page level until later writes land on top of them. VACUUM FULL memories; rewrites the table into new files and drops the old ones, at the cost of an exclusive lock and free disk space roughly the size of the table. pg_repack does the same rewrite without holding the table locked throughout, which is what you want on a box that is also serving the agent.
One more Postgres copy hides behind all of that. The delete's old row content is written into the write-ahead log, and those segments are recycled on their own schedule.
SELECT slot_name, active, restart_lsn FROM pg_replication_slots;An inactive slot left over from an experiment pins every WAL segment since it stalled, which means it also pins the rows you deleted. Drop the slots you do not use with SELECT pg_drop_replication_slot('slot_name');, and check whether an archive_command is shipping WAL to storage that has its own retention.
On Qdrant, delete by payload filter and wait for the write to be applied.
curl -X POST 'http://127.0.0.1:6333/collections/agent_memory/points/delete?wait=true' \
-H 'api-key: <key>' \
-H 'Content-Type: application/json' \
-d '{"filter": {"must": [{"key": "user_id", "match": {"value": "u_8412"}}]}}'That marks the points deleted. The vectors stay inside their segment until the optimizer rewrites that segment, and the optimizer has two conditions before it will bother. Qdrant's shipped config.yaml, as of September 2026, sets deleted_threshold: 0.2, described as the minimal fraction of deleted vectors in a segment, and vacuum_min_vector_number: 1000, the minimal number of vectors in a segment. Forty deleted points inside a segment of fifty thousand is well under that fraction, so no rewrite is triggered and the vectors remain on disk. Lower both for the collection, let the optimizer work, then set them back.
curl -X PATCH 'http://127.0.0.1:6333/collections/agent_memory' \
-H 'api-key: <key>' \
-H 'Content-Type: application/json' \
-d '{"optimizers_config": {"deleted_threshold": 0.001, "vacuum_min_vector_number": 100}}'A collection snapshot taken before the deletion is a full working copy of the deleted points, so list what exists and decide each one under the backup policy in step 4.
curl -s 'http://127.0.0.1:6333/collections/agent_memory/snapshots' -H 'api-key: <key>'
curl -X DELETE 'http://127.0.0.1:6333/collections/agent_memory/snapshots/<snapshot_name>' -H 'api-key: <key>'Step 3: derived data is still that person's data
A summary written from a deleted conversation is that person's data. An extracted trait, something like "prefers replies in Portuguese, works night shifts", is that person's data, even though they never typed those words. A stored embedding is that person's data as well: published work on embedding inversion reconstructs approximate source text from vectors, so treat a vector as the sentence it came from rather than as a one-way hash.
The failure that looks like a haunting is simpler than it seems. You delete the memory rows, keep the transcripts for analytics, and the nightly job reads the transcripts and writes the facts back. Nothing malfunctioned. The pipeline did its job on data you chose to keep, which means an agent that can re-learn a fact from a log you kept has not forgotten anything. The same coupling is what makes pruning agent memories that have gone stale harder than it looks: the derived layer keeps regenerating from the source layer until you cut one of them.
Model weights are the one place you cannot clean up after the fact. If you fine-tuned on memory content, you cannot remove one person from the weights, so pick a position in advance: either memory content never becomes training data, or a retrain is part of the cost of a deletion request. Evaluation sets are the quiet version of the same problem. A fixture file holding a real user's real sentences is a copy with no user id attached, so nothing in your deletion script can find it. Tag fixtures with the subject key at export time, or build them from synthetic text. The discipline that keeps credentials out of an agent's context applies here without changes: what never enters memory never has to be erased from it.
Step 4: write the backup policy before you need it
Editing last month's pg_dump by hand is a bad idea. You damage the only recovery path you have, and you cannot prove afterwards what you changed. So the policy has to be about time and process instead, and it has to be written down where the person restoring a backup will read it.
The first half is a retention window you state out loud. "Deleted from live systems within 72 hours of the request, and out of every backup within 35 days" is a claim you can keep, as long as your retention really is 35 days and nothing holds a yearly archive off to one side. Go and verify that. A quarterly snapshot sitting in an object storage bucket with versioning turned on makes the sentence false, and the same applies to volume snapshots your provider keeps on its own schedule. What to check on a storage VPS before you put personal data on it lists the questions worth asking your host before this becomes urgent.
The second half is a replay step. Keep the erasure log, keys and timestamps only, somewhere that is not the database being restored, and make the restore runbook end by re-applying it. Restoring a copy from before the request brings the subject's rows back, and the replay removes them again.
DELETE FROM memories m USING erasure_log e WHERE m.user_id = e.subject_key;
DELETE FROM messages m USING erasure_log e WHERE m.user_id = e.subject_key;Run both halves. The replay covers the copies you actually restore, and the retention window covers the copies you never touch. What is not a policy is silence: if your runbook says nothing about erasure, the first real restore quietly un-deletes every subject who ever asked.
Step 5: prove that nothing about the subject is retrievable
Run these against your own stack after the deletes, and treat any non-empty result as unfinished work.
- Count by key in every store from your inventory, including the job and audit tables you added yourself.
- Query the vector store semantically, using the subject's known facts as the query text rather than their id, and read the top twenty payloads. Deleting by id and then searching by id proves very little, because the interesting failure is a fact that survived under a different key.
- Start a fresh agent session and ask the questions that used to surface the memory. That is the test the user cares about.
- Search the filesystem for the key and for the distinctive strings, including the export and fixture directories.
- List every backup and snapshot that still contains the subject, with the date each one expires.
grep -rIl 'u_8412' /opt/agent/logs /opt/agent/exports /opt/agent/fixturesAny path that command prints is a copy you have not dealt with yet. The list from step 5 is the honest answer to "is it gone", and it should shrink to nothing on a date you can name.
An auditor asks for six things: the inventory of stores, the request with its identity check and its date, the script that ran and the output it produced, the erasure log entry, the backup retention policy together with the replay job, and the list of processors you notified. Keep the script in git and review it like any other code, because it runs rarely and it deletes irreversibly.
Build for this before the request arrives
- One subject key on every table, every vector payload and every graph node, written at insert time rather than backfilled later.
- Foreign keys with
ON DELETE CASCADEpointing at the subject row, so a table added next year cannot be quietly left out of the deletion path. - A retention window on transcripts and application logs, enforced by a job, because the cheapest data to erase is data you stopped keeping.
- One reviewed script that takes a key and performs every step, the vector store maintenance and the history database included.
- Derived artefacts tagged with the key they came from, evaluation sets included.
All of this is easier on a stack you own end to end. If you are still assembling one, running a Mem0 memory server on your own VPS puts the rows and the audit log in one place you can open with sqlite3 and inspect for yourself, and a pgvector HNSW index on a small VPS keeps the vectors inside the same Postgres transaction as the facts, which means one COMMIT covers the row and its embedding instead of two systems that can disagree about who has been forgotten.
FAQ
Does DELETE in Postgres remove the vector from disk?
No. DELETE marks the row dead, so it stops appearing in query results, while the tuple and its embedding stay in the heap page until vacuum runs. Plain VACUUM then marks that space reusable without rewriting the file. To have the bytes gone, follow the delete with REINDEX INDEX CONCURRENTLY on the vector index, then VACUUM FULL or pg_repack to rewrite the table, and check pg_replication_slots for an inactive slot pinning WAL segments that still hold the old row.
How do I check that a deleted memory is really out of the vector index?
Search semantically instead of by id: embed a sentence built from what you know about the subject, then read the top results. In Qdrant a delete by filter marks points deleted and the vectors stay inside their segment until the optimizer rewrites it, which the defaults deleted_threshold: 0.2 and vacuum_min_vector_number: 1000 can postpone for a long time after a small deletion. On pgvector an unvacuumed index entry does not return the row, because the visibility check drops it, so there the real question is whether reindex and vacuum have actually run.
Do I have to delete a user's data from my backups?
Editing an existing backup archive is impractical and it damages your recovery path, so handle it with policy instead. State a retention window after which every backup holding that person has expired, and add a replay step to the restore runbook that re-applies deletions from an erasure log kept outside the restored database. The defensible claim is "gone from live systems now, and out of every remaining copy by this date".
Is a summary written from a deleted conversation still that user's data?
Yes. A summary, an extracted trait and a stored embedding all describe the person, and work on embedding inversion shows approximate source text can be reconstructed from a vector, so a vector is not anonymous. Derived records need the subject key attached when they are written, because without it your deletion script has no way to find them. Keeping the source transcript carries a second cost: the extraction job reads it again and writes the fact straight back.