Tools/RAG & retrieval

Mem0: what an agent memory layer costs per turn

A review of Mem0: facts extracted from every turn, the April 2026 benchmark table and its platform-only caveat, four cloud tiers and what self-hosting leaves out.

Type
Agent memory
Pricing
Free tier · from $19 per month

··10 min read

  • Agent memory
  • Long-term memory
  • RAG
  • Vector search
A loop that turns conversation into stored facts and reads them back into the prompt

Key takeaways

  • Mem0 turns conversation into stored facts: an LLM extracts, deduplicates and embeds them, so every write costs a model call on top of the storage write.
  • The current algorithm scores 92.5 on LoCoMo and 94.4 on LongMemEval, and the README states those numbers come from the managed platform rather than the open-source SDK.
  • Cloud tiers are Hobby for free, Starter at $19 a month and Pro at $249 a month, with graph memory and Dream consolidation gated behind Pro.
  • Open-source retrieval has no graph memory and boosts on entity overlap alone, so self-hosting buys the API rather than the benchmark tables.
  • The free tier allows 10,000 add and 1,000 retrieval requests a month, which is roughly thirty-three searches a day.

Mem0 is a memory layer for agents: conversation goes in, facts come out, and the facts are fetched back before the next model call. The position taken here is that it is the most carefully built memory API in this market and a bad default at the same time, because an LLM call sits on the write path — remembering is billed per turn, not per user. It earns its place when memories have to outlive sessions and be filtered per user, agent and run; a rolling summary and a table beat it when they do not.

It sits between the application and the model, replacing the habit of appending transcript to the prompt, and competes with Zep, LangMem, Letta and whatever a team builds from a database and a summarisation step. It ships in three shapes under one API: an Apache-2.0 library, a self-hosted server behind docker compose with authentication on by default, and a hosted platform with a dashboard, metered plans and an MCP endpoint.

What it is

Two products share the name. The open-source repository is the mem0ai library for Python and JavaScript, embedded in a process; the platform is a hosted API with scopes, plans and a dashboard, which the same SDK reaches over HTTPS at api.mem0.ai. The homepage listed 62,590 GitHub stars in early October 2026.

  • Published on PyPI as mem0ai, version 2.2.1, under Apache-2.0, with a matching JavaScript package.
  • Three deployment modes in the README: library for a prototype, a docker-compose server with authentication on by default for a team, and the cloud platform for zero operations.
  • Operations are add, search, get_all, update and delete, each scoped by user_id, agent_id, app_id or run_id.
  • Writes are inferred: an LLM extracts durable facts, deduplicates and embeds them, and infer=False stores the raw message instead.
  • Defaults are gpt-5-mini for extraction and text-embedding-3-small for embeddings, with Qdrant as the packaged vector store.
  • Hosted MCP server at mcp.mem0.ai exposing eleven memory tools, authenticated by browser sign-in or by an API key as a bearer token.
  • Compliance listed on the homepage as SOC 2 Type 1, HIPAA and GDPR, with BYOK and on-premises deployment on the Enterprise tier.

How it works

The pipeline is asymmetric, and that asymmetry is the product. On add, Mem0 looks up related memories so the same fact is not stored twice, extracts new facts with an LLM, deduplicates and embeds them, pulls out entities, and writes to a SQL store for facts, a vector store for embeddings and an entity store for links. On search, four signals are scored and fused: semantic similarity, keyword matching, entity overlap, and a temporal signal scored from metadata written at extraction time.

How a memory is written and readA conversation box on the left feeds an extraction box in the middle, labelled one LLM call and ADD only, which produces facts rather than transcripts. An arrow continues to a deduplicate box that embeds the facts and links entities. From there the flow fans out into three stores: a SQL store for facts and metadata, a vector store for embeddings, and an entity store for boosting. The three stores merge into a search box at the bottom, labelled semantic, keyword, entity and time. A long return arrow runs from the search box back to the conversation box, labelled memories into the prompt, which closes the loop. A label at the top right reads one model call on the write path.How a memory is written and readone model call on the write pathConversationuser and assistant turnsExtractionone LLM call, ADD onlyfacts, not transcriptsDeduplicateembed and link entitiesSQL storefacts and metadataVector storeembeddingsEntity storelinks for boostingSearchsemantic, keyword, entity, timememories intothe prompt
Facts are extracted once per add and read back by a scoped search; the open-source build has no entity graph, so its entity signal is overlap on extracted terms.

The documentation is explicit about the split that follows. The platform fuses all four signals with graph-backed entity matching; the open-source build has no graph memory and boosts on entity overlap alone, depending on whichever vector store is configured. Extraction is additive as well — a new fact does not silently overwrite an old one — so corrections have to be explicit update or delete calls, which is the right default for an audit trail and an annoying one for a user who just changed their mind.

Getting started

The hosted path is two calls. Install the SDK, take an API key from the dashboard, and give every call a scope: a memory without a user_id, an agent_id or a run_id will be handed to whoever asks for it next.

from mem0 import MemoryClient

client = MemoryClient(api_key="your-api-key")

messages = [
    {"role": "user", "content": "I'm a vegetarian and allergic to nuts."},
    {"role": "assistant", "content": "Noted."},
]
client.add(messages, user_id="user123")

results = client.search(
    "What are my dietary restrictions?",
    filters={"user_id": "user123"},
)
for r in results["results"]:
    print(r["memory"], r["score"])

The open-source library has the same shape and no server behind it: from mem0 import Memory, then add, search, update and delete against stores you operate yourself. It needs an LLM key and a vector store before the first call, which is the real difference between the two halves — the library is Apache-2.0, and the retrieval that makes the benchmark tables lives on the platform.

Benchmarks

Mem0 publishes its own numbers, which is more than most memory layers do, and the README annotates the method: single-pass retrieval, one call with no agentic loop, a top-200 retrieval budget, on the same production-representative model stack.

BenchmarkBefore April 2026Current scoreTokens per queryp50 latency
LoCoMo71.492.57.0K0.88 s
LongMemEval67.894.46.8K1.09 s
BEAM, 1M tokensnot published64.16.7K1.00 s
BEAM, 10M tokensnot published48.66.9K1.05 s

The caveat sits in the same paragraph: those scores reflect the managed platform, which carries proprietary optimizations that are not in the open-source SDK, and open-source users should expect directionally similar rather than identical numbers. The paper behind it, posted to arXiv in April 2025, reports 91% lower p95 latency and more than 90% token savings against full-context prompting on LoCoMo, a 26% relative gain in LLM-as-a-Judge over OpenAI's memory, and about 2% more from the graph variant.

Pricing

The pricing page meters two counters, add requests and retrieval requests, and end users are unlimited on every tier. The figures below are what the page listed in early October 2026; usage-based pricing exists for traffic that does not fit a tier.

PlanPriceAdd requests per monthRetrieval requests per month
HobbyFree10,0001,000
Starter$19 a month50,0005,000
Pro$249 a month500,00050,000
EnterpriseCustomUnlimitedUnlimited

Two features that carry the marketing — graph memory for entity linking and Dream consolidation — are listed under Pro and Enterprise rather than in the free tiers, and the README's own comparison table describes the self-hosted server's advanced features as teasers. Hobby's 1,000 retrievals work out to about thirty-three searches a day: enough to prove an integration, not enough to run a product.

Where it falls short

The weaknesses are structural rather than cosmetic. Every add is a model call, so write latency and the token bill both scale with conversation volume, and the extraction step can store an inference the user never stated — a derived fact is harder to audit than a logged message. The features that make Mem0 interesting live on the platform, which means the open-source build and the paid product share a name more than a feature set. Lock-in is real but modest: memories are rows readable back through the API, and the move to the v3 API has a published migration guide.

OptionMemory modelOperational burdenRight for
Mem0Extracted facts, scoped by user, agent and runA library to run, or a platform to payMulti-session personalisation at a known cost per conversation
Zep with GraphitiTemporal knowledge graph of episodesA graph store and its upkeepRelationship-heavy memory where order in time matters
LangMemMemory stores and semantic search in the LangChain stackPart of the LangChain toolchainTeams already committed to LangChain
A table of factsRows you write yourself, plus pgvectorYour own SQL and a summary promptShort-lived memory and strict audit requirements

Which of these fits depends on how much memory the product needs, not on which row scores highest on LoCoMo. A support agent with a few dozen durable facts per account does not need a memory layer at all; a consumer assistant holding a year of history cannot afford to re-read the transcript.

Verdict

Adopt it when memory is a product feature that outlives sessions and the write-path model call can be priced into the product. Skip it when a session's context fits into a summary, or when every stored fact needs provenance a human can check, because inference on the way in is the design rather than a side effect.

  1. Start with the library: the API has the same shape, and it keeps the move to the platform open.
  2. Scope every call. An unscoped memory is the one failure mode that crosses users.
  3. Budget the write path before the tier: one extraction call per add is the cost model, whatever the plan costs.
  4. Take the platform for graph memory, temporal ranking and fused retrieval — the parts the open-source build does not have.
  5. Never store secrets or unredacted personal data; retrieval is designed to surface what is in the store.
Avoid storing secrets, raw credentials, or unredacted sensitive data. Mem0 is designed to retrieve stored context.

Sources

  1. Mem0 documentation
  2. Mem0 quickstart
  3. How Mem0 works
  4. Mem0 pricing
  5. Mem0 on GitHub
  6. mem0ai on PyPI
  7. Mem0 research and benchmarks
  8. Mem0 MCP server
  9. Mem0 paper on arXiv

Frequently asked questions

Is Mem0 free to self-host?

The library is Apache-2.0 and the README offers a docker-compose server with authentication on by default, so there is no licence fee. It still needs an LLM for extraction, gpt-5-mini by default, and a vector store, Qdrant by default, so the cost moves to inference and operations rather than to a subscription.

Mem0 or a summary of the transcript in Postgres?

If a product holds one short session per user, a rolling summary plus a table of facts is cheaper and easier to audit, and it has no extraction step that can invent a fact. Mem0 earns its place when memories outlive sessions, have to be filtered per user, agent and run, and are read on the hot path of every request.

Does the open-source version reach the published benchmark scores?

No, and the repository says so: the numbers reflect the managed platform, which contains proprietary optimizations that are not in the open-source SDK. Open-source retrieval has no graph memory and depends on the vector store that is configured.

How does Mem0 reach an agent that already speaks MCP?

Through a hosted server at mcp.mem0.ai over HTTPS, exposing eleven tools including add_memory, search_memories, update_memory and delete_memory. It authenticates with a browser sign-in flow or an API key sent as a bearer token, and nothing runs on the developer's machine.

Sounds like what you need?

Tell me about your project or role – I’d love to hear from you.