Tools/RAG & retrieval

Zep review: agent memory on a temporal graph

Zep is a hosted agent-memory API on a temporal knowledge graph: credits on writes, retrieval free, Flex from $125 a month, Graphiti as the part you can self-host.

Type
Agent memory
Pricing
Apache-2.0 core · Cloud from $50 per month

··11 min read

  • Agent memory
  • Knowledge graph
  • Temporal graph
  • Context engineering
  • RAG
Diagram of how a fact reaches the prompt in Zep: messages and facts are extracted into a per-user context graph of entities and relationships, and retrieval walks the graph to return a context block with the supporting facts.

Key takeaways

  • Zep is a hosted agent-memory API that stores facts in a temporal knowledge graph and returns the context an assistant should see on the next turn.
  • Credits are charged on writes only, one credit per episode up to 350 bytes, while retrieval, storage, threads and users are free.
  • Self-serve plans are Free with 10,000 credits, Flex at $125 for 50,000 credits and Flex Plus at $375 for 200,000, with SOC 2 Type II and a HIPAA BAA reserved for Enterprise.
  • Zep Community Edition was discontinued in April 2025, so Graphiti is the only part that can be self-hosted.
  • Extraction runs a model on every write, which makes cost a function of how talkative the agent is and makes extractor errors silently wrong context.

Zep is an agent-memory service: the application sends messages and facts, the service keeps a temporal knowledge graph per user, and each turn it returns the context an assistant should see next. It is the right purchase for a team whose agents fail because they forget, and the wrong one for a team that has to run everything itself, because the only self-hostable part is the Graphiti library.

It sits between the application and the model, replacing hand-rolled summary buffers, vector stores full of old chat logs and the memory helpers inside LangChain and LlamaIndex. It competes with Mem0, with LangMem, with a Postgres table plus embeddings, and with the long-term memory endpoints the model vendors are adding; the difference is that Zep stores facts with a valid-from and valid-to time instead of a flat list of past messages.

What it is

The product is a managed API: user.add for durable facts, thread.create and thread.add_messages for conversation, graph.get_user_context for retrieval. Ingestion is where the work happens, an extraction pass turns each episode into entities, relations and observations stamped with time, and retrieval assembles a context block together with the facts that support it. The open-source half is Graphiti, an Apache-2.0 library that performs the same temporal graph maintenance, including fact-level incremental updates rather than a nightly full re-index.

  • Licence and ownership: Graphiti is Apache-2.0 and maintained by Zep Software; the Zep memory service itself is a closed, hosted product.
  • Self-hosting: Zep Community Edition was discontinued in April 2025 and its code was moved to a legacy folder, so the graph engine can run locally while the full service cannot.
  • Metering: one credit per episode up to 350 bytes, one more per further 350 bytes or part thereof, and an eighth of a credit per webhook invocation.
  • Not metered: retrieval, storage, threads, users and graph storage cost nothing, so a read-heavy agent stays cheap to keep running.
  • Performance: the vendor reports p95 retrieval latency of 148 ms on a 10,000-node graph and 168 ms on a 100-million-node graph.
  • Benchmarks: vendor-reported 94.7% on LoCoMo and 90.2% on LongMemEval, two long-conversation memory benchmarks.
  • Deployment: cloud, cloud with customer-managed keys, or bring-your-own-cloud inside a customer VPC; SOC 2 Type II and a HIPAA BAA belong to the Enterprise tier.

How it works

Every message or fact sent is an episode. The extraction pass asks a model to identify entities and relations, writes them into the graph with a validity window, and touches only the nodes that episode affects: a correction closes the previous interval instead of overwriting it. Retrieval then walks the graph around the subject of the current question, takes the still-valid facts and the recent episodes, and returns a context string within a token budget.

How a fact reaches the promptMessages and facts written by the application are extracted into a per-user context graph of entities and relationships, each with a validity window. On every turn the retrieval call walks that graph, skips superseded statements and returns a context block plus the supporting facts.context graphentities and facts over timeAliceCRM migrationSnowflakeQ3 deadlineworks onsuperseded byget_user_context
The graph is kept per user and per thread, and retrieval assembles a context block rather than a similarity-ranked list of chunks.

The design choice that matters is temporal rather than vector. A fact has a lifetime, so the graph can answer what is true now, what changed and when, and which of two contradictory statements supersedes the other. A summarising buffer cannot do that without a second system, and pure vector retrieval returns whatever is textually similar, including the stale version of a fact corrected last week.

Getting started

The hosted quickstart is four calls. The snippet below stores a fact, opens a thread, appends a message and asks for the context an assistant should receive.

from zep_cloud.client import Zep

client = Zep(api_key="zep-...")

# durable facts, extracted into the context graph
client.user.add(user_id="alice", fact="Alice leads the migration off the legacy CRM")

# episodic history: writes are charged as credits by size, retrieval is free
client.thread.create(thread_id="t-1", user_id="alice")
client.thread.add_messages(
    thread_id="t-1",
    messages=[{"role": "user", "content": "Where did we stop on the CRM migration?"}],
)

# what the assistant should see on this turn
context = client.graph.get_user_context(user_id="alice", max_facts=25)
print(context.context)

Two things behave differently from a vector store. First, ingestion is asynchronous and charged, so the natural pattern is to send messages as they happen and read context once per turn. Second, retrieval returns prose with its supporting facts rather than chunk ids, which moves the summarisation decision out of your prompt and into the service; convenient, but the quality of what the model now sees is Zep's extraction quality as much as your own prompt.

Performance and cost

Cost is credits, and credits are bytes: a 700-byte message costs two credits, so Flex with 50,000 credits a month holds roughly 25,000 average-sized messages before overage at $25 per 10,000 credits. The unusual part is that retrieval is unmetered, because that is the call an agent makes every turn; the meter runs only when memory is written.

OperationCredit costCharged whenWhat to watch
Episode up to 350 bytes1On every write: message, fact or JSON payloadRaw tool output pasted into messages is the usual overrun
Each further 350 bytes+1Per episode, rounded upA 1,200-byte episode already costs 4 credits
Webhook invocation1/8Where webhooks are enabledChatty automations add up quietly
Retrieval, storage, users0NeverRead-heavy agents stay cheap

The self-serve plans are Free with 10,000 credits a month, Flex at $125 for 50,000 credits and Flex Plus at $375 for 200,000, with overage at $25 per 10,000 and $75 per 40,000 credits respectively. Latency figures are vendor-reported: p95 retrieval of 148 ms at 10,000 nodes and 168 ms at 100 million, which says the graph does not visibly degrade with size and says nothing about the latency of the extraction step that runs before it.

  • Size the plan on writes rather than reads: messages a day, average bytes, times credits, times thirty.
  • Strip raw tool payloads before they become episodes; a summarised tool result costs a fraction of a transcript.
  • Set a fact budget on retrieval so the context block never grows into a system prompt nobody reads.

None of that is unusual for a memory vendor, they all bill writes and promise cheap reads. The specific risk here is the extraction step, because whatever the extractor misses cannot be retrieved later: there is no second pass over the raw messages once they are reduced to facts.

Pricing

The pricing page lists credit-based self-serve plans and a negotiated Enterprise tier. The paid entry point on the current page is Flex at $125 a month, and below it sits only the free tier; Enterprise adds custom credits, guaranteed rate limits and the compliance paperwork.

  • Free: 10,000 credits a month, two projects, one Memory MCP server seat, variable rate limits, no rollover.
  • Flex: $125 a month for 50,000 credits, then $25 per 10,000, 600 requests per minute, five projects, 30-day rollover, community support.
  • Flex Plus: $375 a month for 200,000 credits, then $75 per 40,000, 1,000 requests per minute, observations, webhooks, analytics and seven-day API logs.
  • Enterprise: custom credits and rates, SOC 2 Type II, HIPAA BAA, one-year audit and API logs, DPA for EU customers, deployment inside your own VPC.

Where it shingles

Start with the deployment weakness: since Community Edition was discontinued in April 2025, the graph engine, Graphiti, runs locally but the product does not, so an air-gapped or regulated deployment is an Enterprise conversation and the self-serve tiers do not include SOC 2 Type II or a BAA. Second, memory is only as good as the extraction, and the extraction is a model call, so a hallucinated relation or a missed update is not a queryable error but silently wrong context. Third, credit metering makes cost a function of how talkative the agent is rather than of how many users it serves. Fourth, a graph is a different debugging surface from a chat log: explaining why the model saw a fact requires the graph, not the transcript.

ToolWhat it isWhere it winsWhat you give up
ZepHosted memory API over a temporal context graphFacts with validity windows and vendor-reported p95 under 170 msNo self-hosted product, and every write costs credits
Mem0Open-source memory layer with a hosted optionSimpler API and a server you can run yourselfFlatter memory model with less temporal bookkeeping
LangMem and LangGraphMemory primitives inside a graph frameworkMemory next to the agents and stores you already runStorage, extraction and retrieval have to be assembled by you
Postgres plus embeddingsChat history in a table, similarity search over itNo new vendor and full control of the dataNo temporal reasoning, so superseded facts stay retrievable

The real question is whether forgetting is the failure mode. If agents repeat themselves, lose decisions between sessions or cannot say what changed, a graph with validity windows answers it directly. If retrieval keeps returning the wrong document, memory is the wrong layer altogether, and the same effort spent on chunking and reranking pays back sooner.

Verdict

Zep is a good product with a clear price of admission: it solves forgetting well, and what it costs is the ability to run the whole thing yourself. Teams shipping agents to paying users, where the visible failure is a customer being asked the same question twice, should take it. Teams with a data-residency constraint should build on Graphiti and own the service around it.

  1. Take it when agents fail by forgetting, and the failure shows up as repeated questions or decisions lost between sessions.
  2. Take it when memory writes are human messages: the credit model is cheapest when episodes are short and infrequent.
  3. Do not take it as a self-hosted deployment, since April 2025 only Graphiti is yours to run and everything around it lives in Zep's cloud.
  4. Do not take it when retrieval quality rather than memory is the problem, because no memory layer fixes a bad index.
  5. Budget the extraction step: every write is a model call, so cost tracks how verbose the agent is, not how many users it serves.
Agent memory is a state-management problem wearing an AI costume: a graph that gives facts a lifetime is exactly what a summary buffer cannot provide, and the bill is a model call for every message you decide to remember.

Sources

  1. Zep pricing: plans, credits and limits
  2. Zep documentation
  3. Graphiti on GitHub
  4. Graphiti product page
  5. Announcing a new direction for Zep's open-source strategy
  6. Graphiti: temporal knowledge graphs for AI agents (arXiv)

Frequently asked questions

How does Zep charge for usage?

By credits on writes: an episode up to 350 bytes costs one credit, a 640-byte episode costs two and a 1,200-byte episode costs four. Retrieval, storage, threads, users and graph storage cost nothing, so a read-heavy agent keeps running cheaply.

Can Zep be self-hosted?

Not as a product. Zep Community Edition was discontinued in April 2025 and its code sits in a legacy folder; Graphiti, the Apache-2.0 temporal graph library, can run locally, and the full service is available inside your own VPC only through the Enterprise tier.

How much does Zep cost?

The free tier gives 10,000 credits a month, Flex is $125 a month for 50,000 credits with overage at $25 per 10,000, and Flex Plus is $375 for 200,000 credits. Enterprise pricing is negotiated and adds SOC 2 Type II, a HIPAA BAA and deployment options.

What is the difference between Zep and a vector store for chat history?

A vector store returns what is textually similar, including superseded facts, while Zep stores entities and relations with validity windows, so retrieval can return what is true now and what changed. The trade-off is a model call on every write, and its mistakes become context.

Sounds like what you need?

Tell me about your project or role – I’d love to hear from you.