Tools/RAG & retrieval
Zep review: agent memory on a temporal graph
Zep is a hosted agent-memory API on a temporal knowledge graph: credits on writes, retrieval free, Flex from $125 a month, Graphiti as the part you can self-host.
- Type
- Agent memory
- Pricing
- Apache-2.0 core · Cloud from $50 per month
Balázs Csorba··11 min read
- Agent memory
- Knowledge graph
- Temporal graph
- Context engineering
- RAG

Key takeaways
- Zep is a hosted agent-memory API that stores facts in a temporal knowledge graph and returns the context an assistant should see on the next turn.
- Credits are charged on writes only, one credit per episode up to 350 bytes, while retrieval, storage, threads and users are free.
- Self-serve plans are Free with 10,000 credits, Flex at $125 for 50,000 credits and Flex Plus at $375 for 200,000, with SOC 2 Type II and a HIPAA BAA reserved for Enterprise.
- Zep Community Edition was discontinued in April 2025, so Graphiti is the only part that can be self-hosted.
- Extraction runs a model on every write, which makes cost a function of how talkative the agent is and makes extractor errors silently wrong context.
Zep is an agent-memory service: the application sends messages and facts, the service keeps a temporal knowledge graph per user, and each turn it returns the context an assistant should see next. It is the right purchase for a team whose agents fail because they forget, and the wrong one for a team that has to run everything itself, because the only self-hostable part is the Graphiti library.
It sits between the application and the model, replacing hand-rolled summary buffers, vector stores full of old chat logs and the memory helpers inside LangChain and LlamaIndex. It competes with Mem0, with LangMem, with a Postgres table plus embeddings, and with the long-term memory endpoints the model vendors are adding; the difference is that Zep stores facts with a valid-from and valid-to time instead of a flat list of past messages.
What it is
The product is a managed API: user.add for durable facts, thread.create and thread.add_messages for conversation, graph.get_user_context for retrieval. Ingestion is where the work happens, an extraction pass turns each episode into entities, relations and observations stamped with time, and retrieval assembles a context block together with the facts that support it. The open-source half is Graphiti, an Apache-2.0 library that performs the same temporal graph maintenance, including fact-level incremental updates rather than a nightly full re-index.
- Licence and ownership: Graphiti is Apache-2.0 and maintained by Zep Software; the Zep memory service itself is a closed, hosted product.
- Self-hosting: Zep Community Edition was discontinued in April 2025 and its code was moved to a legacy folder, so the graph engine can run locally while the full service cannot.
- Metering: one credit per episode up to 350 bytes, one more per further 350 bytes or part thereof, and an eighth of a credit per webhook invocation.
- Not metered: retrieval, storage, threads, users and graph storage cost nothing, so a read-heavy agent stays cheap to keep running.
- Performance: the vendor reports p95 retrieval latency of 148 ms on a 10,000-node graph and 168 ms on a 100-million-node graph.
- Benchmarks: vendor-reported 94.7% on LoCoMo and 90.2% on LongMemEval, two long-conversation memory benchmarks.
- Deployment: cloud, cloud with customer-managed keys, or bring-your-own-cloud inside a customer VPC; SOC 2 Type II and a HIPAA BAA belong to the Enterprise tier.
How it works
Every message or fact sent is an episode. The extraction pass asks a model to identify entities and relations, writes them into the graph with a validity window, and touches only the nodes that episode affects: a correction closes the previous interval instead of overwriting it. Retrieval then walks the graph around the subject of the current question, takes the still-valid facts and the recent episodes, and returns a context string within a token budget.
The design choice that matters is temporal rather than vector. A fact has a lifetime, so the graph can answer what is true now, what changed and when, and which of two contradictory statements supersedes the other. A summarising buffer cannot do that without a second system, and pure vector retrieval returns whatever is textually similar, including the stale version of a fact corrected last week.
Getting started
The hosted quickstart is four calls. The snippet below stores a fact, opens a thread, appends a message and asks for the context an assistant should receive.
from zep_cloud.client import Zep
client = Zep(api_key="zep-...")
# durable facts, extracted into the context graph
client.user.add(user_id="alice", fact="Alice leads the migration off the legacy CRM")
# episodic history: writes are charged as credits by size, retrieval is free
client.thread.create(thread_id="t-1", user_id="alice")
client.thread.add_messages(
thread_id="t-1",
messages=[{"role": "user", "content": "Where did we stop on the CRM migration?"}],
)
# what the assistant should see on this turn
context = client.graph.get_user_context(user_id="alice", max_facts=25)
print(context.context)Two things behave differently from a vector store. First, ingestion is asynchronous and charged, so the natural pattern is to send messages as they happen and read context once per turn. Second, retrieval returns prose with its supporting facts rather than chunk ids, which moves the summarisation decision out of your prompt and into the service; convenient, but the quality of what the model now sees is Zep's extraction quality as much as your own prompt.
Performance and cost
Cost is credits, and credits are bytes: a 700-byte message costs two credits, so Flex with 50,000 credits a month holds roughly 25,000 average-sized messages before overage at $25 per 10,000 credits. The unusual part is that retrieval is unmetered, because that is the call an agent makes every turn; the meter runs only when memory is written.
| Operation | Credit cost | Charged when | What to watch |
|---|---|---|---|
| Episode up to 350 bytes | 1 | On every write: message, fact or JSON payload | Raw tool output pasted into messages is the usual overrun |
| Each further 350 bytes | +1 | Per episode, rounded up | A 1,200-byte episode already costs 4 credits |
| Webhook invocation | 1/8 | Where webhooks are enabled | Chatty automations add up quietly |
| Retrieval, storage, users | 0 | Never | Read-heavy agents stay cheap |
The self-serve plans are Free with 10,000 credits a month, Flex at $125 for 50,000 credits and Flex Plus at $375 for 200,000, with overage at $25 per 10,000 and $75 per 40,000 credits respectively. Latency figures are vendor-reported: p95 retrieval of 148 ms at 10,000 nodes and 168 ms at 100 million, which says the graph does not visibly degrade with size and says nothing about the latency of the extraction step that runs before it.
- Size the plan on writes rather than reads: messages a day, average bytes, times credits, times thirty.
- Strip raw tool payloads before they become episodes; a summarised tool result costs a fraction of a transcript.
- Set a fact budget on retrieval so the context block never grows into a system prompt nobody reads.
None of that is unusual for a memory vendor, they all bill writes and promise cheap reads. The specific risk here is the extraction step, because whatever the extractor misses cannot be retrieved later: there is no second pass over the raw messages once they are reduced to facts.
Pricing
The pricing page lists credit-based self-serve plans and a negotiated Enterprise tier. The paid entry point on the current page is Flex at $125 a month, and below it sits only the free tier; Enterprise adds custom credits, guaranteed rate limits and the compliance paperwork.
- Free: 10,000 credits a month, two projects, one Memory MCP server seat, variable rate limits, no rollover.
- Flex: $125 a month for 50,000 credits, then $25 per 10,000, 600 requests per minute, five projects, 30-day rollover, community support.
- Flex Plus: $375 a month for 200,000 credits, then $75 per 40,000, 1,000 requests per minute, observations, webhooks, analytics and seven-day API logs.
- Enterprise: custom credits and rates, SOC 2 Type II, HIPAA BAA, one-year audit and API logs, DPA for EU customers, deployment inside your own VPC.
Where it shingles
Start with the deployment weakness: since Community Edition was discontinued in April 2025, the graph engine, Graphiti, runs locally but the product does not, so an air-gapped or regulated deployment is an Enterprise conversation and the self-serve tiers do not include SOC 2 Type II or a BAA. Second, memory is only as good as the extraction, and the extraction is a model call, so a hallucinated relation or a missed update is not a queryable error but silently wrong context. Third, credit metering makes cost a function of how talkative the agent is rather than of how many users it serves. Fourth, a graph is a different debugging surface from a chat log: explaining why the model saw a fact requires the graph, not the transcript.
| Tool | What it is | Where it wins | What you give up |
|---|---|---|---|
| Zep | Hosted memory API over a temporal context graph | Facts with validity windows and vendor-reported p95 under 170 ms | No self-hosted product, and every write costs credits |
| Mem0 | Open-source memory layer with a hosted option | Simpler API and a server you can run yourself | Flatter memory model with less temporal bookkeeping |
| LangMem and LangGraph | Memory primitives inside a graph framework | Memory next to the agents and stores you already run | Storage, extraction and retrieval have to be assembled by you |
| Postgres plus embeddings | Chat history in a table, similarity search over it | No new vendor and full control of the data | No temporal reasoning, so superseded facts stay retrievable |
The real question is whether forgetting is the failure mode. If agents repeat themselves, lose decisions between sessions or cannot say what changed, a graph with validity windows answers it directly. If retrieval keeps returning the wrong document, memory is the wrong layer altogether, and the same effort spent on chunking and reranking pays back sooner.
Verdict
Zep is a good product with a clear price of admission: it solves forgetting well, and what it costs is the ability to run the whole thing yourself. Teams shipping agents to paying users, where the visible failure is a customer being asked the same question twice, should take it. Teams with a data-residency constraint should build on Graphiti and own the service around it.
- Take it when agents fail by forgetting, and the failure shows up as repeated questions or decisions lost between sessions.
- Take it when memory writes are human messages: the credit model is cheapest when episodes are short and infrequent.
- Do not take it as a self-hosted deployment, since April 2025 only Graphiti is yours to run and everything around it lives in Zep's cloud.
- Do not take it when retrieval quality rather than memory is the problem, because no memory layer fixes a bad index.
- Budget the extraction step: every write is a model call, so cost tracks how verbose the agent is, not how many users it serves.
Agent memory is a state-management problem wearing an AI costume: a graph that gives facts a lifetime is exactly what a summary buffer cannot provide, and the bill is a model call for every message you decide to remember.
Sources
Frequently asked questions
How does Zep charge for usage?
By credits on writes: an episode up to 350 bytes costs one credit, a 640-byte episode costs two and a 1,200-byte episode costs four. Retrieval, storage, threads, users and graph storage cost nothing, so a read-heavy agent keeps running cheaply.
Can Zep be self-hosted?
Not as a product. Zep Community Edition was discontinued in April 2025 and its code sits in a legacy folder; Graphiti, the Apache-2.0 temporal graph library, can run locally, and the full service is available inside your own VPC only through the Enterprise tier.
How much does Zep cost?
The free tier gives 10,000 credits a month, Flex is $125 a month for 50,000 credits with overage at $25 per 10,000, and Flex Plus is $375 for 200,000 credits. Enterprise pricing is negotiated and adds SOC 2 Type II, a HIPAA BAA and deployment options.
What is the difference between Zep and a vector store for chat history?
A vector store returns what is textually similar, including superseded facts, while Zep stores entities and relations with validity windows, so retrieval can return what is true now and what changed. The trade-off is a model call on every write, and its mistakes become context.