[{"data":1,"prerenderedAt":645},["ShallowReactive",2],{"tool-chroma-en":3},{"slug":4,"published":5,"minutes":6,"category":7,"tags":8,"keywords":14,"about":22,"sources":29,"cover":57,"og":58,"expertise":59,"locales":60,"lang":61,"title":64,"description":65,"coverAlt":66,"url":67,"pricing":68,"kind":69,"metaTitle":70,"takeaways":71,"faq":77,"toc":90,"blocks":118,"others":453},"chroma","2026-07-20",10,"rag",[9,10,11,12,13],"Vector search","Embeddings","RAG","Hybrid search","Metadata filter",[15,16,17,18,19,20,21],"chroma vector database","chroma vs qdrant","chroma vs pgvector","hybrid search reciprocal rank fusion","self-hosted vector database","chroma cloud pricing","metadata filtering vector search",[23,26],{"name":24,"url":25},"Chroma","https:\u002F\u002Fgithub.com\u002Fchroma-core\u002Fchroma",{"name":27,"url":28},"Apache License 2.0","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FApache_License",[30,33,36,39,42,45,48,51,54],{"title":31,"url":32},"Chroma documentation: introduction","https:\u002F\u002Fdocs.trychroma.com\u002Fdocs\u002Foverview\u002Fintroduction",{"title":34,"url":35},"Chroma architecture overview","https:\u002F\u002Fdocs.trychroma.com\u002Freference\u002Farchitecture\u002Foverview",{"title":37,"url":38},"Chroma: getting started with the SDK","https:\u002F\u002Fdocs.trychroma.com\u002Fdocs\u002Foverview\u002Fgetting-started",{"title":40,"url":41},"Chroma: query and get","https:\u002F\u002Fdocs.trychroma.com\u002Fdocs\u002Fquerying-collections\u002Fquery-and-get",{"title":43,"url":44},"Chroma: metadata filtering","https:\u002F\u002Fdocs.trychroma.com\u002Fdocs\u002Fquerying-collections\u002Fmetadata-filtering",{"title":46,"url":47},"Chroma Cloud: Search API overview","https:\u002F\u002Fdocs.trychroma.com\u002Fcloud\u002Fsearch-api\u002Foverview",{"title":49,"url":50},"Chroma Cloud: hybrid search with RRF","https:\u002F\u002Fdocs.trychroma.com\u002Fcloud\u002Fsearch-api\u002Fhybrid-search",{"title":52,"url":53},"Chroma Cloud: pricing","https:\u002F\u002Fwww.trychroma.com\u002Fpricing",{"title":55,"url":56},"Chroma changelog","https:\u002F\u002Fwww.trychroma.com\u002Fchangelog","\u002Fimages\u002Fblog\u002Fchroma\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fchroma\u002Fog.jpg","ai-engineer",[61,62,63],"en","de","hu","Chroma: simple vector search with one catch","Chroma is an Apache-2.0 vector database that runs embedded, single-node or as Chroma Cloud. Where it is pleasant, and where the good search features stop at the cloud boundary.","Cover: Chroma, a vector database for retrieval, with the retrieval path drawn as four stages from ingest to ranking","https:\u002F\u002Fwww.trychroma.com","Apache-2.0 · Cloud paid","Vector database","Chroma: simple vector search, cloud-only depth",[72,73,74,75,76],"Chroma is Apache-2.0 with about 27,000 GitHub stars and a single API across Python, TypeScript and Rust clients.","The embedded client starts inside the application process, so a first collection needs no server and no container.","Hybrid search with reciprocal rank fusion exists only in Chroma Cloud; the docs list single-node support as future work.","Cloud pricing is usage-based: $2.50 per GiB written, $0.33 per GiB stored per month, $0.0075 per TiB queried and $0.09 per GiB returned.","The honest engineering verdict: excellent default for a first retrieval layer, weak as the long-term home of a hybrid search stack.",[78,81,84,87],{"q":79,"a":80},"Is Chroma free to self-host?","Yes. The server is Apache-2.0 and runs as a local library, a single `chroma run` process or a Docker container. Chroma Cloud is the paid, managed layer on top, and the two share an API.",{"q":82,"a":83},"Does Chroma support hybrid search?","Only in Chroma Cloud. The Search API exposes reciprocal rank fusion over dense and sparse embeddings, and its documentation states plainly that single-node support is planned for a future release.",{"q":85,"a":86},"How large a corpus can one Chroma instance hold?","The architecture documentation puts single-node Chroma at fewer than 10 million records across a handful of collections. Larger workloads are meant for the distributed deployment behind Chroma Cloud.",{"q":88,"a":89},"Chroma or pgvector?","pgvector wins when the vectors already live in a Postgres row and transactional consistency matters. Chroma wins when retrieval is the product: richer search modes, multi-modal collections and a client that embeds for you.",[91,94,97,100,103,106,109,112,115],{"id":92,"title":93},"what-it-is","What it is",{"id":95,"title":96},"how-it-works","How it works",{"id":98,"title":99},"getting-started","Getting started",{"id":101,"title":102},"hybrid-search","Hybrid search and the cloud line",{"id":104,"title":105},"pricing","Pricing and what it costs per query",{"id":107,"title":108},"self-hosting","Self-hosting: what you actually run",{"id":110,"title":111},"where-it-shingles","Where it falls short",{"id":113,"title":114},"verdict","Verdict",{"id":116,"title":117},"sources","Sources",[119,123,126,129,132,148,149,152,161,164,165,168,171,174,210,211,214,216,227,228,231,299,302,303,306,318,323,324,327,394,397,398,401,416,422,423],{"type":120,"content":121},"paragraph",[122],"Chroma is an open-source vector database, Apache-2.0 licensed, that runs as an embedded library inside an application, as a single `chroma run` process, or as Chroma Cloud. The position this review takes is straightforward: it is the best default for the first retrieval layer of a RAG system, because the friction between an empty repository and a working query is close to zero, and it is a poor long-term home for a serious hybrid search stack, because the interesting parts of the query language are documented as cloud-only.",{"type":120,"content":124},[125],"In the stack it occupies one slot: storage and retrieval for embeddings. It competes with pgvector when the application already runs Postgres, with Qdrant or Milvus when retrieval needs its own service, and with Pinecone when nobody wants to operate anything. Chroma's distinctive bet is not the index but the ergonomics: an embedding function attached to a collection, a filter language that needs no index declarations, and clients for Python, TypeScript and Rust that all speak the same verbs.",{"type":127,"level":128,"id":92,"text":93},"heading",2,{"type":120,"content":130},[131],"The data model has three levels. A tenant holds databases, a database holds collections, and a collection is the unit of storage and querying: each item has an id, an embedding, and optionally a document and a metadata object. Access control, quota and billing are scoped at the tenant level, which matters for multi-tenant SaaS and is quietly absent from most single-node vector stores.",{"type":133,"ordered":134,"items":135},"list",false,[136,138,140,142,144,146],[137],"Licence: Apache 2.0, with roughly 27,000 stars on the GitHub repository and a tagged 1.5.9 release from May 2026.",[139],"Three deployment modes: an embedded library, a single node, and a distributed deployment that Chroma Cloud operates.",[141],"Three clients — Python, TypeScript and Rust — with the same collection, add, query and get shape, plus an async HTTP client.",[143],"Four search modes in one collection: dense vectors, sparse vectors, full-text and regex over documents, and metadata filtering.",[145],"Embedding functions as an interface, so a collection embeds on write and `query_texts` needs no model call in application code.",[147],"Two query APIs: the classic `query` and `get`, and a newer Search API with ranking expressions, available in Chroma Cloud.",{"type":127,"level":128,"id":95,"text":96},{"type":120,"content":150},[151],"Chroma delegates durability to subsystems it does not have to reinvent: SQLite locally, and cloud object storage in the distributed build, where hot data sits in SSD caches. The vendor's argument is economic rather than exotic — vectors are large and memory is expensive, so keeping the source of truth in object storage at a fraction of the cost of RAM is what makes the cloud price list possible.",{"type":153,"attrs":154,"inner":158,"caption":159},"diagram",{"viewBox":155,"role":156,"aria-labelledby":157},"0 0 720 372","img","chroma-t chroma-d","\u003Ctitle id=\"chroma-t\">The Chroma retrieval path\u003C\u002Ftitle>\u003Cdesc id=\"chroma-d\">Four stages. Ingest: chunks arrive with ids, documents and metadata, and an embedding function produces the vectors. Index: one collection holds a vector index, a sparse index, a full-text index and a metadata index side by side. Query: a vector, a where filter on metadata and a where_document filter on the stored text narrow the candidate set before ranking. Rank: the classic query API returns distances, while the cloud Search API can fuse dense and sparse rankings with reciprocal rank fusion. Below them a bar: single node keeps everything in one process over SQLite, Chroma Cloud runs a distributed engine on object storage with an SSD cache.\u003C\u002Fdesc>\u003Ctext x=\"20\" y=\"28\" class=\"d-title\">The Chroma retrieval path\u003C\u002Ftext>\u003Ctext x=\"700\" y=\"28\" text-anchor=\"end\" class=\"d-label\">four stages\u003C\u002Ftext>\u003Crect x=\"20\" y=\"46\" width=\"164\" height=\"64\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"102\" y=\"73\" text-anchor=\"middle\" class=\"d-text\">Ingest\u003C\u002Ftext>\u003Ctext x=\"102\" y=\"95\" text-anchor=\"middle\" class=\"d-small\">chunks, ids, metadata\u003C\u002Ftext>\u003Cpath d=\"M102 110 V124\" class=\"d-line\" \u002F>\u003Crect x=\"20\" y=\"124\" width=\"164\" height=\"150\" rx=\"10\" class=\"d-box d-dash\" \u002F>\u003Ctext x=\"102\" y=\"152\" text-anchor=\"middle\" class=\"d-small\">embedding function\u003C\u002Ftext>\u003Ctext x=\"102\" y=\"178\" text-anchor=\"middle\" class=\"d-small\">id required\u003C\u002Ftext>\u003Ctext x=\"102\" y=\"204\" text-anchor=\"middle\" class=\"d-small\">documents optional\u003C\u002Ftext>\u003Ctext x=\"102\" y=\"230\" text-anchor=\"middle\" class=\"d-small\">upsert, not just add\u003C\u002Ftext>\u003Crect x=\"192\" y=\"46\" width=\"164\" height=\"64\" rx=\"10\" class=\"d-sky\" \u002F>\u003Ctext x=\"274\" y=\"73\" text-anchor=\"middle\" class=\"d-text\">Index\u003C\u002Ftext>\u003Ctext x=\"274\" y=\"95\" text-anchor=\"middle\" class=\"d-small\">four indexes, one collection\u003C\u002Ftext>\u003Cpath d=\"M274 110 V124\" class=\"d-line\" \u002F>\u003Crect x=\"192\" y=\"124\" width=\"164\" height=\"150\" rx=\"10\" class=\"d-box d-dash\" \u002F>\u003Ctext x=\"274\" y=\"152\" text-anchor=\"middle\" class=\"d-small\">vector index\u003C\u002Ftext>\u003Ctext x=\"274\" y=\"178\" text-anchor=\"middle\" class=\"d-small\">sparse vector\u003C\u002Ftext>\u003Ctext x=\"274\" y=\"204\" text-anchor=\"middle\" class=\"d-small\">full text, regex\u003C\u002Ftext>\u003Ctext x=\"274\" y=\"230\" text-anchor=\"middle\" class=\"d-small\">metadata\u003C\u002Ftext>\u003Crect x=\"364\" y=\"46\" width=\"164\" height=\"64\" rx=\"10\" class=\"d-accent\" \u002F>\u003Ctext x=\"446\" y=\"73\" text-anchor=\"middle\" class=\"d-text\">Query\u003C\u002Ftext>\u003Ctext x=\"446\" y=\"95\" text-anchor=\"middle\" class=\"d-small\">vector, where, document\u003C\u002Ftext>\u003Cpath d=\"M446 110 V124\" class=\"d-line\" \u002F>\u003Crect x=\"364\" y=\"124\" width=\"164\" height=\"150\" rx=\"10\" class=\"d-box d-dash\" \u002F>\u003Ctext x=\"446\" y=\"152\" text-anchor=\"middle\" class=\"d-small\">query, get\u003C\u002Ftext>\u003Ctext x=\"446\" y=\"178\" text-anchor=\"middle\" class=\"d-small\">n_results, include\u003C\u002Ftext>\u003Ctext x=\"446\" y=\"204\" text-anchor=\"middle\" class=\"d-small\">$in, $and, $or\u003C\u002Ftext>\u003Ctext x=\"446\" y=\"230\" text-anchor=\"middle\" class=\"d-small\">$contains on arrays\u003C\u002Ftext>\u003Crect x=\"536\" y=\"46\" width=\"164\" height=\"64\" rx=\"10\" class=\"d-gold\" \u002F>\u003Ctext x=\"618\" y=\"73\" text-anchor=\"middle\" class=\"d-text\">Rank\u003C\u002Ftext>\u003Ctext x=\"618\" y=\"95\" text-anchor=\"middle\" class=\"d-small\">distances or fused scores\u003C\u002Ftext>\u003Cpath d=\"M618 110 V124\" class=\"d-line\" \u002F>\u003Crect x=\"536\" y=\"124\" width=\"164\" height=\"150\" rx=\"10\" class=\"d-box d-dash\" \u002F>\u003Ctext x=\"618\" y=\"152\" text-anchor=\"middle\" class=\"d-small\">classic: distances\u003C\u002Ftext>\u003Ctext x=\"618\" y=\"178\" text-anchor=\"middle\" class=\"d-small\">Search: ascending\u003C\u002Ftext>\u003Ctext x=\"618\" y=\"204\" text-anchor=\"middle\" class=\"d-small\">RRF, cloud only\u003C\u002Ftext>\u003Ctext x=\"618\" y=\"230\" text-anchor=\"middle\" class=\"d-small\">return_rank is required\u003C\u002Ftext>\u003Crect x=\"20\" y=\"292\" width=\"680\" height=\"64\" rx=\"10\" class=\"d-mint\" \u002F>\u003Ctext x=\"360\" y=\"319\" text-anchor=\"middle\" class=\"d-text\">One process over SQLite on a single node, or a distributed engine on object storage\u003C\u002Ftext>\u003Ctext x=\"360\" y=\"341\" text-anchor=\"middle\" class=\"d-small\">Same API in both modes · the Search API is the exception, and it is the point of difference\u003C\u002Ftext>",[160],"The API is identical in both modes until the Search API enters, and that is exactly where the two products stop being interchangeable.",{"type":120,"content":162},[163],"What the vendor does not publish is a public benchmark with methodology behind it. The Cloud documentation claims that production systems exceed 90 per cent recall and that the storage design makes Chroma Cloud an order of magnitude cheaper than alternatives; both are plausible and neither is reproducible from the docs. Treat them as a starting point for your own evaluation, not as a result.",{"type":127,"level":128,"id":98,"text":99},{"type":120,"content":166},[167],"The embedded client starts a server inside the process and loses everything when the program exits, which is exactly what you want for a test and exactly what you do not want for production. A minimal query needs a client, a collection, an upsert and a query — no server, no container, no index tuning.",{"type":169,"code":170},"code","import chromadb\n\nclient = chromadb.Client()\ncollection = client.get_or_create_collection(\"docs\")\n\ncollection.upsert(\n    ids=[\"d1\", \"d2\"],\n    documents=[\n        \"Refunds are issued within five business days.\",\n        \"Support answers within one working day.\",\n    ],\n    metadatas=[{\"topic\": \"billing\"}, {\"topic\": \"support\"}],\n)\n\nhits = collection.query(\n    query_texts=[\"how fast is a refund?\"],\n    where={\"topic\": \"billing\"},\n    n_results=1,\n)\n\n# Results are column-major: one list per query, one entry per result.\nprint(hits[\"documents\"][0][0])",{"type":120,"content":172},[173],"Three details in that snippet cause most of the friction later. Results come back column-major, so every consumer has to zip parallel arrays rather than iterate records. `n_results` defaults to 10, which silently returns nothing useful on a small test collection. And the filter goes in `where` against metadata, while text matching against the stored document goes in `where_document` with `$contains` or a regex.",{"type":175,"variant":176,"title":177,"body":178},"callout","note","Filter operators",[179],[180,181,184,185,188,189,184,192,195,196,184,199,202,203,184,206,209],"The metadata filter supports comparison operators, ",{"tag":169,"children":182},[183],"$gt"," and ",{"tag":169,"children":186},[187],"$lte",", membership with ",{"tag":169,"children":190},[191],"$in",{"tag":169,"children":193},[194],"$nin",", boolean composition with ",{"tag":169,"children":197},[198],"$and",{"tag":169,"children":200},[201],"$or",", and, since February 2026, ",{"tag":169,"children":204},[205],"$contains",{"tag":169,"children":207},[208],"$not_contains"," over array metadata fields. No index declaration is needed, which is convenient until a filter is slow and there is nothing to tune.",{"type":127,"level":128,"id":101,"text":102},{"type":120,"content":212},[213],"The Search API replaces `query` and `get` with a composable expression: `Search` builds the filter and the limit, `Knn` supplies a ranking, and `Rrf` fuses several rankings. Reciprocal rank fusion scores each candidate as the negative sum of weight divided by the smoothing constant plus its rank, with k defaulting to 60, which is why it works across dense and sparse results without normalising two different score scales.",{"type":169,"code":215},"from chromadb import Search, K, Knn, Rrf\n\ndense = Knn(\n    query=\"how fast is a refund?\",\n    key=\"#embedding\",\n    return_rank=True,\n    limit=200,\n)\nsparse = Knn(\n    query=\"how fast is a refund?\",\n    key=\"sparse_embedding\",\n    return_rank=True,\n    limit=200,\n)\n\nsearch = (\n    Search()\n    .where(K(\"topic\") == \"billing\")\n    .rank(Rrf(ranks=[dense, sparse], weights=[0.7, 0.3], k=60))\n    .limit(10)\n    .select(K.DOCUMENT, K.SCORE)\n)\n\nrows = collection.search(search).rows()[0]\nfor row in rows:\n    print(row[\"score\"], row[\"document\"][:60])",{"type":175,"variant":217,"title":218,"body":219},"warn","Three traps in the Search API",[220],[221,222,226],"It is documented as Chroma Cloud only, with single-node support planned for a future release, so this code does not run against a self-hosted server. ",{"tag":223,"children":224},"strong",[225],"return_rank=True"," is mandatory on every Knn inside an Rrf — omit it and the components contribute distances instead of ranks, producing a plausible, silently wrong ranking. And scores are sorted ascending, so lower is better, unlike the distances the classic API returns.",{"type":127,"level":128,"id":104,"text":105},{"type":120,"content":229},[230],"Chroma Cloud bills four separate meters, and the interesting one is not the vector store. Writes are $2.50 per GiB, storage is $0.33 per GiB per month, queries are $0.0075 per TiB, and network egress is $0.09 per GiB returned. A Starter plan costs nothing per month and includes $5 in credits, ten databases and ten team members; Team is $250 per month with $100 in credits, a hundred databases, thirty members and SOC II.",{"type":232,"head":233,"rows":253},"table",[234,239,244,249],[235,236,237,238],"Cloud plan","Price","Usage meter","Included",[240,241,242,243],"Starter","$0 per month","write, storage, query, network","$5 credits, 10 databases, 10 members",[245,246,247,248],"Team","$250 per month","same four meters","$100 credits, 100 databases, 30 members, SOC II",[250,251,247,252],"Enterprise","Custom","unlimited databases, single tenant, BYOC, SLAs",[254,263,272,281,290],[255,257,259,261],[256],"Write",[258],"per GiB written",[260],"$2.50",[262],"billed once per ingest",[264,266,268,270],[265],"Storage",[267],"per GiB per month",[269],"$0.33",[271],"vectors plus documents plus metadata",[273,275,277,279],[274],"Query",[276],"per TiB queried",[278],"$0.0075",[280],"effectively free at most corpus sizes",[282,284,286,288],[283],"Network",[285],"per GiB returned",[287],"$0.09",[289],"the meter that punishes large results",[291,293,295,297],[292],"Self-hosted",[294],"$0",[296],"your own disk",[298],"Apache-2.0, single node, no Search API",{"type":120,"content":300},[301],"The vendor's own calculator is instructive. At 1536 dimensions with 8 KiB documents across 500 collections, one million documents cost about $34 to write, six million stored documents about $27 a month, and ten million queries about $19 — roughly $79 a month for a corpus of six million chunks. Note the direction of the economics: 1 GiB of text becomes about 15 GiB of vectors, so storage is billed on the number that inflates fastest, and returning whole documents rather than short chunks is what pushes the egress meter up.",{"type":127,"level":128,"id":107,"text":108},{"type":120,"content":304},[305],"Self-hosting is not one thing. The architecture documentation separates three modes with different ceilings, and the difference between them is larger than the branding suggests.",{"type":133,"ordered":134,"items":307},[308,310,312,314,316],[309],"Embedded: `chromadb.Client()` runs in-process, keeps data in memory or on a local path, and dies with the program. Fine for tests, not for a service.",[311],"Single node: `chroma run --path` behind an `HttpClient`, documented at fewer than 10 million records across a handful of collections. Durable, simple, single point of failure.",[313],"Distributed: the deployment Chroma Cloud operates, with object storage persistence and SSD caches, pinned to a single region per database.",[315],"Data residency: Chroma Cloud runs in AWS us-east-1 and, since April 2026, GCP europe-west1. The region is fixed at creation and moving means creating a new database and reindexing.",[317],"Compliance: Chroma Cloud is SOC 2 Type II certified; customer-managed encryption keys arrived in December 2025 and private networking in January 2026.",{"type":175,"variant":176,"title":319,"body":320},"Region choice is permanent",[321],[322],"The EU region is offered on all plans and carries full API parity for collections, search and forking — but Chroma Sync, the Chroma CLI and the Search Agent are US-only at launch. If a plan depends on those, the region decision made at creation time cannot be undone.",{"type":127,"level":128,"id":110,"text":111},{"type":120,"content":325},[326],"The weaknesses come first, because they decide the choice. Hybrid ranking is cloud-only. Metadata filtering is convenient rather than tunable, and a filter that slows down gives you no index to adjust. The result shape is column-major, which adds a small but permanent tax on every consumer. And the price list punishes exactly the pattern that retrieval-augmented generation encourages: returning a large context window per query.",{"type":232,"head":328,"rows":352},[329,333,337,342,347],[330,331,12,332],"","Deployment","Lock-in profile",[24,334,335,336],"embedded, single node, cloud","RRF in Chroma Cloud only","Apache-2.0 server; the Search API is not",[338,339,340,341],"pgvector","extension inside existing Postgres","manual: vector index plus SQL full text","Postgres licence; nothing to leave",[343,344,345,346],"Qdrant","self-hosted or Cloud","native dense plus sparse","Apache-2.0; filter-heavy design to learn",[348,349,350,351],"Pinecone","managed only","native in the hosted API","no self-hosted path",[353,363,374,383],[354,356,358,360,361],[355],"Licence",[357],"Apache 2.0",[359],"PostgreSQL licence",[357],[362],"proprietary",[364,366,368,370,372],[365],"Operational cost",[367],"lowest",[369],"none, same instance",[371],"a service to run",[373],"none",[375,377,379,381,382],[376],"Dimensional limits",[378],"none documented",[380],"2,000 on vector, 4,000 on halfvec",[378],[378],[384,386,388,390,392],[385],"Best fit",[387],"first retrieval layer",[389],"vectors beside the row",[391],"filter-heavy retrieval",[393],"no operations at all",{"type":120,"content":395},[396],"The comparison that matters is with pgvector. If the corpus already lives in a Postgres table and the vectors belong in the same transaction, a dedicated vector store is an extra service to back up, secure and monitor for a capability Postgres can already provide, with the hard ceiling of 2,000 dimensions on the standard vector type and 4,000 with half-precision. Chroma earns its place when retrieval is the product: multi-modal collections, regex over stored documents, collection forking for experiments, and an embedding function that removes a model call from application code.",{"type":127,"level":128,"id":113,"text":114},{"type":120,"content":399},[400],"Chroma is a well-judged default with a specific blind spot. The ergonomics are real: three clients, one API, a filter language that needs no schema, and a path from `import chromadb` to a ranked result in about ten lines. The blind spot is equally real: the query features that separate a demo from a retrieval system are the ones you cannot self-host today. That is a reasonable trade for a first layer and a poor one for a system expected to last.",{"type":133,"ordered":402,"items":403},true,[404,406,408,410,412,414],[405],"Pick it when the retrieval layer is still being designed and the fastest path to a measurable baseline matters most.",[407],"Pick it when embeddings, documents and metadata must live in one collection and the team is multi-tenant from the start.",[409],"Pick it for prototyping that will later be replaced; the collection API is small enough that a rewrite is a day, not a quarter.",[411],"Skip it when hybrid search has to run inside your own network boundary today — the Search API is not available there.",[413],"Skip it when the corpus is already in Postgres and pgvector's dimension ceiling does not bite.",[415],"Reconsider it once the corpus passes a few million chunks and single-node ceilings, or once egress per GiB returned dominates the bill.",{"type":175,"variant":417,"title":418,"body":419},"tip","The one thing to get right",[420],[421],"Attach an explicit embedding function to the collection and pin the model version in your own config. The default embedding function is convenient and will change under you, and re-embedding a corpus is the one migration in this stack that is genuinely expensive.",{"type":127,"level":128,"id":116,"text":117},{"type":133,"ordered":402,"items":424},[425,429,432,435,438,441,444,447,450],[426],{"tag":427,"href":32,"children":428},"a",[31],[430],{"tag":427,"href":35,"children":431},[34],[433],{"tag":427,"href":38,"children":434},[37],[436],{"tag":427,"href":41,"children":437},[40],[439],{"tag":427,"href":44,"children":440},[43],[442],{"tag":427,"href":47,"children":443},[46],[445],{"tag":427,"href":50,"children":446},[49],[448],{"tag":427,"href":53,"children":449},[52],[451],{"tag":427,"href":56,"children":452},[55],[454,503,555,595],{"slug":455,"published":456,"minutes":457,"category":7,"tags":458,"keywords":463,"about":472,"sources":476,"cover":495,"og":496,"expertise":59,"locales":497,"lang":61,"title":498,"description":499,"coverAlt":500,"url":501,"pricing":502,"kind":459},"zep","2026-10-06",11,[459,460,461,462,11],"Agent memory","Knowledge graph","Temporal graph","Context engineering",[464,465,466,467,468,469,470,471],"zep ai","zep agent memory","graphiti knowledge graph","zep pricing","zep vs mem0","long-term memory for agents","temporal knowledge graph","zep cloud",[473],{"name":474,"url":475},"Zep","https:\u002F\u002Fwww.getzep.com\u002F",[477,480,483,486,489,492],{"title":478,"url":479},"Zep pricing: plans, credits and limits","https:\u002F\u002Fwww.getzep.com\u002Fpricing",{"title":481,"url":482},"Zep documentation","https:\u002F\u002Fhelp.getzep.com\u002F",{"title":484,"url":485},"Graphiti on GitHub","https:\u002F\u002Fgithub.com\u002Fgetzep\u002Fgraphiti",{"title":487,"url":488},"Graphiti product page","https:\u002F\u002Fwww.getzep.com\u002Fplatform\u002Fgraphiti\u002F",{"title":490,"url":491},"Announcing a new direction for Zep's open-source strategy","https:\u002F\u002Fwww.getzep.com\u002Fblog\u002Fannouncing-a-new-direction-for-zeps-open-source-strategy\u002F",{"title":493,"url":494},"Graphiti: temporal knowledge graphs for AI agents (arXiv)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2501.13956","\u002Fimages\u002Fblog\u002Fzep\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fzep\u002Fog.jpg",[61,62,63],"Zep review: agent memory on a temporal graph","Zep is a hosted agent-memory API on a temporal knowledge graph: credits on writes, retrieval free, Flex from $125 a month, Graphiti as the part you can self-host.","Diagram of how a fact reaches the prompt in Zep: messages and facts are extracted into a per-user context graph of entities and relationships, and retrieval walks the graph to return a context block with the supporting facts.","https:\u002F\u002Fwww.getzep.com","Apache-2.0 core · Cloud from $50 per month",{"slug":504,"published":505,"minutes":506,"category":7,"tags":507,"keywords":509,"about":517,"sources":523,"cover":548,"og":549,"expertise":59,"locales":550,"lang":61,"title":551,"description":552,"coverAlt":553,"url":554,"pricing":68,"kind":69},"lancedb","2026-09-21",9,[9,12,508,11],"Embedded database",[504,510,511,512,513,514,515,516],"lancedb review","lance vector database","embedded vector database","lancedb vs qdrant","hybrid search rrf","lancedb indexing","lance data format",[518,520],{"name":69,"url":519},"https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FVector_database",{"name":521,"url":522},"Apache Arrow","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FApache_Arrow",[524,527,530,533,536,539,542,545],{"title":525,"url":526},"LanceDB quickstart","https:\u002F\u002Fdocs.lancedb.com\u002Fquickstart",{"title":528,"url":529},"LanceDB vector indexes","https:\u002F\u002Fdocs.lancedb.com\u002Findexing\u002Fvector-index",{"title":531,"url":532},"LanceDB indexing guide","https:\u002F\u002Fdocs.lancedb.com\u002Findexing\u002Findex",{"title":534,"url":535},"LanceDB hybrid search","https:\u002F\u002Fdocs.lancedb.com\u002Fsearch\u002Fhybrid-search",{"title":537,"url":538},"LanceDB Enterprise","https:\u002F\u002Fdocs.lancedb.com\u002Fenterprise",{"title":540,"url":541},"LanceDB frequently asked questions","https:\u002F\u002Fdocs.lancedb.com\u002Ffaq\u002Ffaq-oss",{"title":543,"url":544},"LanceDB pricing","https:\u002F\u002Flancedb.com\u002Fpricing",{"title":546,"url":547},"LanceDB on PyPI","https:\u002F\u002Fpypi.org\u002Fproject\u002Flancedb\u002F","\u002Fimages\u002Fblog\u002Flancedb\u002Fcover.webp","\u002Fimages\u002Fblog\u002Flancedb\u002Fog.jpg",[61,62,63],"LanceDB: vector search that starts as a library","A review of LanceDB: an Apache-2.0 embedded vector library, its IVF and HNSW index choices, hybrid search with rank fusion, and what the Enterprise tier adds.","Cover art for the LanceDB review: one Lance table feeding a vector index and a full-text index into a fused ranking","https:\u002F\u002Flancedb.com",{"slug":338,"published":505,"minutes":6,"category":7,"tags":556,"keywords":560,"about":567,"sources":573,"cover":588,"og":589,"expertise":59,"locales":590,"lang":61,"title":591,"description":592,"coverAlt":593,"url":569,"pricing":359,"kind":594},[9,557,558,11,559],"Postgres","HNSW","Quantisation",[338,561,562,563,564,565,566],"pgvector vs qdrant","postgres vector search","hnsw index postgres","iterative index scans","binary quantization postgres","vector database postgres",[568,570],{"name":338,"url":569},"https:\u002F\u002Fgithub.com\u002Fpgvector\u002Fpgvector",{"name":571,"url":572},"PostgreSQL","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FPostgreSQL",[574,576,579,582,585],{"title":575,"url":569},"pgvector README",{"title":577,"url":578},"pgvector changelog","https:\u002F\u002Fgithub.com\u002Fpgvector\u002Fpgvector\u002Fblob\u002Fmaster\u002FCHANGELOG.md",{"title":580,"url":581},"PostgreSQL news: pgvector 0.8.2 released","https:\u002F\u002Fwww.postgresql.org\u002Fabout\u002Fnews\u002Fpgvector-082-released-3245\u002F",{"title":583,"url":584},"AWS: Scale pgvector with binary quantization","https:\u002F\u002Faws.amazon.com\u002Fblogs\u002Fdatabase\u002Fscale-pgvector-with-binary-quantization-on-amazon-aurora-postgresql\u002F",{"title":586,"url":587},"pgvector licence","https:\u002F\u002Fgithub.com\u002Fpgvector\u002Fpgvector\u002Fblob\u002Fmaster\u002FLICENSE","\u002Fimages\u002Fblog\u002Fpgvector\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fpgvector\u002Fog.jpg",[61,62,63],"pgvector, reviewed: the vector database you do not have to run","A review of pgvector 0.8.7: iterative scans for filtered search, HNSW and IVFFlat, binary quantisation at 100M vectors, and the CVE that made index builds a patch item.","A query enters at the top and splits into an exact sequential scan, an HNSW graph walk and an IVFFlat probe; a band below shows an iterative scan continuing until the limit is full.","Vector database extension",{"slug":596,"published":597,"minutes":6,"category":7,"tags":598,"keywords":600,"about":606,"sources":610,"cover":638,"og":639,"expertise":59,"locales":640,"lang":61,"title":641,"description":642,"coverAlt":643,"url":609,"pricing":644,"kind":459},"mem0","2026-09-17",[459,599,11,9],"Long-term memory",[596,601,602,603,604,469,605],"mem0 review","agent memory layer","mem0 self-hosted","mem0 pricing","mem0 alternatives",[607],{"name":608,"url":609},"Mem0","https:\u002F\u002Fmem0.ai",[611,614,617,620,623,626,629,632,635],{"title":612,"url":613},"Mem0 documentation","https:\u002F\u002Fdocs.mem0.ai\u002Fintroduction",{"title":615,"url":616},"Mem0 quickstart","https:\u002F\u002Fdocs.mem0.ai\u002Fquickstart",{"title":618,"url":619},"How Mem0 works","https:\u002F\u002Fdocs.mem0.ai\u002Fcore-concepts\u002Fhow-it-works",{"title":621,"url":622},"Mem0 pricing","https:\u002F\u002Fmem0.ai\u002Fpricing",{"title":624,"url":625},"Mem0 on GitHub","https:\u002F\u002Fgithub.com\u002Fmem0ai\u002Fmem0",{"title":627,"url":628},"mem0ai on PyPI","https:\u002F\u002Fpypi.org\u002Fproject\u002Fmem0ai\u002F",{"title":630,"url":631},"Mem0 research and benchmarks","https:\u002F\u002Fmem0.ai\u002Fresearch",{"title":633,"url":634},"Mem0 MCP server","https:\u002F\u002Fdocs.mem0.ai\u002Fplatform\u002Fmem0-mcp",{"title":636,"url":637},"Mem0 paper on arXiv","https:\u002F\u002Farxiv.org\u002Fabs\u002F2504.19413","\u002Fimages\u002Fblog\u002Fmem0\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fmem0\u002Fog.jpg",[61,62,63],"Mem0: what an agent memory layer costs per turn","A review of Mem0: facts extracted from every turn, the April 2026 benchmark table and its platform-only caveat, four cloud tiers and what self-hosting leaves out.","A loop that turns conversation into stored facts and reads them back into the prompt","Free tier · from $19 per month",1791383548852]