[{"data":1,"prerenderedAt":694},["ShallowReactive",2],{"tool-qdrant-en":3},{"slug":4,"published":5,"minutes":6,"category":7,"tags":8,"keywords":14,"about":22,"sources":32,"cover":72,"og":73,"expertise":74,"locales":75,"lang":76,"title":79,"description":80,"coverAlt":81,"url":82,"pricing":83,"kind":24,"metaTitle":84,"takeaways":85,"faq":91,"toc":104,"blocks":129,"others":503},"qdrant","2026-05-26",10,"rag",[9,10,11,12,13],"Vector search","HNSW","Quantisation","Hybrid search","Filtering",[4,15,16,17,18,19,20,21],"qdrant vs pinecone","vector database comparison","hnsw index","filtered vector search","vector quantisation","hybrid search rrf","rag vector store",[23,26,29],{"name":24,"url":25},"Vector database","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FVector_database",{"name":27,"url":28},"Nearest neighbor search","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FNearest_neighbor_search",{"name":30,"url":31},"Retrieval-augmented generation","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FRetrieval-augmented_generation",[33,36,39,42,45,48,51,54,57,60,63,66,69],{"title":34,"url":35},"Qdrant documentation","https:\u002F\u002Fqdrant.tech\u002Fdocumentation\u002F",{"title":37,"url":38},"Qdrant: points","https:\u002F\u002Fqdrant.tech\u002Fdocumentation\u002Fmanage-data\u002Fpoints\u002F",{"title":40,"url":41},"Qdrant: collections","https:\u002F\u002Fqdrant.tech\u002Fdocumentation\u002Fconcepts\u002Fcollections\u002F",{"title":43,"url":44},"Qdrant: indexing","https:\u002F\u002Fqdrant.tech\u002Fdocumentation\u002Fmanage-data\u002Findexing\u002F",{"title":46,"url":47},"Qdrant: quantisation","https:\u002F\u002Fqdrant.tech\u002Fdocumentation\u002Fmanage-data\u002Fquantization\u002F",{"title":49,"url":50},"Qdrant: capacity planning","https:\u002F\u002Fqdrant.tech\u002Fdocumentation\u002Fcapacity-planning\u002F",{"title":52,"url":53},"Qdrant: optimise performance","https:\u002F\u002Fqdrant.tech\u002Fdocumentation\u002Fops-optimization\u002Foptimize\u002F",{"title":55,"url":56},"Qdrant: hybrid and multi-stage queries","https:\u002F\u002Fqdrant.tech\u002Fdocumentation\u002Fconcepts\u002Fhybrid-queries\u002F",{"title":58,"url":59},"Qdrant: filtering","https:\u002F\u002Fqdrant.tech\u002Fdocumentation\u002Fconcepts\u002Ffiltering\u002F",{"title":61,"url":62},"Qdrant: security and access control","https:\u002F\u002Fqdrant.tech\u002Fdocumentation\u002Fsecurity\u002F",{"title":64,"url":65},"Qdrant pricing","https:\u002F\u002Fqdrant.tech\u002Fpricing\u002F",{"title":67,"url":68},"Qdrant on GitHub","https:\u002F\u002Fgithub.com\u002Fqdrant\u002Fqdrant",{"title":70,"url":71},"Qdrant vector search benchmarks","https:\u002F\u002Fqdrant.tech\u002Fbenchmarks\u002F","\u002Fimages\u002Fblog\u002Fqdrant\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fqdrant\u002Fog.jpg","ai-engineer",[76,77,78],"en","de","hu","Qdrant: a vector database built around filtering","A review of Qdrant: filterable HNSW, four quantisation methods, the memory tiers in v1.19 and the operations nobody publishes any more.","Cover art for the Qdrant review: a query crossing a filterable HNSW graph into a collection of points","https:\u002F\u002Fqdrant.tech","Apache-2.0 · Cloud from $25 per month","Qdrant review: the filter-first vector database · Balázs Csorba",[86,87,88,89,90],"The current release is 1.19.2, published 5 October 2026; the memory-tier parameter that drives every sizing decision arrived in 1.19.","Filtering is the reason to pick it: a payload index feeds the HNSW traversal, but payload indexes must exist before ingestion or the graph has to be rebuilt.","TurboQuant compresses up to 32 times, is asymmetric by default and rescores the top-k against full-precision vectors, which is a latency cost on every query.","A self-hosted instance is open to every network interface with no authentication until an API key, TLS and a network bind are configured by hand.","The vendor benchmark page still shows runs from January and June 2024, so its comparative numbers should not carry a purchase decision in 2026.",[92,95,98,101],{"q":93,"a":94},"Is Qdrant free to use in production?","Yes. The server is Apache-2.0 licensed, and self-hosting carries no licence fee, no feature gate and no call-home. The trade is operations: authentication, TLS, backups, shard rebalancing and upgrades all become your problem. Qdrant Cloud is the managed alternative, with a free tier limited to one node, 1GB RAM and 4GB disk.",{"q":96,"a":97},"Qdrant or pgvector?","Below a few million vectors, on a database you already run, pgvector wins on everything except throughput: it is one extension rather than a service. Qdrant earns its keep when a metadata filter sits on every query, because it can feed that filter into the HNSW traversal instead of retrieving candidates and discarding them.",{"q":99,"a":100},"How much does Qdrant Cloud cost?","The vendor publishes no entry price. Its pricing page describes hourly usage billing for vCPU, memory, storage, backup storage and inference tokens and links to a calculator; the free tier is 1GB RAM and 4GB disk, the Standard tier carries a 99.5% uptime SLA and the Premium tier adds SSO, private VPC links and 99.9%. The widely quoted figure of about $25 a month for the smallest paid cluster is a third-party estimate, not a published rate.",{"q":102,"a":103},"What happens when I turn quantisation on?","The compressed vectors are stored next to the originals, so nothing is destroyed and quantisation can be switched off. Rescoring of the top-k against full-precision vectors is on by default for binary quantisation and for TurboQuant at 1, 1.5 and 2 bits, and off for scalar and product quantisation. The production checklist tells you to re-benchmark retrieval quality afterwards, because some embedding models quantise badly.",[105,108,111,114,117,120,123,126],{"id":106,"title":107},"what-it-is","What it is",{"id":109,"title":110},"how-it-works","How a query actually runs",{"id":112,"title":113},"getting-started","Getting started",{"id":115,"title":116},"memory-and-quantisation","Memory and quantisation",{"id":118,"title":119},"self-hosting-and-security","Self-hosting and security",{"id":121,"title":122},"where-it-weakens","Where it weakens",{"id":124,"title":125},"verdict","Verdict",{"id":127,"title":128},"sources","Sources",[130,134,137,140,143,178,179,182,191,194,195,198,200,207,229,230,233,290,306,316,317,320,332,337,338,341,353,420,423,435,436,439,452,456,457],{"type":131,"content":132},"paragraph",[133],"Qdrant is a vector database written in Rust, released under Apache-2.0, and built around one decision most competitors treat as an afterthought: a metadata filter is not something you apply after retrieval, it is something the index is walked with. The result is the most convincing filtered-search story in the open-source world, and a genuinely simple deployment story — one container, one REST and gRPC API, official clients in six languages. The cost of that focus is that dense search has exactly one index implementation, HNSW, and that the managed tier is priced on metering rather than on a published rate. Recommended for filter-heavy retrieval, less so as a general-purpose store.",{"type":131,"content":135},[136],"It sits in the same layer as pgvector, Milvus, Weaviate and Pinecone, but the competing proposition is different. Weaviate sells an integrated AI-native store with built-in hybrid search and modules; Milvus sells horizontal scale to billions; Pinecone sells zero operations. Qdrant sells memory efficiency and filter throughput on a single node, which is a narrower claim and an easier one to hold to.",{"type":138,"level":139,"id":106,"text":107},"heading",2,{"type":131,"content":141},[142],"The data model is small on purpose. A point is a record with a vector and an optional JSON payload. A collection is a named set of points that share a dimensionality and a distance metric. Named vectors let one point hold several vectors with their own size and metric, which is how hybrid search is expressed. Everything else in the product is machinery for finding the nearest point faster without scanning.",{"type":144,"ordered":145,"items":146},"list",false,[147,149,156,158,168,170,176],[148],"Licence Apache-2.0, written in Rust, one repository with roughly 35,000 stars and 7,100 commits.",[150,151,155],"Current release 1.19.2, published 5 October 2026; release 1.19 added the ",{"tag":152,"children":153},"code",[154],"memory"," tiers that control where vectors, indexes and payloads live.",[157],"Distance metrics are dot product, cosine, Euclidean and Manhattan; cosine is implemented as a dot product over normalised vectors, normalised on upload.",[159,160,163,164,167],"Dense search uses HNSW only, with ",{"tag":152,"children":161},[162],"m"," defaulting to 16 and ",{"tag":152,"children":165},[166],"ef_construct"," to 100, both overridable per collection and per named vector.",[169],"Payload index types are keyword, integer, float, bool, geo, datetime, text and uuid, each created before ingestion for the filterable HNSW to use it.",[171,172,175],"Hybrid and multi-stage search arrived in 1.10 through the Query API: ",{"tag":152,"children":173},[174],"prefetch"," sub-requests fused with RRF or DBSF, and prefetches can nest.",[177],"Official clients exist for Python, TypeScript, Rust, Go, Java and .NET over REST on 6333 and gRPC on 6334.",{"type":138,"level":139,"id":109,"text":110},{"type":131,"content":180},[181],"The interesting engineering is in the query planner, and it splits into three cases. A filter so strict that it matches very little data is better served by a full scan than by a graph walk. A filter so weak that it matches most of the collection can use the HNSW graph as it is. Everything in between — which is where tenant-scoped and language-scoped retrieval lives — is the case a plain vector index plus post-filtering handles badly, and the case the filterable HNSW was built for.",{"type":183,"attrs":184,"inner":188,"caption":189},"diagram",{"viewBox":185,"role":186,"aria-labelledby":187},"0 0 720 372","img","d-qd-t d-qd-d","\u003Ctitle id=\"d-qd-t\">Three paths a filtered query can take through Qdrant\u003C\u002Ftitle>\u003Cdesc id=\"d-qd-d\">A query carrying a vector and a filter enters at the top and splits three ways. Left: the filter is strict and few points match, so the payload index is used for a full scan and the HNSW graph is skipped; the default full scan threshold is 10,000 kilobytes, where one kilobyte is one vector of size 256. Middle: the filter is weak and most points match, so the HNSW graph is used as it is and filtering happens afterwards. Right: the filter is in the middle, the case that matters in production, so the payload index feeds the HNSW walk, and the payload indexes have to exist before the data is ingested or the graph has to be rebuilt. A band below: with quantisation enabled the graph is walked over compressed vectors and the top candidates are rescored against the originals, with rescore and oversampling as per-query settings.\u003C\u002Fdesc>\u003Ctext x=\"20\" y=\"26\" class=\"d-title\">Three paths a filtered query can take\u003C\u002Ftext>\u003Ctext x=\"700\" y=\"26\" text-anchor=\"end\" class=\"d-label\">decided by the filter, not the vector\u003C\u002Ftext>\u003Crect x=\"270\" y=\"44\" width=\"180\" height=\"50\" rx=\"10\" class=\"d-accent\" \u002F>\u003Ctext x=\"360\" y=\"66\" text-anchor=\"middle\" class=\"d-text\">Query\u003C\u002Ftext>\u003Ctext x=\"360\" y=\"84\" text-anchor=\"middle\" class=\"d-small\">vector + filter\u003C\u002Ftext>\u003Cpath d=\"M126 100 V112 H360 V124\" class=\"d-line\" \u002F>\u003Cpath d=\"M360 100 V124\" class=\"d-line\" \u002F>\u003Cpath d=\"M594 100 V112 H360\" class=\"d-line\" \u002F>\u003Crect x=\"20\" y=\"124\" width=\"213\" height=\"132\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"126\" y=\"150\" text-anchor=\"middle\" class=\"d-text\">Strict filter\u003C\u002Ftext>\u003Ctext x=\"126\" y=\"176\" text-anchor=\"middle\" class=\"d-small\">payload index\u003C\u002Ftext>\u003Ctext x=\"126\" y=\"198\" text-anchor=\"middle\" class=\"d-small\">plus a full scan\u003C\u002Ftext>\u003Ctext x=\"126\" y=\"220\" text-anchor=\"middle\" class=\"d-small\">no graph walk\u003C\u002Ftext>\u003Ctext x=\"126\" y=\"242\" text-anchor=\"middle\" class=\"d-small\">default 10,000 KB\u003C\u002Ftext>\u003Crect x=\"253\" y=\"124\" width=\"214\" height=\"132\" rx=\"10\" class=\"d-sky\" \u002F>\u003Ctext x=\"360\" y=\"150\" text-anchor=\"middle\" class=\"d-text\">Weak filter\u003C\u002Ftext>\u003Ctext x=\"360\" y=\"176\" text-anchor=\"middle\" class=\"d-small\">HNSW as it is\u003C\u002Ftext>\u003Ctext x=\"360\" y=\"198\" text-anchor=\"middle\" class=\"d-small\">filter afterwards\u003C\u002Ftext>\u003Ctext x=\"360\" y=\"220\" text-anchor=\"middle\" class=\"d-small\">cheap, broad\u003C\u002Ftext>\u003Ctext x=\"360\" y=\"242\" text-anchor=\"middle\" class=\"d-small\">the easy case\u003C\u002Ftext>\u003Crect x=\"486\" y=\"124\" width=\"213\" height=\"132\" rx=\"10\" class=\"d-gold\" \u002F>\u003Ctext x=\"592\" y=\"150\" text-anchor=\"middle\" class=\"d-text\">Middle ground\u003C\u002Ftext>\u003Ctext x=\"592\" y=\"176\" text-anchor=\"middle\" class=\"d-small\">filterable HNSW\u003C\u002Ftext>\u003Ctext x=\"592\" y=\"198\" text-anchor=\"middle\" class=\"d-small\">index feeds the walk\u003C\u002Ftext>\u003Ctext x=\"592\" y=\"220\" text-anchor=\"middle\" class=\"d-small\">tenants, languages\u003C\u002Ftext>\u003Ctext x=\"592\" y=\"242\" text-anchor=\"middle\" class=\"d-small\">index before ingest\u003C\u002Ftext>\u003Crect x=\"20\" y=\"280\" width=\"679\" height=\"72\" rx=\"10\" class=\"d-box d-dash\" \u002F>\u003Ctext x=\"40\" y=\"308\" class=\"d-text\">With quantisation on\u003C\u002Ftext>\u003Ctext x=\"40\" y=\"332\" class=\"d-small\">the walk runs over compressed vectors and the top candidates are rescored against the originals;\u003C\u002Ftext>\u003Ctext x=\"40\" y=\"350\" class=\"d-small\">rescore and oversampling are per-query settings, so the recall and latency trade is tunable at query time\u003C\u002Ftext>",[190],"The filter decides the execution path, and the default thresholds are configuration you inherit rather than choices you make.",{"type":131,"content":192},[193],"One operational detail in that middle column causes more downtime than any query tuning: payload indexes only help the HNSW graph if they existed before the data arrived. Add a tenant field to an existing collection and the graph has to be rebuilt to become filter-aware, which on a large collection is measured in hours. Design the payload schema before the first upsert, not after the first incident.",{"type":138,"level":139,"id":112,"text":113},{"type":131,"content":196},[197],"One container on 6333 with no authentication is enough to start, and that is also the single most important thing to fix before it leaves a laptop. The configuration below creates a collection with TurboQuant at one bit, indexes the tenant field before ingesting anything, and runs a filtered query through the Query API:",{"type":152,"code":199},"from qdrant_client import QdrantClient, models\n\nclient = QdrantClient(url=\"http:\u002F\u002Flocalhost:6333\")\n\nclient.create_collection(\n    collection_name=\"chunks\",\n    vectors_config=models.VectorParams(size=1024, distance=models.Distance.COSINE),\n    quantization_config=models.TurboQuantization(\n        turbo=models.TurboQuantQuantizationConfig(bits=models.TurboQuantBitSize.BITS1),\n    ),\n)\n\n# Before ingestion: the HNSW graph can only be filter-aware for\n# fields that already have a payload index.\nclient.create_payload_index(\n    collection_name=\"chunks\",\n    field_name=\"tenant\",\n    field_schema=models.PayloadSchemaType.KEYWORD,\n)\n\nclient.upsert(\n    collection_name=\"chunks\",\n    points=[models.PointStruct(id=1, vector=[0.1] * 1024, payload={\"tenant\": \"acme\"})],\n)\n\nhits = client.query_points(\n    collection_name=\"chunks\",\n    query=[0.1] * 1024,\n    query_filter=models.Filter(\n        must=[models.FieldCondition(key=\"tenant\", match=models.MatchValue(value=\"acme\"))],\n    ),\n    limit=10,\n    search_params=models.SearchParams(quantization=models.QuantizationSearchParams(oversampling=2.0)),\n).points\n",{"type":131,"content":201},[202,203,206],"Two things in that snippet are worth expanding. Quantisation is a collection setting applied at indexation time, and the compressed vectors live beside the originals, so the originals are still there for a rescore. The ",{"tag":152,"children":204},[205],"oversampling"," parameter asks the quantised index for twice the candidates before rescoring, which is the knob to turn when compressed search starts dropping results you know should be there.",{"type":208,"variant":209,"title":210,"body":211},"callout","note","Read the capacity page before the first node",[212],[213,214,217,218,221,222,224,225,228],"The vendor's own sizing rules are short enough to be worth memorising. Vectors and the HNSW indexes are ",{"tag":152,"children":215},[216],"cached"," by default, meaning pre-loaded but evictable; set them to ",{"tag":152,"children":219},[220],"cold"," to live on disk. Payloads default to ",{"tag":152,"children":223},[220],", payload indexes and quantised vectors to ",{"tag":152,"children":226},[227],"pinned",". Budget roughly twice the size of the indexed fields for payload indexes, and add 1.5 for payload storage overhead.",{"type":138,"level":139,"id":115,"text":116},{"type":131,"content":231},[232],"Quantisation is the single highest-leverage change before production, and Qdrant now offers four methods with genuinely different trade-offs rather than one binary switch. The production checklist calls it one of the three things worth doing first, alongside sizing RAM honestly and indexing the fields you filter on.",{"type":234,"head":235,"rows":246},"table",[236,238,240,242,244],[237],"Method",[239],"Compression",[241],"Rescores by default",[243],"Where it fits",[245],"Cost",[247,258,269,280],[248,250,252,254,256],[249],"TurboQuant",[251],"Up to 32x, 4 bits down to 1 bit",[253],"Yes, at 1, 1.5 and 2 bits",[255],"The default choice since 1.18",[257],"Asymmetric, so queries stay full precision",[259,261,263,265,267],[260],"Scalar",[262],"4x, float32 to int8",[264],"No",[266],"Lowest-risk compression",[268],"Needs a quantile to bound outliers",[270,272,274,276,278],[271],"Binary",[273],"Up to 32x, 1 to 2 bits per component",[275],"Yes",[277],"High-dimensional centred embeddings",[279],"Fails on distributions it was not built for",[281,283,285,286,288],[282],"Product",[284],"Up to 64x, 256 centroids per chunk",[264],[287],"Memory is the only goal",[289],"Largest accuracy loss of the four",{"type":131,"content":291},[292,293,295,296,298,299,302,303,305],"Beyond the method itself, 1.19 added a memory-tier parameter for the quantised copy, which is what makes quantisation usable as a RAM reduction rather than only a speed trick. Put the originals in the ",{"tag":152,"children":294},[220]," tier and the quantised vectors in ",{"tag":152,"children":297},[227],", and search touches disk only while rescoring the top candidates. The ",{"tag":152,"children":300},[301],"turbo4"," datatype, which stores each dimension on disk as four bits instead of a float, shrinks the disk side further at the price of recall. Inline storage in the HNSW index, available since 1.16, cuts I\u002FO further but only pays off with vectors and index in the ",{"tag":152,"children":304},[220]," tier and quantisation enabled; keep it to at most four bits per dimension or the index balloons.",{"type":208,"variant":307,"title":308,"body":309},"warn","Quantisation is a recall decision, not a storage decision",[310],[311,312,315],"The documentation is explicit that some embedding models cannot be quantised efficiently, and the production checklist asks you to re-benchmark retrieval quality after enabling it rather than assuming. Treat recall as an acceptance criterion with a number attached, and keep a way to turn quantisation off at query time: the ",{"tag":152,"children":313},[314],"ignore"," search parameter does exactly that without a collection rebuild.",{"type":138,"level":139,"id":118,"text":119},{"type":131,"content":318},[319],"The security documentation opens by saying that self-hosted open-source deployments are not secure by default and are not production-ready, and that a default instance is open to all network interfaces with no authentication configured. That is a more honest security page than most database vendors publish, and it comes with a specific checklist.",{"type":144,"ordered":145,"items":321},[322,324,326,328,330],[323],"Three API key types: admin, read-only for query-only services, and granular keys scoped per collection with read or write rights.",[325],"TLS for traffic in both directions, plus a network bind to a private interface; bind to 127.0.0.1 while developing locally.",[327],"Audit logging of API operations to a file, for forensics and compliance evidence rather than for insight.",[329],"Everything works the same on Qdrant Cloud, where these controls are on by default — which is the real argument for the managed tier.",[331],"Community, Standard and Premium support tiers differ in response time (four hours for a full outage on the free tier, one hour on Standard) rather than in features.",{"type":208,"variant":209,"title":333,"body":334},"Migration is a stated non-issue, mostly",[335],[336],"The company says a migration tool and documentation exist for moving from a self-hosted deployment to Qdrant Cloud, and that migrating the other way is just a matter of running the container. That is true of the vectors. It is not true of everything you built around the engine: the Query API's fusion strategies and your HNSW tuning are the parts worth writing down before you move.",{"type":138,"level":139,"id":121,"text":122},{"type":131,"content":339},[340],"Five weaknesses are worth stating plainly, because they are the ones that decide against it.",{"type":144,"ordered":145,"items":342},[343,345,347,349,351],[344],"One dense index. The documentation says Qdrant only uses HNSW for dense vectors. There is no IVF, no disk-based graph and no GPU index, so a corpus that will not fit a machine's RAM needs an architecture change rather than a setting.",[346],"Payload indexes are built before ingestion or not at all. Correctness of the optimisation depends on a migration you forget to schedule.",[348],"Horizontal scaling is real but not free. Sharding, replication factors and shard-key-aware reads add configuration that has to be right, and the capacity page asks you to decide all of it before provisioning.",[350],"The managed tier publishes no rate. Billing is hourly on vCPU, memory, storage, backups and inference tokens, with a calculator instead of a price list, and serverless is still listed as coming soon.",[352],"The public benchmark is stale. The vendor's comparison page is labelled January and June 2024, so its conclusion that Qdrant leads on throughput and latency describes software from two years ago.",{"type":234,"head":354,"rows":365},[355,357,359,361,363],[356],"Engine",[358],"Licence",[360],"Dense index options",[362],"Where filtering happens",[364],"Operational shape",[366,377,388,398,409],[367,369,371,373,375],[368],"Qdrant",[370],"Apache-2.0",[372],"HNSW only, filterable",[374],"Payload index feeds the graph walk",[376],"One container, or managed cloud",[378,380,382,384,386],[379],"pgvector",[381],"PostgreSQL Licence",[383],"HNSW, IVFFlat",[385],"SQL WHERE on the same table",[387],"An extension in a database you already run",[389,391,392,394,396],[390],"Milvus",[370],[393],"HNSW, IVF, DiskANN, SCANN, GPU",[395],"Scalar index inside the engine",[397],"Distributed services, heavy to operate",[399,401,403,405,407],[400],"Weaviate",[402],"BSD-3-Clause",[404],"HNSW, flat, dynamic",[406],"Inverted index, native BM25 hybrid",[408],"Single binary, optional cluster",[410,412,414,416,418],[411],"Pinecone",[413],"Proprietary, managed only",[415],"Proprietary serverless index",[417],"Server-side filtering on the service",[419],"Nothing to run, per-query billing",{"type":131,"content":421},[422],"The one-line version: pgvector if you already run Postgres and the corpus fits; Weaviate if native hybrid search is the requirement; Milvus when the numbers genuinely reach hundreds of millions; Pinecone when nobody will operate anything. Qdrant is the pick in between, where a metadata filter on every query matters more than the ceiling does.",{"type":208,"variant":307,"title":424,"body":425},"Measure your own recall",[426],[427,428,431,432,434],"Every comparison in this space is a vendor running its own harness, which is why the useful number is the one you produce: the fraction of true nearest neighbours returned at your ",{"tag":152,"children":429},[430],"ef"," and your ",{"tag":152,"children":433},[162],", with and without quantisation. The search API accepts an exact mode for precisely this comparison, and running it once against your own embeddings costs an afternoon and settles the argument.",{"type":138,"level":139,"id":124,"text":125},{"type":131,"content":437},[438],"Qdrant is the best open-source answer to filtered vector search, and its weaknesses are all in areas where it is not trying to win. If your queries carry a tenant, a language or a date — and in production they nearly always do — the filterable HNSW is a real architectural advantage rather than a feature checkbox. Choose it knowing that dense search has one index, that the payload schema has to be right on day one, and that you will be running the security checklist yourself.",{"type":144,"ordered":440,"items":441},true,[442,444,446,448,450],[443],"Choose it when a metadata filter is on essentially every query and the corpus fits on one machine with headroom.",[445],"Choose it when you want no licence conversation: Apache-2.0, no feature gate, no call-home, no usage reporting.",[447],"Choose it when your team can own an API key, a TLS certificate, a backup schedule and a version upgrade.",[449],"Do not choose it when the corpus will exceed RAM and you have no appetite for a migration to a distributed engine.",[451],"Do not choose it on a benchmark. Run the exact-mode recall check on your own embeddings, then decide.",{"type":453,"content":454},"quote",[455],"Self-hosted open-source deployments are not secure by default and are not production-ready. By default, all self-deployed Qdrant instances are open to all network interfaces and have no authentication configured.",{"type":138,"level":139,"id":127,"text":128},{"type":144,"ordered":440,"items":458},[459,463,466,469,473,477,480,483,486,490,493,497,500],[460],{"tag":461,"href":35,"children":462},"a",[34],[464],{"tag":461,"href":38,"children":465},[37],[467],{"tag":461,"href":41,"children":468},[40],[470],{"tag":461,"href":44,"children":471},[472],"Qdrant: indexing, payload indexes and the filterable HNSW",[474],{"tag":461,"href":47,"children":475},[476],"Qdrant: quantisation methods and memory tiers",[478],{"tag":461,"href":50,"children":479},[49],[481],{"tag":461,"href":53,"children":482},[52],[484],{"tag":461,"href":56,"children":485},[55],[487],{"tag":461,"href":59,"children":488},[489],"Qdrant: filtering clauses",[491],{"tag":461,"href":62,"children":492},[61],[494],{"tag":461,"href":65,"children":495},[496],"Qdrant pricing: free, standard and premium tiers",[498],{"tag":461,"href":68,"children":499},[67],[501],{"tag":461,"href":71,"children":502},[70],[504,554,605,644],{"slug":505,"published":506,"minutes":507,"category":7,"tags":508,"keywords":514,"about":523,"sources":527,"cover":546,"og":547,"expertise":74,"locales":548,"lang":76,"title":549,"description":550,"coverAlt":551,"url":552,"pricing":553,"kind":509},"zep","2026-10-06",11,[509,510,511,512,513],"Agent memory","Knowledge graph","Temporal graph","Context engineering","RAG",[515,516,517,518,519,520,521,522],"zep ai","zep agent memory","graphiti knowledge graph","zep pricing","zep vs mem0","long-term memory for agents","temporal knowledge graph","zep cloud",[524],{"name":525,"url":526},"Zep","https:\u002F\u002Fwww.getzep.com\u002F",[528,531,534,537,540,543],{"title":529,"url":530},"Zep pricing: plans, credits and limits","https:\u002F\u002Fwww.getzep.com\u002Fpricing",{"title":532,"url":533},"Zep documentation","https:\u002F\u002Fhelp.getzep.com\u002F",{"title":535,"url":536},"Graphiti on GitHub","https:\u002F\u002Fgithub.com\u002Fgetzep\u002Fgraphiti",{"title":538,"url":539},"Graphiti product page","https:\u002F\u002Fwww.getzep.com\u002Fplatform\u002Fgraphiti\u002F",{"title":541,"url":542},"Announcing a new direction for Zep's open-source strategy","https:\u002F\u002Fwww.getzep.com\u002Fblog\u002Fannouncing-a-new-direction-for-zeps-open-source-strategy\u002F",{"title":544,"url":545},"Graphiti: temporal knowledge graphs for AI agents (arXiv)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2501.13956","\u002Fimages\u002Fblog\u002Fzep\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fzep\u002Fog.jpg",[76,77,78],"Zep review: agent memory on a temporal graph","Zep is a hosted agent-memory API on a temporal knowledge graph: credits on writes, retrieval free, Flex from $125 a month, Graphiti as the part you can self-host.","Diagram of how a fact reaches the prompt in Zep: messages and facts are extracted into a per-user context graph of entities and relationships, and retrieval walks the graph to return a context block with the supporting facts.","https:\u002F\u002Fwww.getzep.com","Apache-2.0 core · Cloud from $50 per month",{"slug":555,"published":556,"minutes":557,"category":7,"tags":558,"keywords":560,"about":567,"sources":572,"cover":597,"og":598,"expertise":74,"locales":599,"lang":76,"title":600,"description":601,"coverAlt":602,"url":603,"pricing":604,"kind":24},"lancedb","2026-09-21",9,[9,12,559,513],"Embedded database",[555,561,562,563,564,20,565,566],"lancedb review","lance vector database","embedded vector database","lancedb vs qdrant","lancedb indexing","lance data format",[568,569],{"name":24,"url":25},{"name":570,"url":571},"Apache Arrow","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FApache_Arrow",[573,576,579,582,585,588,591,594],{"title":574,"url":575},"LanceDB quickstart","https:\u002F\u002Fdocs.lancedb.com\u002Fquickstart",{"title":577,"url":578},"LanceDB vector indexes","https:\u002F\u002Fdocs.lancedb.com\u002Findexing\u002Fvector-index",{"title":580,"url":581},"LanceDB indexing guide","https:\u002F\u002Fdocs.lancedb.com\u002Findexing\u002Findex",{"title":583,"url":584},"LanceDB hybrid search","https:\u002F\u002Fdocs.lancedb.com\u002Fsearch\u002Fhybrid-search",{"title":586,"url":587},"LanceDB Enterprise","https:\u002F\u002Fdocs.lancedb.com\u002Fenterprise",{"title":589,"url":590},"LanceDB frequently asked questions","https:\u002F\u002Fdocs.lancedb.com\u002Ffaq\u002Ffaq-oss",{"title":592,"url":593},"LanceDB pricing","https:\u002F\u002Flancedb.com\u002Fpricing",{"title":595,"url":596},"LanceDB on PyPI","https:\u002F\u002Fpypi.org\u002Fproject\u002Flancedb\u002F","\u002Fimages\u002Fblog\u002Flancedb\u002Fcover.webp","\u002Fimages\u002Fblog\u002Flancedb\u002Fog.jpg",[76,77,78],"LanceDB: vector search that starts as a library","A review of LanceDB: an Apache-2.0 embedded vector library, its IVF and HNSW index choices, hybrid search with rank fusion, and what the Enterprise tier adds.","Cover art for the LanceDB review: one Lance table feeding a vector index and a full-text index into a fused ranking","https:\u002F\u002Flancedb.com","Apache-2.0 · Cloud paid",{"slug":379,"published":556,"minutes":6,"category":7,"tags":606,"keywords":608,"about":615,"sources":621,"cover":636,"og":637,"expertise":74,"locales":638,"lang":76,"title":639,"description":640,"coverAlt":641,"url":617,"pricing":642,"kind":643},[9,607,10,513,11],"Postgres",[379,609,610,611,612,613,614],"pgvector vs qdrant","postgres vector search","hnsw index postgres","iterative index scans","binary quantization postgres","vector database postgres",[616,618],{"name":379,"url":617},"https:\u002F\u002Fgithub.com\u002Fpgvector\u002Fpgvector",{"name":619,"url":620},"PostgreSQL","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FPostgreSQL",[622,624,627,630,633],{"title":623,"url":617},"pgvector README",{"title":625,"url":626},"pgvector changelog","https:\u002F\u002Fgithub.com\u002Fpgvector\u002Fpgvector\u002Fblob\u002Fmaster\u002FCHANGELOG.md",{"title":628,"url":629},"PostgreSQL news: pgvector 0.8.2 released","https:\u002F\u002Fwww.postgresql.org\u002Fabout\u002Fnews\u002Fpgvector-082-released-3245\u002F",{"title":631,"url":632},"AWS: Scale pgvector with binary quantization","https:\u002F\u002Faws.amazon.com\u002Fblogs\u002Fdatabase\u002Fscale-pgvector-with-binary-quantization-on-amazon-aurora-postgresql\u002F",{"title":634,"url":635},"pgvector licence","https:\u002F\u002Fgithub.com\u002Fpgvector\u002Fpgvector\u002Fblob\u002Fmaster\u002FLICENSE","\u002Fimages\u002Fblog\u002Fpgvector\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fpgvector\u002Fog.jpg",[76,77,78],"pgvector, reviewed: the vector database you do not have to run","A review of pgvector 0.8.7: iterative scans for filtered search, HNSW and IVFFlat, binary quantisation at 100M vectors, and the CVE that made index builds a patch item.","A query enters at the top and splits into an exact sequential scan, an HNSW graph walk and an IVFFlat probe; a band below shows an iterative scan continuing until the limit is full.","PostgreSQL licence","Vector database extension",{"slug":645,"published":646,"minutes":6,"category":7,"tags":647,"keywords":649,"about":655,"sources":659,"cover":687,"og":688,"expertise":74,"locales":689,"lang":76,"title":690,"description":691,"coverAlt":692,"url":658,"pricing":693,"kind":509},"mem0","2026-09-17",[509,648,513,9],"Long-term memory",[645,650,651,652,653,520,654],"mem0 review","agent memory layer","mem0 self-hosted","mem0 pricing","mem0 alternatives",[656],{"name":657,"url":658},"Mem0","https:\u002F\u002Fmem0.ai",[660,663,666,669,672,675,678,681,684],{"title":661,"url":662},"Mem0 documentation","https:\u002F\u002Fdocs.mem0.ai\u002Fintroduction",{"title":664,"url":665},"Mem0 quickstart","https:\u002F\u002Fdocs.mem0.ai\u002Fquickstart",{"title":667,"url":668},"How Mem0 works","https:\u002F\u002Fdocs.mem0.ai\u002Fcore-concepts\u002Fhow-it-works",{"title":670,"url":671},"Mem0 pricing","https:\u002F\u002Fmem0.ai\u002Fpricing",{"title":673,"url":674},"Mem0 on GitHub","https:\u002F\u002Fgithub.com\u002Fmem0ai\u002Fmem0",{"title":676,"url":677},"mem0ai on PyPI","https:\u002F\u002Fpypi.org\u002Fproject\u002Fmem0ai\u002F",{"title":679,"url":680},"Mem0 research and benchmarks","https:\u002F\u002Fmem0.ai\u002Fresearch",{"title":682,"url":683},"Mem0 MCP server","https:\u002F\u002Fdocs.mem0.ai\u002Fplatform\u002Fmem0-mcp",{"title":685,"url":686},"Mem0 paper on arXiv","https:\u002F\u002Farxiv.org\u002Fabs\u002F2504.19413","\u002Fimages\u002Fblog\u002Fmem0\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fmem0\u002Fog.jpg",[76,77,78],"Mem0: what an agent memory layer costs per turn","A review of Mem0: facts extracted from every turn, the April 2026 benchmark table and its platform-only caveat, four cloud tiers and what self-hosting leaves out.","A loop that turns conversation into stored facts and reads them back into the prompt","Free tier · from $19 per month",1791383549036]