Tools/RAG & retrieval
Qdrant: a vector database built around filtering
A review of Qdrant: filterable HNSW, four quantisation methods, the memory tiers in v1.19 and the operations nobody publishes any more.
- Type
- Vector database
- Pricing
- Apache-2.0 · Cloud from $25 per month
Balázs Csorba··10 min read
- Vector search
- HNSW
- Quantisation
- Hybrid search
- Filtering

Key takeaways
- The current release is 1.19.2, published 5 October 2026; the memory-tier parameter that drives every sizing decision arrived in 1.19.
- Filtering is the reason to pick it: a payload index feeds the HNSW traversal, but payload indexes must exist before ingestion or the graph has to be rebuilt.
- TurboQuant compresses up to 32 times, is asymmetric by default and rescores the top-k against full-precision vectors, which is a latency cost on every query.
- A self-hosted instance is open to every network interface with no authentication until an API key, TLS and a network bind are configured by hand.
- The vendor benchmark page still shows runs from January and June 2024, so its comparative numbers should not carry a purchase decision in 2026.
Qdrant is a vector database written in Rust, released under Apache-2.0, and built around one decision most competitors treat as an afterthought: a metadata filter is not something you apply after retrieval, it is something the index is walked with. The result is the most convincing filtered-search story in the open-source world, and a genuinely simple deployment story — one container, one REST and gRPC API, official clients in six languages. The cost of that focus is that dense search has exactly one index implementation, HNSW, and that the managed tier is priced on metering rather than on a published rate. Recommended for filter-heavy retrieval, less so as a general-purpose store.
It sits in the same layer as pgvector, Milvus, Weaviate and Pinecone, but the competing proposition is different. Weaviate sells an integrated AI-native store with built-in hybrid search and modules; Milvus sells horizontal scale to billions; Pinecone sells zero operations. Qdrant sells memory efficiency and filter throughput on a single node, which is a narrower claim and an easier one to hold to.
What it is
The data model is small on purpose. A point is a record with a vector and an optional JSON payload. A collection is a named set of points that share a dimensionality and a distance metric. Named vectors let one point hold several vectors with their own size and metric, which is how hybrid search is expressed. Everything else in the product is machinery for finding the nearest point faster without scanning.
- Licence Apache-2.0, written in Rust, one repository with roughly 35,000 stars and 7,100 commits.
- Current release 1.19.2, published 5 October 2026; release 1.19 added the
memorytiers that control where vectors, indexes and payloads live. - Distance metrics are dot product, cosine, Euclidean and Manhattan; cosine is implemented as a dot product over normalised vectors, normalised on upload.
- Dense search uses HNSW only, with
mdefaulting to 16 andef_constructto 100, both overridable per collection and per named vector. - Payload index types are keyword, integer, float, bool, geo, datetime, text and uuid, each created before ingestion for the filterable HNSW to use it.
- Hybrid and multi-stage search arrived in 1.10 through the Query API:
prefetchsub-requests fused with RRF or DBSF, and prefetches can nest. - Official clients exist for Python, TypeScript, Rust, Go, Java and .NET over REST on 6333 and gRPC on 6334.
How a query actually runs
The interesting engineering is in the query planner, and it splits into three cases. A filter so strict that it matches very little data is better served by a full scan than by a graph walk. A filter so weak that it matches most of the collection can use the HNSW graph as it is. Everything in between — which is where tenant-scoped and language-scoped retrieval lives — is the case a plain vector index plus post-filtering handles badly, and the case the filterable HNSW was built for.
One operational detail in that middle column causes more downtime than any query tuning: payload indexes only help the HNSW graph if they existed before the data arrived. Add a tenant field to an existing collection and the graph has to be rebuilt to become filter-aware, which on a large collection is measured in hours. Design the payload schema before the first upsert, not after the first incident.
Getting started
One container on 6333 with no authentication is enough to start, and that is also the single most important thing to fix before it leaves a laptop. The configuration below creates a collection with TurboQuant at one bit, indexes the tenant field before ingesting anything, and runs a filtered query through the Query API:
from qdrant_client import QdrantClient, models
client = QdrantClient(url="http://localhost:6333")
client.create_collection(
collection_name="chunks",
vectors_config=models.VectorParams(size=1024, distance=models.Distance.COSINE),
quantization_config=models.TurboQuantization(
turbo=models.TurboQuantQuantizationConfig(bits=models.TurboQuantBitSize.BITS1),
),
)
# Before ingestion: the HNSW graph can only be filter-aware for
# fields that already have a payload index.
client.create_payload_index(
collection_name="chunks",
field_name="tenant",
field_schema=models.PayloadSchemaType.KEYWORD,
)
client.upsert(
collection_name="chunks",
points=[models.PointStruct(id=1, vector=[0.1] * 1024, payload={"tenant": "acme"})],
)
hits = client.query_points(
collection_name="chunks",
query=[0.1] * 1024,
query_filter=models.Filter(
must=[models.FieldCondition(key="tenant", match=models.MatchValue(value="acme"))],
),
limit=10,
search_params=models.SearchParams(quantization=models.QuantizationSearchParams(oversampling=2.0)),
).points
Two things in that snippet are worth expanding. Quantisation is a collection setting applied at indexation time, and the compressed vectors live beside the originals, so the originals are still there for a rescore. The oversampling parameter asks the quantised index for twice the candidates before rescoring, which is the knob to turn when compressed search starts dropping results you know should be there.
Memory and quantisation
Quantisation is the single highest-leverage change before production, and Qdrant now offers four methods with genuinely different trade-offs rather than one binary switch. The production checklist calls it one of the three things worth doing first, alongside sizing RAM honestly and indexing the fields you filter on.
| Method | Compression | Rescores by default | Where it fits | Cost |
|---|---|---|---|---|
| TurboQuant | Up to 32x, 4 bits down to 1 bit | Yes, at 1, 1.5 and 2 bits | The default choice since 1.18 | Asymmetric, so queries stay full precision |
| Scalar | 4x, float32 to int8 | No | Lowest-risk compression | Needs a quantile to bound outliers |
| Binary | Up to 32x, 1 to 2 bits per component | Yes | High-dimensional centred embeddings | Fails on distributions it was not built for |
| Product | Up to 64x, 256 centroids per chunk | No | Memory is the only goal | Largest accuracy loss of the four |
Beyond the method itself, 1.19 added a memory-tier parameter for the quantised copy, which is what makes quantisation usable as a RAM reduction rather than only a speed trick. Put the originals in the cold tier and the quantised vectors in pinned, and search touches disk only while rescoring the top candidates. The turbo4 datatype, which stores each dimension on disk as four bits instead of a float, shrinks the disk side further at the price of recall. Inline storage in the HNSW index, available since 1.16, cuts I/O further but only pays off with vectors and index in the cold tier and quantisation enabled; keep it to at most four bits per dimension or the index balloons.
Self-hosting and security
The security documentation opens by saying that self-hosted open-source deployments are not secure by default and are not production-ready, and that a default instance is open to all network interfaces with no authentication configured. That is a more honest security page than most database vendors publish, and it comes with a specific checklist.
- Three API key types: admin, read-only for query-only services, and granular keys scoped per collection with read or write rights.
- TLS for traffic in both directions, plus a network bind to a private interface; bind to 127.0.0.1 while developing locally.
- Audit logging of API operations to a file, for forensics and compliance evidence rather than for insight.
- Everything works the same on Qdrant Cloud, where these controls are on by default — which is the real argument for the managed tier.
- Community, Standard and Premium support tiers differ in response time (four hours for a full outage on the free tier, one hour on Standard) rather than in features.
Where it weakens
Five weaknesses are worth stating plainly, because they are the ones that decide against it.
- One dense index. The documentation says Qdrant only uses HNSW for dense vectors. There is no IVF, no disk-based graph and no GPU index, so a corpus that will not fit a machine's RAM needs an architecture change rather than a setting.
- Payload indexes are built before ingestion or not at all. Correctness of the optimisation depends on a migration you forget to schedule.
- Horizontal scaling is real but not free. Sharding, replication factors and shard-key-aware reads add configuration that has to be right, and the capacity page asks you to decide all of it before provisioning.
- The managed tier publishes no rate. Billing is hourly on vCPU, memory, storage, backups and inference tokens, with a calculator instead of a price list, and serverless is still listed as coming soon.
- The public benchmark is stale. The vendor's comparison page is labelled January and June 2024, so its conclusion that Qdrant leads on throughput and latency describes software from two years ago.
| Engine | Licence | Dense index options | Where filtering happens | Operational shape |
|---|---|---|---|---|
| Qdrant | Apache-2.0 | HNSW only, filterable | Payload index feeds the graph walk | One container, or managed cloud |
| pgvector | PostgreSQL Licence | HNSW, IVFFlat | SQL WHERE on the same table | An extension in a database you already run |
| Milvus | Apache-2.0 | HNSW, IVF, DiskANN, SCANN, GPU | Scalar index inside the engine | Distributed services, heavy to operate |
| Weaviate | BSD-3-Clause | HNSW, flat, dynamic | Inverted index, native BM25 hybrid | Single binary, optional cluster |
| Pinecone | Proprietary, managed only | Proprietary serverless index | Server-side filtering on the service | Nothing to run, per-query billing |
The one-line version: pgvector if you already run Postgres and the corpus fits; Weaviate if native hybrid search is the requirement; Milvus when the numbers genuinely reach hundreds of millions; Pinecone when nobody will operate anything. Qdrant is the pick in between, where a metadata filter on every query matters more than the ceiling does.
Verdict
Qdrant is the best open-source answer to filtered vector search, and its weaknesses are all in areas where it is not trying to win. If your queries carry a tenant, a language or a date — and in production they nearly always do — the filterable HNSW is a real architectural advantage rather than a feature checkbox. Choose it knowing that dense search has one index, that the payload schema has to be right on day one, and that you will be running the security checklist yourself.
- Choose it when a metadata filter is on essentially every query and the corpus fits on one machine with headroom.
- Choose it when you want no licence conversation: Apache-2.0, no feature gate, no call-home, no usage reporting.
- Choose it when your team can own an API key, a TLS certificate, a backup schedule and a version upgrade.
- Do not choose it when the corpus will exceed RAM and you have no appetite for a migration to a distributed engine.
- Do not choose it on a benchmark. Run the exact-mode recall check on your own embeddings, then decide.
Self-hosted open-source deployments are not secure by default and are not production-ready. By default, all self-deployed Qdrant instances are open to all network interfaces and have no authentication configured.
Sources
- Qdrant documentation
- Qdrant: points
- Qdrant: collections
- Qdrant: indexing, payload indexes and the filterable HNSW
- Qdrant: quantisation methods and memory tiers
- Qdrant: capacity planning
- Qdrant: optimise performance
- Qdrant: hybrid and multi-stage queries
- Qdrant: filtering clauses
- Qdrant: security and access control
- Qdrant pricing: free, standard and premium tiers
- Qdrant on GitHub
- Qdrant vector search benchmarks
Frequently asked questions
Is Qdrant free to use in production?
Yes. The server is Apache-2.0 licensed, and self-hosting carries no licence fee, no feature gate and no call-home. The trade is operations: authentication, TLS, backups, shard rebalancing and upgrades all become your problem. Qdrant Cloud is the managed alternative, with a free tier limited to one node, 1GB RAM and 4GB disk.
Qdrant or pgvector?
Below a few million vectors, on a database you already run, pgvector wins on everything except throughput: it is one extension rather than a service. Qdrant earns its keep when a metadata filter sits on every query, because it can feed that filter into the HNSW traversal instead of retrieving candidates and discarding them.
How much does Qdrant Cloud cost?
The vendor publishes no entry price. Its pricing page describes hourly usage billing for vCPU, memory, storage, backup storage and inference tokens and links to a calculator; the free tier is 1GB RAM and 4GB disk, the Standard tier carries a 99.5% uptime SLA and the Premium tier adds SSO, private VPC links and 99.9%. The widely quoted figure of about $25 a month for the smallest paid cluster is a third-party estimate, not a published rate.
What happens when I turn quantisation on?
The compressed vectors are stored next to the originals, so nothing is destroyed and quantisation can be switched off. Rescoring of the top-k against full-precision vectors is on by default for binary quantisation and for TurboQuant at 1, 1.5 and 2 bits, and off for scalar and product quantisation. The production checklist tells you to re-benchmark retrieval quality afterwards, because some embedding models quantise badly.