Tools/RAG & retrieval
Milvus review: the most complete vector database to operate
Milvus 3.0.2 is the most complete open-source vector database and the heaviest to run. A review of its architecture, hybrid search, costs and where it should not be used.
- Type
- Vector database
- Pricing
- Apache-2.0 · Zilliz Cloud free tier
Balázs Csorba··10 min read
- Vector search
- Hybrid search
- BM25 full text
- Distributed
- RAG

Key takeaways
- Milvus 3.0.2, released 20 September 2026, is current while the 2.6 branch is still maintained in parallel.
- Dense, sparse and server-side BM25 vectors live in one collection, and reranking now happens inside the search request.
- Self-hosting means running etcd, object storage and a WAL layer; the distributed mode is a Kubernetes deployment.
- Zilliz Cloud lists dedicated compute at $0.273 per CU-hour and storage at $0.025 per GB-month, with a 5 GB free tier.
- Below tens of millions of vectors the architecture is a cost rather than a capability.
Milvus is an open-source vector database for similarity search over embeddings, written in Go and C++ and developed under LF AI & Data with Zilliz as its main contributor. The position of this review: at scale it is the most complete engine open source offers, and it is also the most expensive one to operate, so the real question is who runs the cluster, not which index type wins.
It occupies the retrieval layer of a RAG or search stack: the store that holds vectors, metadata and, since 3.0, long text, and that answers top-k queries under filters. It competes with Qdrant and Weaviate as self-hostable peers, with Pinecone as the managed-only service, and with pgvector for teams that would rather not operate another database. Zilliz Cloud, sold by the company that contributes most of the code, is the managed twin of the Apache-licensed project.
What it is
Three deployment shapes exist, and they are not equivalent. Milvus Lite is a local file opened by the Python client for experiments; standalone is one node plus its dependencies; the distributed mode is the real product, a Kubernetes deployment with storage and compute disaggregated.
- Apache-2.0 licence, an LF AI & Data project with Zilliz as the main contributor; written in Go and C++, with search kernels built on FAISS, HNSW, DiskANN and SCANN.
- Current release 3.0.2, published 20 September 2026, while the 2.6 branch is still maintained in parallel at 2.6.25.
- Index menu: HNSW, IVF, FLAT, SCANN, DiskANN, GPU indexes such as NVIDIA CAGRA, plus quantisation and mmap for memory-bound data.
- Dense vectors, learned sparse vectors and server-side BM25 in one collection, with hybrid search and reranking inside a single request.
- Multi-tenancy at database, collection, partition or partition-key level, behind authentication, TLS and RBAC.
- 3.0 adds external collections over Parquet, Lance, Iceberg and Vortex, snapshots, and online schema change with backfill.
- The integrations a retrieval stack expects: LangChain, LlamaIndex, Attu for administration, Prometheus and Grafana for monitoring, plus Spark and Kafka connectors.
How it works
Milvus separates the data plane from the control plane in four layers. Stateless proxies accept and reduce requests; exactly one coordinator is active at a time and schedules DDL, routing, query and compaction work; worker nodes execute without holding data of their own; storage is shared. The documentation describes Woodpecker as a zero-disk write-ahead log that writes straight to object storage, which takes local disk management off the write path.
A write is logged to the WAL first, becomes queryable in the streaming node as growing data, and stays there until compaction seals it; the data node then builds indexes and the query node loads them. A search runs against growing data locally and against sealed segments in parallel, with results reduced at three levels before the proxy returns them. Every hop is a place where consistency level and replica placement change the latency that comes back.
Getting started
The shortest path is Milvus Lite through the Python client: one pip install and a file name, no server, no etcd, no object store. The same client then points at a server or a Zilliz Cloud endpoint by changing uri and token, which is why prototypes usually move without a rewrite.
# pip install -U pymilvus — Milvus Lite stores everything in one local file
from pymilvus import MilvusClient
client = MilvusClient(uri="./milvus_demo.db")
client.create_collection(
collection_name="papers",
dimension=768, # must match the embedding model
auto_id=True,
metric_type="COSINE",
)
client.insert(collection_name="papers", data=[
{"vector": v, "title": t, "year": y} for v, t, y in rows
])
hits = client.search(
collection_name="papers",
data=[query_vector],
limit=5,
filter="year >= 2023",
output_fields=["title", "year"],
)
print([(h["entity"]["title"], round(h["distance"], 3)) for h in hits[0]])The simplified client hides schema, index parameters and dynamic fields, and that is the right level for a prototype. Production has to make these choices explicitly: metric type, index type, quantisation, mmap, partition keys for tenancy, and a consistency level per request.
Hybrid search and scale
Milvus keeps dense vectors, learned sparse vectors and BM25 output in the same collection, so one request can run several vector searches and merge them. Reranking moved into the server with 3.0: the Function Chain API composes score transformation, model-based reranking and candidate trimming inside a single search call, and weighted reciprocal-rank fusion arrived in 3.0.1.
- BM25 is computed server-side from raw text, so the application never ships tokens to the database and back.
- Sparse search in 3.0 is rebuilt around SINDI, with Block-Max WAND and Block-Max MaxScore selectable per workload.
- Faceted search on the ANN path returns the top facet values with COUNT and AVG in the same request instead of an over-fetch in the client.
- TEXT fields keep values under 64 KB inline and larger ones in partition-level LOB files, so source text and vectors are read from one store.
The 3.0.0 release notes report two internal numbers: the compressed BM25 index is roughly three times smaller than the 2.6 sparse index at comparable recall, and SINDI reaches up to about ten times the QPS of MaxScore on learned sparse embeddings. Both are the vendor’s own measurements. The more informative figure sits in 3.0.2, where an atomic refcount hotspot that accounted for around 48% of leaf CPU time in search was removed: filtered search had been paying that tax, and most production vector workloads are filtered.
Operating it and paying for it
Self-hosting Milvus means owning coordinator failover, replica placement, compaction behaviour and index-build capacity. Zilliz Cloud sells the same engine with those decisions taken off the table, and its pricing pages are explicit about what is metered.
- Free tier: 5 GB of storage, 2.5 million vCUs per month and up to 5 collections, with community support only.
- Dedicated serving compute lists at $0.273 per CU-hour for performance- and capacity-optimised clusters, and at $0.41 for tiered-storage clusters.
- Storage lists at $0.025 per GB-month for dedicated clusters and is billed hourly; backups are also $0.025 per GB-month.
- Enterprise starts at $197 per month with a 99.95% uptime SLA, audit logs, SSO and VPC peering.
- On-demand compute for lake-scale query and index jobs lists at $0.41 per CU-hour and is billed by the CU-minute.
| Plan | Compute | Storage | Positioned for |
|---|---|---|---|
| Free | 2.5M vCUs per month included | 5 GB | learning and small prototypes |
| Standard, serverless | usage-based, system-managed scaling | usage-based | prototypes and test environments |
| Standard, dedicated | $0.273 per CU-hour | $0.025 per GB-month | steady production load |
| Enterprise | from $197 per month | $0.025 per GB-month | production with an SLA and SSO |
Where it shingles
The weaknesses come first. Milvus is a distributed system with a coordinator, a WAL, an object store and index workers, and that surface shows up as operational work long before it shows up as capability. Schema changes, index rebuilds and compaction tuning are ordinary tasks with ordinary failure modes; 3.0.2 alone ships fixes for a replica whose channels all landed on one query node and left it unserviceable, and for WAL fencing that stalled writes for 45 to 60 seconds.
- Setup cost: the distributed mode is a Kubernetes deployment with operators, not a docker run.
- Small workloads pay for the architecture: below tens of millions of vectors, a single-binary store is simpler and usually quicker to query.
- Version skew is expensive: 2.6 and 3.0 are maintained in parallel, Storage V3 is off by default, and enabling it removes the rollback path to 2.6.
- The public comparison set is mostly vendor material; independent numbers at a fixed recall target are scarce.
| System | Deployment | Hybrid retrieval | Operational burden |
|---|---|---|---|
| Milvus | Lite file, Docker, or a Kubernetes cluster | dense, sparse and BM25 in one collection | coordinator, WAL and object storage to run |
| Qdrant | a single container or its own cloud | HNSW plus sparse vectors and payload filters | one stateful service |
| pgvector | an extension inside an existing Postgres | vectors next to SQL, no native BM25 | nothing beyond the database |
| Pinecone | managed only, no self-hosting | dense and sparse vectors on the service | nothing to run |
Read that table as a comparison of attention, not of features. Milvus wins when the workload is large, filtered and multi-tenant, and loses everywhere else, because the coordinator, the WAL and the index workers all need someone on call.
Verdict
Milvus is the right engine for a team that already runs distributed systems and has a workload that justifies them. It is the wrong first choice for a product still finding its retrieval quality, where iteration speed matters more than tail latency.
- Take it when you need hundreds of millions of vectors, strict tenant isolation, or dense and sparse retrieval in one store.
- Take it self-hosted only if someone on the team already operates stateful Kubernetes services; otherwise start on Zilliz Cloud and keep the migration path open.
- Skip it for a RAG prototype under a few million chunks: Lite for the experiment, then a single-binary store for production.
- Skip it if Postgres already holds the data and vector search is a side feature; pgvector keeps one backup story and one set of credentials.
- Whichever way you go, pin the version and read the release notes: 3.0 changed storage-format defaults and the new indexes are opt-in.
A vector database you cannot operate is not a cheaper database. It is an unpaid operations contract with an index attached.
Sources
Frequently asked questions
Is Milvus free to use?
The Milvus code is Apache-2.0 and self-hosting costs only infrastructure. Zilliz Cloud charges for the managed service and lists a free tier with 5 GB of storage and 2.5 million vCUs per month.
What is the difference between Milvus and Zilliz Cloud?
Milvus is the open-source project under LF AI & Data; Zilliz Cloud runs the same engine as a managed service with serverless, dedicated and bring-your-own-cloud options. Zilliz is the main contributor to the project.
Can Milvus replace Elasticsearch for full-text search?
Milvus computes BM25 server-side and can hold vectors, sparse vectors and text in one collection, which is enough for hybrid retrieval in a RAG stack. It is not a log analytics platform, so existing Elasticsearch estates are rarely replaced wholesale.
Does Milvus run on a laptop?
Yes, through Milvus Lite, which installs with pip and stores everything in a local file. Lite is for prototyping; standalone and distributed modes need etcd, object storage and a WAL layer.