Tools/RAG & retrieval
Weaviate: a vector database that has to win on search, not only on similarity
Weaviate review: hybrid BM25 and vector search in one query, HNSW and the disk-based HFresh index, quantisation choices, a built-in MCP server and what the licence keys now cover.
- Type
- Vector database
- Pricing
- BSD-3 · Cloud from $25 per month
Balázs Csorba··10 min read
- Vector search
- Hybrid search
- HNSW
- Quantization
- Multi-tenancy

Key takeaways
- Weaviate answers one query with BM25 keyword search, vector similarity and structured filters at the same time, which is the reason to pick it over a bare nearest-neighbour index.
- HNSW is memory-bound: the documentation puts a node at 2 to 12 kB, so a million 1536-dimension vectors is 2 to 12 GB of index in RAM and a hundred million is 200 to 1200 GB.
- Index type and quantisation width are effectively creation-time decisions. Rotational quantisation bits are fixed the first time RQ is enabled and cannot be migrated afterwards.
- The vendor benchmark for DBPedia with OpenAI ada002 embeddings reports 97.24 percent recall@10 at 5639 queries per second and 4.43 ms p99 on a single 16 vCPU, 128 GB machine, unfiltered.
- The repository is BSD-3 outside the wl directory, and v1.40 gates Namespaces and deduplicated backups behind a Weaviate licence key in the same binary.
Weaviate is a vector database that stores objects and their embeddings side by side and answers one query with BM25 keyword search, vector similarity and structured filters at once. It is the most complete search engine in the open-source category. The reason to hesitate is not retrieval quality: it is the memory bill, and the number of index and quantisation decisions that have to be made before the first import.
It competes with Qdrant on memory-efficient indexes, with Pinecone on managed operations, and with pgvector on the argument that a team which already runs a database should not add a second one. Weaviate's answer is that it is a full database first, replication, backups, multi-tenancy, RBAC and incremental schema changes included, with nearest-neighbour search attached.
What it is
The project comes out of the Dutch company Weaviate B.V., is written in Go, and the current release is 1.40.0, tagged on 7 October 2026. It ships official clients for Python, JavaScript, Java, Go and C#, and speaks REST, gRPC and GraphQL. The core facts worth knowing before a comparison:
- Hybrid search fuses BM25 and vector results in a single query, with an alpha parameter weighting the two halves.
- Four vector index types: flat, HNSW, dynamic (a flat index that upgrades itself to HNSW) and HFresh, the disk-based index introduced in 1.36 and generally available in 1.38.
- Quantisation covers scalar, product, binary and rotational schemes; 4-bit rotational quantisation is a preview, and v1.40 adds RQ-4 to HNSW indexes.
- A built-in MCP server, generally available since 1.38, exposes four tools at /v1/mcp on the REST port.
- Multi-tenancy, replication, backups, incremental backups, object TTL and collection aliases are all in the open-source build rather than the paid one.
How it works
A query meets two index families. Object properties live in inverted indexes, the same BM25 machinery an inverted-index search engine uses, so a filter narrows the candidate set before anything expensive happens. Vectors live in the vector index, where HNSW walks a layered graph held in memory. What comes back is fused, optionally boosted and optionally reranked.
Two consequences follow. Filtered search stays cheap because the filter runs first. And the expensive part is always the vector index, which is where the operational cost sits. QUERY_HYBRID_MAXIMUM_RESULTS defaults to 200, so each half of a hybrid query retrieves at least that many candidates before fusion; it is the first knob to turn down when a query is slower than its benchmark suggested.
Getting started
The Python client is the one to reach for. A collection with an HNSW index, quantisation, hybrid search, a filter, a booster and diversity selection, in about thirty lines:
from datetime import timedelta
import weaviate
from weaviate.classes.config import Configure, VectorDistances
from weaviate.classes.query import Boost, Diversity, Filter
client = weaviate.connect_to_local() # or connect_to_cloud(cluster_url, auth)
client.collections.create(
"Product",
vector_config=Configure.Vectors.text2vec_openai(
source_properties=["name", "description"],
vector_index_config=Configure.VectorIndex.hnsw(
distance_metric=VectorDistances.COSINE,
quantizer=Configure.VectorIndex.Quantizer.rq(bits=8),
),
),
)
products = client.collections.get("Product")
products.data.insert_many([
{"name": "Kestrel Wireless Headphones", "in_stock": True, "released": "2026-09-20"},
{"name": "Aurora Wireless Headphones", "in_stock": False, "released": "2025-11-02"},
{"name": "Nimbus Wireless Headphones", "in_stock": True, "released": "2026-10-05"},
])
boost = Boost.blend(
[
Boost.filter(Filter.by_property("in_stock").equal(True), weight=2.0),
Boost.time_decay("released", scale=timedelta(days=30)),
],
weight=0.3, # 30 percent boost, 70 percent original relevance
depth=200, # re-score the top 200 candidates
)
hits = products.query.hybrid(
query="wireless headphones",
limit=4,
filters=Filter.by_property("released").greater_than("2025-01-01"),
boost=boost,
diversity_selection=Diversity.mmr(limit=4, balance=0.3),
)
for hit in hits.objects:
print(hit.properties["name"])
Four details in that snippet decide how the page looks. Boost.blend never removes a result, it only re-sorts, so a negative condition weight demotes instead of filtering. depth sets how many candidates are rescored and the outer weight decides how much of the final score the boost owns. MMR needs client 4.23.0 or newer, and its balance default is 0.0, which means pure diversity rather than a neutral midpoint. If boost and rerank are combined, the reranker runs last and has the final word.
Indexes, memory and the bill
HNSW holds nodes and edges in memory, and the documentation is unusually direct about the cost: a node is 2 to 12 kB depending on dimensionality, so one million vectors is 2 to 12 GB and a hundred million is 200 to 1200 GB, with edges adding roughly 200 bytes per vector. That is a number for a capacity plan, not a throughput target.
| Index | Memory | Search behaviour | Use it when |
|---|---|---|---|
| Flat | Very low | Exact, linear scan | Small collections and per-tenant datasets in a multi-tenant setup |
| HNSW | High, everything resident | Fastest; logarithmic in the graph | Large collections with high query throughput |
| HFresh | Low, disk-backed postings | Reads a few posting lists, then rescores | Memory is the binding constraint; tolerates slightly higher p99 |
HFresh is the interesting one for a large corpus. It groups vectors into on-disk postings, keeps a small 8-bit-quantised HNSW index over their centroids in memory, and stores the postings themselves at 1-bit. Search reads only the postings the centroid index selects, then rescores the candidates against uncompressed vectors. The documentation is careful that it supports only cosine and l2-squared distances, that dot product is not available, and that it is not designed to beat HNSW on raw throughput. Read the recall parameters as a latency dial: searchProbe sets how many posting lists a query visits, replicas how many lists each vector joins, and maxPostingSizeKB the cluster size.
Quantisation is the other lever. Rotational quantisation rotates a vector so its values spread evenly across the dimensions, then stores each dimension as a small integer. At 4 bits and 1536 dimensions a vector costs 784 bytes against 6144 for raw float32, a factor of 7.84 rather than a round 8, because the 16-byte rotation header stays. Compression cuts memory and compute but not the number of dimensions Cloud bills for: Weaviate Cloud prices from $0.00465 per million vector dimensions per month on Flex, storage from $0.12 per GiB, with a $45 monthly minimum, and more aggressive compression shows up as a lower per-dimension rate rather than a smaller bill.
The vendor benchmark is worth reading with its methodology attached. On DBPedia embedded with OpenAI ada002, one million objects at 1536 dimensions and cosine distance, the recommended configuration of efConstruction 256, maxConnections 16 and ef 96 yields 97.24 percent recall@10 at 5639 queries per second, 2.80 ms mean latency and 4.43 ms p99. That is 10,000 unfiltered searches on one GCP n4-highmem-16 instance with 16 vCPU and 128 GB, driven by the Go client from the same VPC, with every matched object read back from disk. The scripts are open source, which is the part that matters.
Where it shingle
The weaknesses first, because they decide whether this is your database. Growth on HNSW is a memory curve rather than a horizontal one: a hundred million 1536-dimension vectors is a multi-hundred-gigabyte planning exercise, and adding nodes does not shrink one shard's index. Deletion is asynchronous, so a search straight after a delete can still return the object. The API surface is broad enough that GraphQL, gRPC, gRPC-Web and a fourth experimental REST search API coexist, and the 1.39 release notes state that reference selection in that new REST API is being replaced, so code written against it will change.
| Database | Index model | Self-host footprint | Where it hurts |
|---|---|---|---|
| Weaviate | HNSW, flat, dynamic, HFresh | A database: replication, backups, RBAC, multi-tenancy | Memory-bound growth; more schema surface to learn |
| Qdrant | HNSW with on-disk and scalar quantisation | A focused vector engine with filtering | Fewer of the database features Weaviate ships |
| pgvector | Postgres indexes: HNSW, IVFFlat | None: it is an extension on a database you already run | Recall and tuning are Postgres tuning now |
| Pinecone | Managed only | Nothing to run | No self-hosting, and a bill that scales with read units |
Two operational notes finish the picture. Async replication was rebuilt in 1.38 to run cluster-wide from a single scheduler and is on by default for every replicated collection, which is a reliability win and a background-load cost at the same time. And in 1.40 the new Namespaces feature, which adds control-plane and data isolation between users sharing one cluster, is gated behind a Weaviate licence key.
MCP access and the licence boundary
The MCP server is what most teams will meet first. It is a Streamable HTTP server at /v1/mcp on the REST port, disabled by default when self-hosting, always enabled in Weaviate Cloud, and authenticated with an API key as a bearer token. It exposes four tools: weaviate-collections-get-config for schemas, weaviate-tenants-list, weaviate-query-hybrid for a hybrid search with an alpha that defaults to 0.75, and weaviate-objects-upsert for writes. Permissions are the normal RBAC roles, checked when a tool is called.
On licensing the repository LICENSE is unambiguous: code outside the wl directory is BSD-3-Clause, code inside it is Copyright Weaviate B.V. and available only under a separate enterprise licence unlocked with a licence key, and the BSD licence does not grant the right to use those features or to circumvent the key. v1.40 puts Namespaces and deduplicated backups on that side of the line. The database itself remains BSD-3 and self-hostable without a usage limit, so the only real question is how much of the roadmap ends up behind the key.
Verdict
Weaviate is the vector database to choose when retrieval quality and operational completeness matter more than the smallest possible memory footprint, and when the team can afford to treat index type, quantisation width and ef as real parameters rather than defaults. It is the wrong answer for a team that already runs Postgres and wants embeddings next to the rows, and the wrong answer for a hundred-million-vector corpus on a fixed memory budget.
- Choose it if queries need filters and keywords as well as vectors, and you want that in one round trip rather than three.
- Choose it if you need multi-tenancy, replication, backups and RBAC without bolting four services onto a bare vector index.
- Choose it if you can self-host a BSD-3 database and want a managed option behind the same API for when you cannot.
- Avoid it if the corpus is large enough that HNSW memory dominates the bill and your queries would tolerate a slightly higher p99; that is HFresh's job, or a purpose-built engine's.
- Avoid it if you expect the whole roadmap to stay behind the BSD licence. Check which features your version gates behind a licence key before you design around them.
Sources
Frequently asked questions
Is Weaviate still free to self-host?
The database is BSD-3-Clause and can be self-hosted with no usage limit. The repository LICENSE is explicit that code inside the wl directory is proprietary and unlocked with an enterprise licence key, and v1.40 puts Namespaces and deduplicated backups behind that key. Cloud pricing has three tiers: a free tier, Flex from $45 per month and Premium from $400 per month.
HNSW or HFresh for a large collection?
HNSW is the fastest option and the default, but it holds the whole graph in memory. HFresh keeps only a compressed centroid index in RAM and reads posting lists from disk, which the documentation says is not designed to beat HNSW on raw throughput. Choose HFresh when memory is the binding constraint and the workload tolerates a slightly higher p99.
Does quantisation reduce the Weaviate Cloud bill?
No. Cloud billing is based on the number of stored vector dimensions, plus storage and backups, and compression does not change the dimension count. It reduces the memory and compute Weaviate needs, which shows up in the per-dimension list rate rather than in the number of dimensions billed.
Can an agent query Weaviate over MCP?
Yes, since v1.38 the database ships an MCP server at /v1/mcp on the REST port with four tools: weaviate-collections-get-config, weaviate-tenants-list, weaviate-query-hybrid and weaviate-objects-upsert. It is disabled by default when self-hosting and always enabled in Weaviate Cloud, where the write tool is exposed unless the cluster's Enable MCP Read-Only switch is set.
How large can a HNSW collection get before memory becomes the limit?
The documentation sizes an HNSW node at 2 to 12 kB depending on dimensionality, plus about 200 bytes of edges per vector. That is 2 to 12 GB at a million vectors and 200 to 1200 GB at a hundred million. Rotational quantisation or HFresh are the two levers; the default vector cache is capped at 1e12 objects per collection and is a separate constraint during import.