Tools/RAG & retrieval
LanceDB: vector search that starts as a library
A review of LanceDB: an Apache-2.0 embedded vector library, its IVF and HNSW index choices, hybrid search with rank fusion, and what the Enterprise tier adds.
- Type
- Vector database
- Pricing
- Apache-2.0 · Cloud paid
Balázs Csorba··9 min read
- Vector search
- Hybrid search
- Embedded database
- RAG

Key takeaways
- LanceDB 0.40.0, published 7 October 2026, is an Apache-2.0 Rust library embedded in the application process, with Python, JavaScript and Rust clients and no server to run.
- Below roughly a million rows a vector index is optional: the vendor's FAQ measures 100,000 pairs of 1,000-dimensional vectors at under 20 ms and recommends skipping the index for small tables.
- HNSW is not a top-level index here — it exists only inside IVF partitions — and the docs warn that HNSW-backed indexes show higher latency variance under metadata filters.
- Hybrid search is the strongest feature: BM25 full-text search and vector search merged by a reciprocal rank fusion reranker, with prefiltering on by default.
- The vendor's own comparison puts the open-source build at 10 to 50 queries per second and 500 to 1,000 ms from object storage, against up to 10,000 queries and 50 to 200 ms on Enterprise.
LanceDB is an embedded vector database: a Rust library that links into the application process and stores vectors, metadata and source text in the Lance columnar format. There is no server to deploy, no cluster to size and no connection pool to tune — the strongest argument for it, and the reason it should be compared as much with SQLite as with Qdrant.
It competes with the embedded options — SQLite with a vector extension, Chroma, an in-process index — and, once pointed at object storage, with the hosted services. It replaces the pattern of running a dedicated search cluster for a corpus that fits on one machine.
What LanceDB is
The current release is 0.40.0, published on 7 October 2026 and requiring Python 3.10 or newer, with JavaScript and Rust clients against the same Rust core. Data lives wherever the connection URI points: a local directory, s3://, gs://, az://, or db:// for the Enterprise cluster, and the same Lance files are readable by both editions.
- Licence: Apache-2.0 for the library; Enterprise is a commercial product sold as managed or bring-your-own-cloud.
- Shape: embedded in your process — no daemon, no sharding, no separate query language.
- Storage: the Lance format holds vectors, metadata and raw data in one table, with versioning and zero-copy reads through Apache Arrow.
- Indexes:
IVF_RQ,IVF_PQ,IVF_HNSW_SQandIVF_HNSW_FLATfor vectors, BM25 for full text, plus scalar indexes. - Search: vector, full-text and hybrid queries with reranking, prefiltering by default, distance bounds and an exact-scan escape hatch.
- Scale guidance: comfortable on a single node; the FAQ targets roughly 10 to 50 billion rows and 10 to 30 TB before Enterprise is the answer.
- Clients: Python, JavaScript and Rust, installed with pip, npm or cargo.
How a query runs
A search without an index is a scan: every vector is compared against the query and the closest k are returned, which is exact and fast enough while the table is small. An IVF index spends training time clustering vectors into partitions, so a query compares against a few centroids first and then brute-forces inside those partitions; nprobes, which defaults to 20, is how many partitions are opened. HNSW then sits inside each partition as a second-level graph, which is why the documentation can say that HNSW is not a top-level index in LanceDB.
Full-text search is a separate BM25 index built with create_fts_index, and hybrid search runs both halves and merges them: by default with an RRF reranker, which turns each list into ranks and adds them, so a result that ranks well in either half wins without either score scale dominating. Filters passed to where are prefilters by default, applied before scoring; prefilter=False moves the filter after the sub-queries, which can return fewer than limit rows.
The storage layer decides the latency profile. On local disk reads are memory-mapped and quick; pointed at S3, GCS or Azure Blob, every cold read is a network round trip, and the vendor's own comparison puts that at 500 to 1,000 ms for the open-source build against 50 to 200 ms on Enterprise, where an NVMe cache absorbs the repeat reads.
Which index to build
The index choice is a compression decision first, and the documentation is direct about the trade:
| Priority | Index | Compression | Note |
|---|---|---|---|
| Maximum compression | IVF_RQ | About 1/32 of raw size | RaBitQ quantisation over IVF |
| Accuracy at 256 dimensions or fewer | IVF_PQ | 1/64 to 1/16 of raw size | Product quantisation, recall tuned with refine_factor |
| Best recall-to-latency trade-off | IVF_HNSW_SQ | A little over 1/4 of raw size | IVF partitions with HNSW inside, scalar quantisation |
| Highest recall, no quantisation | IVF_HNSW_FLAT | Raw size plus graph overhead | The expensive, faithful option |
Two rules from the docs matter more than the table. HNSW never appears alone: it is only ever a substructure inside IVF partitions, so there is no plain HNSW index to create. And if the workload carries metadata filters, the docs tell you to prefer IVF_RQ or IVF_PQ, because the HNSW-backed variants show higher latency variance in filtered searches.
Getting started
Install the package, point at a directory and the table is on disk. The snippet below builds both indexes and runs the hybrid query the documentation recommends, with the metadata filter applied before scoring.
import lancedb
from lancedb.rerankers import RRFReranker
db = lancedb.connect("./data") # a local directory, no server
table = db.open_table("documents")
table.create_fts_index("text") # BM25, built in the background
results = (
table.search(query_type="hybrid")
.vector(embed(query)) # your embedding model
.text("refunds within 14 days")
.where("lang = 'en'", prefilter=True) # default: filter before scoring
.rerank(RRFReranker()) # the default hybrid reranker
.limit(5)
.to_list()
)
for row in results:
print(round(row["_relevance_score"], 3), row["text"][:80])
Nothing in that snippet needs a server, a container or an API key, and the same code runs against s3:// by changing the connect URI. Full-text and vector index builds return immediately and finish in the background; wait_for_index together with index_stats is how a job checks that nothing is left unindexed.
Where it shingles
The weaknesses follow from the architecture. One process means one host: the vendor's own comparison caps the open-source build at 10 to 50 queries per second with no cache, and every maintenance task — compaction, reindexing, index fragmentation after deletes — is a job someone has to schedule. Concurrent writes are bounded by how many times a writer will retry a commit, and Python users are told not to fork.
| Engine | Licence | Index options | Operational shape |
|---|---|---|---|
| LanceDB | Apache-2.0, embedded | IVF and IVF-HNSW, BM25, scalar | A library inside your process |
| Qdrant | Apache-2.0, server | Filterable HNSW, scalar and product quantisation | A container or a managed cloud |
| pgvector | PostgreSQL licence | HNSW and IVFFlat inside Postgres | An extension in a database you already run |
| Weaviate | BSD-3-Clause | HNSW, flat and dynamic | Single binary, optional cluster |
The honest summary: LanceDB is the cheapest thing to run and the most maintenance to own. The vendor's table puts single-process throughput at 10 to 50 queries per second and object-storage latency at 500 to 1,000 ms, with distributed search and platform-managed compaction still marked as coming soon on the Enterprise side — worth knowing before a purchase decision rests on the roadmap.
API asymmetry is the other papercut: on an Enterprise RemoteTable the table-level to_arrow and to_pandas calls are refused, so materialisation has to go through the query builder, and an operation can fail on a service-level policy instead of on the data. Code that moves from OSS to Enterprise is close to the same, but not the same.
Pricing
The library is Apache-2.0 and free in every sense that matters: pip install, no account, no metering. The pricing page does not list a rate card — it is a contact form — and the Enterprise tier is sold as managed or bring-your-own-cloud with SOC 2 Type II, HIPAA coverage and OpenTelemetry metrics and traces. The one concrete number the vendor publishes is a benchmark of roughly $779 a month for 100 million vectors.
- Open source: the library, the indexes, hybrid search and the CLI maintenance calls, under Apache-2.0, with community support.
- Enterprise: managed or BYOC, distributed query nodes, an NVMe cache, platform-run indexing and compaction, and compliance under SOC 2 Type II and HIPAA.
- Coming soon, in the vendor's own table: distributed search, distributed indexing and compaction.
Verdict
LanceDB is the right answer when the corpus and the traffic fit one machine, and the wrong answer when they do not — the vendor says as much in its own comparison table. Its strength is that retrieval becomes a library call instead of an infrastructure project, and its cost is that every operational duty of a database lands on the application team. The opinionated reading: for a RAG pipeline under a million chunks, running a separate vector database is an unnecessary service to operate, and LanceDB is what that pipeline should reach for first.
- Use it for a prototype, a single-node RAG pipeline or an edge deployment, where a database server would be the heaviest component in the stack.
- Use it when the data is already in object storage and the Lance format's versioning and zero-copy reads remove a second copy of the corpus.
- Think twice when queries must stay under 100 ms from S3 with real concurrency — the vendor's own figure for OSS is 500 to 1,000 ms and 10 to 50 queries per second.
- Budget for maintenance: optimize, compaction and reindexing are unscheduled work in the open-source build, and index fragmentation grows with every delete.
- Choose the Enterprise tier only for the distributed parts; the API differences, from RemoteTable materialisation limits to cluster-side guardrails, are what lock the application in.
Sources
Frequently asked questions
Do I need a vector index in LanceDB?
Not at first. The FAQ puts brute-force search at under 20 ms for 100,000 pairs of 1,000-dimensional vectors and says a vector index becomes worthwhile beyond roughly one million rows or higher dimensions; below that, scanning is usually fast enough.
Why is there no plain HNSW index?
In LanceDB, HNSW is a substructure inside IVF partitions rather than a top-level index, which combines IVF scalability with HNSW recall. The available types are IVF_HNSW_FLAT, IVF_HNSW_PQ and IVF_HNSW_SQ, alongside the unquantised and quantised IVF variants.
How do you keep an index healthy in the open-source build?
Index builds run asynchronously and appended rows stay outside the index until they are folded in with optimize(). create_index returns immediately, wait_for_index waits for the coverage to be complete, and fast_search() skips the slower fallback path over rows that are not indexed yet.
What does LanceDB Enterprise cost?
There is no public rate card: the pricing page is a contact form, and the Enterprise tier is sold as a managed deployment or bring-your-own-cloud with SOC 2 Type II and HIPAA coverage. The vendor's published benchmark works out at roughly $779 a month for 100 million vectors.