> pgvector, Qdrant, Weaviate, Milvus, Pinecone, OpenSearch or Elasticsearch? A practical 2026 guide to filtering, hybrid search, scale, cost and EU hosting.
>
> Web page: https://balazscsorba.com/blog/pgvector-vs-vector-databases · Language: English · Also available in: [Deutsch](https://balazscsorba.com/de/blog/pgvector-vs-vector-databases.md) · [Magyar](https://balazscsorba.com/hu/blog/pgvector-vs-vector-databases.md)
> Author: Balázs Csorba · Published: 2026-10-02 · Keywords: pgvector vs vector database, pgvector vs Qdrant, pgvector vs Pinecone, best vector database 2026, pgvector HNSW iterative scan, pgvector halfvec, OpenSearch vs Elasticsearch vector search, vector database EU hosting, hybrid search Postgres, Weaviate vs Milvus

[Blog](https://balazscsorba.com/blog)/RAG & retrieval

# pgvector or a vector database? How to choose vector storage in 2026

pgvector, Qdrant, Weaviate, Milvus, Pinecone, OpenSearch or Elasticsearch? A practical 2026 guide to filtering, hybrid search, scale, cost and EU hosting.

[Balázs Csorba](https://balazscsorba.com/about)·October 2, 2026·13 min read

-   pgvector
-   Vector databases
-   RAG
-   Hybrid search
-   EU hosting

![Diagram: a decision path from your data to pgvector in Postgres, a search engine with vector fields, or a dedicated vector database.](https://balazscsorba.com/images/blog/pgvector-vs-vector-databases/cover.webp?v=df702aff6b)

## Key takeaways

-   Start where your data already lives: if your records are in Postgres, pgvector with HNSW, halfvec and iterative scans covers most RAG workloads without a second system to run.
-   Filtering decides the choice more than raw speed. Test your real filters, because a filter applied after an approximate index scan can return too few results.
-   Memory is the scale threshold you can calculate: a float32 vector with 1,536 dimensions takes 6,144 bytes before any index overhead, and halfvec halves that.
-   If you already run OpenSearch or Elasticsearch for keyword search, adding vector fields is often cheaper than introducing a new database.
-   Pick a dedicated vector database when vectors are the product: very large corpora, many tenants, or a team that owns search. Check EU regions and operations before you sign.

On this page

1.  [The short answer](https://balazscsorba.com/#short-answer)
2.  [What pgvector can do today](https://balazscsorba.com/#pgvector-today)
3.  [Filtering is where choices get decided](https://balazscsorba.com/#filtering)
4.  [Hybrid search: built in, or build it yourself](https://balazscsorba.com/#hybrid-search)
5.  [A decision diagram](https://balazscsorba.com/#decision-diagram)
6.  [The options side by side](https://balazscsorba.com/#comparison)
7.  [Scale thresholds and cost](https://balazscsorba.com/#scale-and-cost)
8.  [EU hosting and data protection](https://balazscsorba.com/#eu-hosting)
9.  [Operational burden](https://balazscsorba.com/#operations)
10.  [A checklist before you decide](https://balazscsorba.com/#checklist)
11.  [What I would do](https://balazscsorba.com/#what-i-would-do)
12.  [Sources](https://balazscsorba.com/#sources)

Every RAG project reaches the same meeting. Somebody opens a slide with six vector database logos, somebody else says "we already have Postgres", and the discussion turns into taste. In 2026 that discussion is less about which engine is fastest, and more about what you already operate, how your queries are filtered, and who gets paged when the index falls over.

I have shipped retrieval on top of relational databases, search engines and dedicated vector stores. My honest summary is that the choice rarely hinges on a benchmark. It hinges on **filters, hybrid search, memory, hosting and operations**, in roughly that order. This article walks through those five, with a decision diagram and a comparison table at the end.

One note on evidence. Everything concrete below comes from the vendors' own documentation and release notes, which I read in the first days of October 2026. I deliberately do not quote benchmark numbers: vector benchmarks depend on dataset, recall target, hardware and filters, and the ones that circulate are mostly vendor-run. Where I give a threshold, I tell you it is my judgement.

## The short answer

If I had to give one sentence: **use the store your team already knows how to run, and leave it only when you can name the measured limit that forces you out.** For most European B2B teams that means Postgres with pgvector, or the search engine they already operate.

**My default order**

First choice: pgvector, when your source data is in Postgres and the vectors fit in memory.

Second: OpenSearch or Elasticsearch, when keyword search, facets and logs already live there.

Third: a dedicated vector database, when vectors, filtering or tenancy are the heart of the product.

The rest of the article explains why, and how to find out which situation you are in. If you are still designing the retrieval layer itself, read [my RAG pipeline guide on chunking, hybrid search and reranking](https://balazscsorba.com/blog/rag-pipeline-chunking-hybrid-search-reranking) first: storage is the smaller decision.

## What pgvector can do today

pgvector has changed a lot since the early IVFFlat-only days. The [changelog](https://github.com/pgvector/pgvector/blob/master/CHANGELOG.md) shows the milestones that matter for RAG, and the latest release I saw was 0.8.7 on 1 October 2026.

-   **HNSW indexes** arrived in 0.5.0 (August 2023), together with parallel IVFFlat builds.
-   **halfvec and sparsevec** arrived in 0.7.0 (April 2024), along with binary quantization functions and indexing for the bit type.
-   **Iterative index scans** arrived in 0.8.0 (October 2024), which is the feature that makes filtered queries behave.
-   **Patch releases matter.** 0.8.3 (June 2026) fixed possible index corruption with HNSW vacuuming and 0.8.4 fixed an HNSW repair error, so stay on the latest 0.8.x patch.

The [README](https://github.com/pgvector/pgvector/blob/master/README.md) documents the limits you will design around. A vector column can be indexed up to 2,000 dimensions, halfvec up to 4,000, bit up to 64,000, and sparsevec up to 1,000 non-zero elements. HNSW defaults are m = 16, ef\_construction = 64 and ef\_search = 40. The index builds fastest when the graph fits in \`maintenance\_work\_mem\`, and halfvec lets you index the same embeddings at half the storage by indexing an expression such as \`embedding::halfvec(1536)\` and querying with the same cast.

What you get beyond the index is the real argument for pgvector: embeddings sit next to the rows they describe. Joins, transactions, row-level security, backups and point-in-time recovery are the ones you already run. A deleted customer is deleted in one place, which matters for GDPR. What you give up is isolation: vector queries and index builds compete with your OLTP traffic for memory and CPU unless you use a replica.

If you outgrow plain pgvector but want to stay in Postgres, [pgvectorscale](https://github.com/timescale/pgvectorscale) adds a StreamingDiskANN index, statistical binary quantization and label-based filtered search under the PostgreSQL licence. At the time of writing its managed offering on Timescale Cloud was a private beta, and the vendor benchmarks it publishes are exactly the kind of numbers I would reproduce on your data before believing.

## Filtering is where choices get decided

Real RAG queries are never "nearest neighbours of this vector". They are "nearest neighbours among documents this user may see, in this language, from this year". An approximate index walks a graph, and a filter that is applied afterwards throws results away.

pgvector is explicit about it. With the default hnsw.ef\_search of 40, a query with a WHERE clause that matches about 10 percent of rows will typically keep only around four of the 40 candidates. Iterative scans fix this: set \`hnsw.iterative\_scan\` to \`strict\_order\` or \`relaxed\_order\` and the index keeps scanning until enough rows match or \`hnsw.max\_scan\_tuples\` (default 20,000) is reached. With \`relaxed\_order\` the results can be slightly out of order, and the README shows how to restore the order with a materialised CTE.

-   **Few distinct filter values:** the README suggests partial indexes.
-   **Many distinct values, such as one tenant per customer:** partition the table.
-   **Low match rate:** an ordinary B-tree index on the filter column, so Postgres can choose an exact scan.

Dedicated systems treat filtering as a design goal. Qdrant [recommends payload indexes](https://qdrant.tech/documentation/concepts/filtering/) on every field you filter by, and supports nested must, should and must\_not conditions plus range, geo and full-text conditions. OpenSearch documents [efficient filtering](https://docs.opensearch.org/latest/vector-search/filter-search-knn/efficient-knn-filtering/) in which both Faiss and Lucene engines apply the filter during graph traversal, so restrictive filters still return an accurate top-k. Elasticsearch describes its kNN filter as a pre-filter applied during the approximate search. Pinecone filters on record metadata inside its query executors.

The practical test is simple and I run it on every project: take the five most selective real filters, run 200 real queries each at the recall you need, and look at how many results come back and how latency behaves. Pair it with a labelled evaluation set, as I describe in [LLM evals for product features](https://balazscsorba.com/blog/llm-evals-for-product-features), so you measure answer quality and not only speed.

## Hybrid search: built in, or build it yourself

Dense vectors miss exact tokens such as part numbers, error codes and names, and B2B catalogues are full of them. Almost every serious RAG system ends up combining keyword and vector retrieval, as I argue in [RAG in 2026: hybrid, agentic and long-context](https://balazscsorba.com/blog/rag-2026-hybrid-agentic-long-context).

The engines differ in how much of that you get for free. Qdrant [supports sparse and dense vectors](https://qdrant.tech/documentation/concepts/hybrid-queries/) in one query, with Reciprocal Rank Fusion and Distribution-Based Score Fusion, nested prefetch stages and multi-vector support for ColBERT-style re-scoring. Weaviate runs BM25 and vector search in parallel and fuses them, with relative score fusion as the default and an alpha parameter that defaults to 0.75. Milvus can search several vector fields, dense and sparse, in one collection. Pinecone supports sparse vectors, hybrid queries and BM25-based full-text search. OpenSearch and Elasticsearch are search engines first, so BM25 and vectors live in one index.

Postgres gives you the pieces: full-text search with tsvector and ranked vector results from pgvector, fused with Reciprocal Rank Fusion in SQL. That is about thirty lines you own and can test, and it is a good trade when you value one transactional store. If your team would rather not own the fusion code and the relevance tuning around it, that is a legitimate reason to choose Weaviate or Qdrant, or to stay in a search engine.

## A decision diagram

This is the order in which I ask the questions. It is deliberately biased towards the systems you already run.

Ask in order and stop at the first yes. The point is to avoid adding a system, not to avoid choosing one.

Two caveats. The first "yes" does not end the conversation if the corpus is enormous: do the memory arithmetic in the scale section. And "start with pgvector" for the fallback is my bias for small teams, because it is the cheapest option to reverse: an export of vectors and metadata is all a migration needs.

## The options side by side

This table compresses what I verified in the documentation. Versions are the latest releases I saw on GitHub at the start of October 2026: Qdrant 1.19.1, Weaviate 1.39.8, Milvus 3.0.2, OpenSearch 3.9.0 and pgvector 0.8.7.

Option

Strongest at

Filtering

Hybrid search

Watch out for

**pgvector**

Vectors next to relational data, one system

WHERE plus iterative scans, partial indexes, partitions

Full-text search plus your own fusion in SQL

Index memory, competes with OLTP, indexed dimension limits

**Qdrant**

Filter-heavy retrieval, flexible hybrid queries

Payload indexes, nested boolean conditions

Sparse and dense, RRF, DBSF, prefetch

A second system to sync and secure

**Weaviate**

Built-in hybrid search, multi-tenancy

Filtered vector search; I did not verify details, test yours

BM25 plus vector, relative score fusion by default

Pricing scales with vector dimensions in the cloud

**Milvus**

Very large, distributed workloads

Metadata filters; I did not verify details, test yours

Multiple dense and sparse vector fields

Distributed deployment on Kubernetes is real operational work

**Pinecone**

Managed, serverless, no servers to run

Metadata filters inside query executors

Sparse vectors, hybrid, BM25 full text

Serverless only in the docs I read, region fixed at creation

**OpenSearch**

One engine for keywords, facets and vectors

Filtering during Faiss or Lucene graph traversal

BM25 and vectors in one index

Cluster tuning, JVM and shard planning

**Elasticsearch**

Same, with Elastic tooling and BBQ quantization

Pre-filter during approximate kNN

BM25 and kNN in one index

Cluster sizing and operations

A few details behind the cells. Weaviate offers HNSW, flat, dynamic and HFresh index types, where dynamic switches from flat to HNSW above a threshold (default 10,000 objects) and suits many small tenants. Milvus documents HNSW, IVF, DiskANN, ScaNN and GPU indexes, and a stateless, decoupled architecture. Elasticsearch documents that new indices with float vectors of 384 dimensions or more default to BBQ HNSW. OpenSearch supports Lucene (HNSW) and Faiss (HNSW and IVF) engines.

## Scale thresholds and cost

The one threshold I can give you without a benchmark is arithmetic. A float32 vector takes 4 bytes per dimension. At 1,536 dimensions that is 6,144 bytes, so 1 million vectors need about 6.1 GB and 10 million about 61 GB before graph links, metadata and replicas. halfvec halves it to roughly 3 GB and 31 GB, and binary quantization shrinks it far more at the cost of recall that you must measure.

HNSW wants its graph in memory, so the practical question is whether that memory fits on the box you would rent anyway. My rule of thumb, and it is judgement and not measurement: up to a few million chunks, a well-sized Postgres instance is rarely the bottleneck. Somewhere in the tens of millions, or when index builds start hurting your primary, I begin comparing disk-based or quantized options in dedicated systems.

-   **Self-hosted Postgres or managed Postgres:** cost is the instance you already pay for plus extra RAM. No new vendor.
-   **Search engine cluster:** you pay per node for memory and disk, but share it with keyword search and logs.
-   **Dedicated cloud service:** you pay for capacity or usage. Weaviate bills mainly by vector dimensions, so lower-dimension or compressed embeddings cut the bill directly.
-   **Hidden cost:** a second system means a second sync pipeline, a second access model and a second incident channel.

Cost also depends on the embeddings you choose. A smaller model, or a model that supports shortened vectors, reduces storage everywhere. I cover the wider levers in [LLM cost, latency, prompt caching and routing](https://balazscsorba.com/blog/llm-cost-latency-prompt-caching-routing).

## EU hosting and data protection

For European clients, "where do the vectors live" is a procurement question. Embeddings are derived from your documents and can leak information about them, so treat them as personal data when the source is.

-   **Postgres:** any EU region of a managed provider, or your own servers. Simplest to audit.
-   **Pinecone:** documents AWS eu-west-1 (Ireland) and eu-central-1 (Frankfurt) and GCP europe-west4 (Netherlands) on Builder plans and above. The Starter plan is limited to AWS us-east-1, and the region cannot be changed after creation.
-   **Weaviate Cloud:** EU regions on shared and dedicated deployments, with a bring-your-own-cloud option listed as coming soon.
-   **Qdrant:** managed clusters on AWS, GCP and Azure, plus Hybrid Cloud, a self-managed option.
-   **Self-hosted Milvus, Qdrant, Weaviate, OpenSearch:** you decide the data centre, and you carry the operations.

A region in Frankfurt is necessary but not always sufficient: a US-headquartered provider can still raise questions about foreign access, and the embedding model call is a second data flow. My GDPR notes in [GDPR and LLM APIs: EU data residency](https://balazscsorba.com/blog/gdpr-llm-api-eu-data-residency) cover that part. Check the current region list on the vendor's page at signing time, since these change.

## Operational burden

This is the cost that benchmarks never show. Ask four questions of any option.

-   **Backups and restore:** can you restore vectors and metadata to a consistent point? Postgres has this solved. For others, test the restore, not only the backup.
-   **Re-indexing:** changing the embedding model means re-embedding the corpus. Can you run the new index beside the old one and switch over?
-   **Upgrades:** pgvector patches such as 0.8.3 fixed index corruption cases. Who watches the release notes?
-   **On-call:** a distributed system on Kubernetes is a different commitment from an extension in a database you already run.

In my experience the cheapest operations come from running one fewer system. That is the whole argument for pgvector and for search engines, and it is why a managed dedicated service is the right answer when your team has no appetite for running any of them.

## A checklist before you decide

1.  Write down your top five filters, including tenant and permission checks.
2.  Count chunks, dimensions and growth, then do the memory arithmetic for float32 and halfvec.
3.  Build a labelled evaluation set of at least a few dozen real questions.
4.  Run the same queries against your top two candidates with your real filters and compare recall and latency.
5.  Decide who owns fusion and relevance tuning for hybrid search.
6.  Confirm EU region, backup, restore and upgrade procedure in writing.
7.  Plan the re-embedding path before you load the first million vectors.

If two options tie, pick the one with fewer moving parts. You can always move later: vectors and metadata export cleanly.

## What I would do

For a typical European B2B project with a catalogue or knowledge base in the low millions of chunks, I would start with pgvector, halfvec, an HNSW index and iterative scans, add Postgres full-text search for hybrid retrieval, and write the evaluation set on day one. If the client already runs OpenSearch, I would put the vectors there.

I would move to a dedicated vector database when a measured limit appears: filtered recall that iterative scans cannot rescue, index builds that hurt production, or tenancy and scale that Postgres should not carry. And I would choose it for its operating model, managed or self-hosted in the EU, as much as for its features. If you want help making that call on your own data, see [my AI engineering work](https://balazscsorba.com/expertise/ai-engineer).

## Sources

1.  [pgvector README (index limits, HNSW defaults, iterative scans, filtering, halfvec)](https://github.com/pgvector/pgvector/blob/master/README.md)
2.  [pgvector CHANGELOG (0.4.0 to 0.8.7)](https://github.com/pgvector/pgvector/blob/master/CHANGELOG.md)
3.  [pgvectorscale: StreamingDiskANN, statistical binary quantization, filtered search](https://github.com/timescale/pgvectorscale)
4.  [Qdrant documentation: Filtering](https://qdrant.tech/documentation/concepts/filtering/)
5.  [Qdrant documentation: Hybrid queries](https://qdrant.tech/documentation/concepts/hybrid-queries/)
6.  [Qdrant documentation: Create a cluster (providers, free tier, Hybrid Cloud)](https://qdrant.tech/documentation/cloud/create-cluster/)
7.  [Weaviate documentation: Hybrid search](https://docs.weaviate.io/weaviate/concepts/search/hybrid-search)
8.  [Weaviate documentation: Vector index types](https://docs.weaviate.io/weaviate/concepts/vector-index)
9.  [Weaviate Cloud pricing and deployment options](https://weaviate.io/pricing)
10.  [Milvus documentation: Overview](https://milvus.io/docs/overview.md)
11.  [Pinecone documentation: Database architecture](https://docs.pinecone.io/guides/get-started/database-architecture)
12.  [Pinecone documentation: Create an index (clouds, regions, sparse and hybrid)](https://docs.pinecone.io/guides/index-data/create-an-index)
13.  [OpenSearch documentation: Methods and engines](https://docs.opensearch.org/latest/mappings/supported-field-types/knn-methods-engines/)
14.  [OpenSearch documentation: Efficient k-NN filtering](https://docs.opensearch.org/latest/vector-search/filter-search-knn/efficient-knn-filtering/)
15.  [Elasticsearch documentation: Dense vector search](https://www.elastic.co/docs/solutions/search/vector/dense-vector)
16.  [Elasticsearch documentation: kNN query (filter as pre-filter)](https://www.elastic.co/docs/reference/query-languages/query-dsl/query-dsl-knn-query)
17.  [GitHub releases: Qdrant, Weaviate, Milvus, OpenSearch, pgvectorscale (versions as of 1 October 2026)](https://github.com/qdrant/qdrant/releases)

## Frequently asked questions

Is pgvector good enough for production RAG?

For many workloads, yes. pgvector supports HNSW and IVFFlat indexes, halfvec and sparsevec types, and since version 0.8.0 iterative index scans for filtered queries. If your data is already in Postgres and the index fits in memory, it is a sound default. Load-test with your own filters and data before you commit.

When should I use a dedicated vector database instead of pgvector?

When the vector workload outgrows what you want to run next to your transactional data: a very large corpus, heavy filtering across many tenants, specialised index types, or a team that treats search as its own product. Then Qdrant, Weaviate, Milvus or Pinecone give you purpose-built features and independent scaling.

What are iterative index scans in pgvector?

Iterative scans, added in pgvector 0.8.0, let an approximate index keep scanning until enough rows pass your WHERE filter. Without them, an HNSW query visits about hnsw.ef\_search candidates (default 40) and filters afterwards, so a selective filter can return fewer rows than you asked for. You enable them with hnsw.iterative\_scan.

Can I do hybrid search in Postgres?

Yes. pgvector combines with PostgreSQL full-text search, and you can merge the two ranked lists with Reciprocal Rank Fusion in SQL or re-rank with a cross-encoder. Dedicated systems such as Qdrant, Weaviate and Milvus ship hybrid queries as a built-in feature, which saves you writing the fusion yourself.

How do I keep vector data in the EU?

Choose a managed service with an EU region or self-host in an EU data centre. Pinecone lists AWS Frankfurt and Ireland and GCP Netherlands, and Weaviate Cloud supports EU regions. Pin the region at creation, because Pinecone documents that it cannot be changed afterwards, and check where your embedding model runs too.

Do I need a vector database for a small RAG app?

Usually not. A few hundred thousand chunks fit comfortably in Postgres with pgvector or even in an in-process index. Start with the simplest store that supports your filters, measure retrieval quality with an evaluation set, and migrate only when a measured limit forces you to.

Written by Balázs Csorba

Senior fullstack & AI engineer in Styria, Austria – 10+ years of Vue, Nuxt, Node.js and PHP, now building tooling for AI agents.

[AI engineering & MCP servers →](https://balazscsorba.com/expertise/ai-engineer)[About me →](https://balazscsorba.com/about)

## More articles

-   [Reducing LLM hallucinations in production: grounding, citations and knowing when to say no](https://balazscsorba.com/blog/llm-hallucination-grounding-citations)
-   [GraphRAG and knowledge-graph RAG: when a graph beats vector search](https://balazscsorba.com/blog/graphrag-knowledge-graph-rag)
-   [Evaluating RAG: retrieval metrics, faithfulness and how to tell which half failed](https://balazscsorba.com/blog/rag-evaluation-metrics)
-   [Semantic product search for B2B shops: part numbers, hybrid retrieval and what to measure](https://balazscsorba.com/blog/semantic-product-search-b2b)

## Sounds like what you need?

Tell me about your project or role – I’d love to hear from you.

[Book a call](mailto:contact@balazscsorba.com) [Connect on LinkedIn](https://www.linkedin.com/in/balazs-csorba)
