[{"data":1,"prerenderedAt":911},["ShallowReactive",2],{"blog-pgvector-vs-vector-databases-en":3},{"slug":4,"published":5,"minutes":6,"category":7,"tags":8,"keywords":14,"about":25,"sources":35,"cover":87,"og":88,"expertise":89,"locales":90,"lang":91,"title":94,"description":95,"coverAlt":96,"metaTitle":97,"takeaways":98,"faq":104,"toc":123,"blocks":160,"others":634},"pgvector-vs-vector-databases","2026-10-02",13,"rag",[9,10,11,12,13],"pgvector","Vector databases","RAG","Hybrid search","EU hosting",[15,16,17,18,19,20,21,22,23,24],"pgvector vs vector database","pgvector vs Qdrant","pgvector vs Pinecone","best vector database 2026","pgvector HNSW iterative scan","pgvector halfvec","OpenSearch vs Elasticsearch vector search","vector database EU hosting","hybrid search Postgres","Weaviate vs Milvus",[26,29,32],{"name":27,"url":28},"Vector database","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FVector_database",{"name":30,"url":31},"PostgreSQL","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FPostgreSQL",{"name":33,"url":34},"Retrieval-augmented generation","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FRetrieval-augmented_generation",[36,39,42,45,48,51,54,57,60,63,66,69,72,75,78,81,84],{"title":37,"url":38},"pgvector README (index limits, HNSW defaults, iterative scans, filtering, halfvec)","https:\u002F\u002Fgithub.com\u002Fpgvector\u002Fpgvector\u002Fblob\u002Fmaster\u002FREADME.md",{"title":40,"url":41},"pgvector CHANGELOG (0.4.0 to 0.8.7)","https:\u002F\u002Fgithub.com\u002Fpgvector\u002Fpgvector\u002Fblob\u002Fmaster\u002FCHANGELOG.md",{"title":43,"url":44},"pgvectorscale: StreamingDiskANN, statistical binary quantization, filtered search","https:\u002F\u002Fgithub.com\u002Ftimescale\u002Fpgvectorscale",{"title":46,"url":47},"Qdrant documentation: Filtering","https:\u002F\u002Fqdrant.tech\u002Fdocumentation\u002Fconcepts\u002Ffiltering\u002F",{"title":49,"url":50},"Qdrant documentation: Hybrid queries","https:\u002F\u002Fqdrant.tech\u002Fdocumentation\u002Fconcepts\u002Fhybrid-queries\u002F",{"title":52,"url":53},"Qdrant documentation: Create a cluster (providers, free tier, Hybrid Cloud)","https:\u002F\u002Fqdrant.tech\u002Fdocumentation\u002Fcloud\u002Fcreate-cluster\u002F",{"title":55,"url":56},"Weaviate documentation: Hybrid search","https:\u002F\u002Fdocs.weaviate.io\u002Fweaviate\u002Fconcepts\u002Fsearch\u002Fhybrid-search",{"title":58,"url":59},"Weaviate documentation: Vector index types","https:\u002F\u002Fdocs.weaviate.io\u002Fweaviate\u002Fconcepts\u002Fvector-index",{"title":61,"url":62},"Weaviate Cloud pricing and deployment options","https:\u002F\u002Fweaviate.io\u002Fpricing",{"title":64,"url":65},"Milvus documentation: Overview","https:\u002F\u002Fmilvus.io\u002Fdocs\u002Foverview.md",{"title":67,"url":68},"Pinecone documentation: Database architecture","https:\u002F\u002Fdocs.pinecone.io\u002Fguides\u002Fget-started\u002Fdatabase-architecture",{"title":70,"url":71},"Pinecone documentation: Create an index (clouds, regions, sparse and hybrid)","https:\u002F\u002Fdocs.pinecone.io\u002Fguides\u002Findex-data\u002Fcreate-an-index",{"title":73,"url":74},"OpenSearch documentation: Methods and engines","https:\u002F\u002Fdocs.opensearch.org\u002Flatest\u002Fmappings\u002Fsupported-field-types\u002Fknn-methods-engines\u002F",{"title":76,"url":77},"OpenSearch documentation: Efficient k-NN filtering","https:\u002F\u002Fdocs.opensearch.org\u002Flatest\u002Fvector-search\u002Ffilter-search-knn\u002Fefficient-knn-filtering\u002F",{"title":79,"url":80},"Elasticsearch documentation: Dense vector search","https:\u002F\u002Fwww.elastic.co\u002Fdocs\u002Fsolutions\u002Fsearch\u002Fvector\u002Fdense-vector",{"title":82,"url":83},"Elasticsearch documentation: kNN query (filter as pre-filter)","https:\u002F\u002Fwww.elastic.co\u002Fdocs\u002Freference\u002Fquery-languages\u002Fquery-dsl\u002Fquery-dsl-knn-query",{"title":85,"url":86},"GitHub releases: Qdrant, Weaviate, Milvus, OpenSearch, pgvectorscale (versions as of 1 October 2026)","https:\u002F\u002Fgithub.com\u002Fqdrant\u002Fqdrant\u002Freleases","\u002Fimages\u002Fblog\u002Fpgvector-vs-vector-databases\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fpgvector-vs-vector-databases\u002Fog.jpg","ai-engineer",[91,92,93],"en","de","hu","pgvector or a vector database? How to choose vector storage in 2026","pgvector, Qdrant, Weaviate, Milvus, Pinecone, OpenSearch or Elasticsearch? A practical 2026 guide to filtering, hybrid search, scale, cost and EU hosting.","Diagram: a decision path from your data to pgvector in Postgres, a search engine with vector fields, or a dedicated vector database.","pgvector vs vector databases in 2026 · Balázs Csorba",[99,100,101,102,103],"Start where your data already lives: if your records are in Postgres, pgvector with HNSW, halfvec and iterative scans covers most RAG workloads without a second system to run.","Filtering decides the choice more than raw speed. Test your real filters, because a filter applied after an approximate index scan can return too few results.","Memory is the scale threshold you can calculate: a float32 vector with 1,536 dimensions takes 6,144 bytes before any index overhead, and halfvec halves that.","If you already run OpenSearch or Elasticsearch for keyword search, adding vector fields is often cheaper than introducing a new database.","Pick a dedicated vector database when vectors are the product: very large corpora, many tenants, or a team that owns search. Check EU regions and operations before you sign.",[105,108,111,114,117,120],{"q":106,"a":107},"Is pgvector good enough for production RAG?","For many workloads, yes. pgvector supports HNSW and IVFFlat indexes, halfvec and sparsevec types, and since version 0.8.0 iterative index scans for filtered queries. If your data is already in Postgres and the index fits in memory, it is a sound default. Load-test with your own filters and data before you commit.",{"q":109,"a":110},"When should I use a dedicated vector database instead of pgvector?","When the vector workload outgrows what you want to run next to your transactional data: a very large corpus, heavy filtering across many tenants, specialised index types, or a team that treats search as its own product. Then Qdrant, Weaviate, Milvus or Pinecone give you purpose-built features and independent scaling.",{"q":112,"a":113},"What are iterative index scans in pgvector?","Iterative scans, added in pgvector 0.8.0, let an approximate index keep scanning until enough rows pass your WHERE filter. Without them, an HNSW query visits about hnsw.ef_search candidates (default 40) and filters afterwards, so a selective filter can return fewer rows than you asked for. You enable them with hnsw.iterative_scan.",{"q":115,"a":116},"Can I do hybrid search in Postgres?","Yes. pgvector combines with PostgreSQL full-text search, and you can merge the two ranked lists with Reciprocal Rank Fusion in SQL or re-rank with a cross-encoder. Dedicated systems such as Qdrant, Weaviate and Milvus ship hybrid queries as a built-in feature, which saves you writing the fusion yourself.",{"q":118,"a":119},"How do I keep vector data in the EU?","Choose a managed service with an EU region or self-host in an EU data centre. Pinecone lists AWS Frankfurt and Ireland and GCP Netherlands, and Weaviate Cloud supports EU regions. Pin the region at creation, because Pinecone documents that it cannot be changed afterwards, and check where your embedding model runs too.",{"q":121,"a":122},"Do I need a vector database for a small RAG app?","Usually not. A few hundred thousand chunks fit comfortably in Postgres with pgvector or even in an in-process index. Start with the simplest store that supports your filters, measure retrieval quality with an evaluation set, and migrate only when a measured limit forces you to.",[124,127,130,133,136,139,142,145,148,151,154,157],{"id":125,"title":126},"short-answer","The short answer",{"id":128,"title":129},"pgvector-today","What pgvector can do today",{"id":131,"title":132},"filtering","Filtering is where choices get decided",{"id":134,"title":135},"hybrid-search","Hybrid search: built in, or build it yourself",{"id":137,"title":138},"decision-diagram","A decision diagram",{"id":140,"title":141},"comparison","The options side by side",{"id":143,"title":144},"scale-and-cost","Scale thresholds and cost",{"id":146,"title":147},"eu-hosting","EU hosting and data protection",{"id":149,"title":150},"operations","Operational burden",{"id":152,"title":153},"checklist","A checklist before you decide",{"id":155,"title":156},"what-i-would-do","What I would do",{"id":158,"title":159},"sources","Sources",[161,165,173,176,179,186,197,206,207,215,239,246,249,256,257,260,263,280,291,299,300,308,315,318,319,322,331,334,335,338,441,444,445,448,451,473,480,481,484,511,519,520,523,545,548,549,566,569,570,573,580,581],{"type":162,"content":163},"paragraph",[164],"Every RAG project reaches the same meeting. Somebody opens a slide with six vector database logos, somebody else says \"we already have Postgres\", and the discussion turns into taste. In 2026 that discussion is less about which engine is fastest, and more about what you already operate, how your queries are filtered, and who gets paged when the index falls over.",{"type":162,"content":166},[167,168,172],"I have shipped retrieval on top of relational databases, search engines and dedicated vector stores. My honest summary is that the choice rarely hinges on a benchmark. It hinges on ",{"tag":169,"children":170},"strong",[171],"filters, hybrid search, memory, hosting and operations",", in roughly that order. This article walks through those five, with a decision diagram and a comparison table at the end.",{"type":162,"content":174},[175],"One note on evidence. Everything concrete below comes from the vendors' own documentation and release notes, which I read in the first days of October 2026. I deliberately do not quote benchmark numbers: vector benchmarks depend on dataset, recall target, hardware and filters, and the ones that circulate are mostly vendor-run. Where I give a threshold, I tell you it is my judgement.",{"type":177,"level":178,"id":125,"text":126},"heading",2,{"type":162,"content":180},[181,182,185],"If I had to give one sentence: ",{"tag":169,"children":183},[184],"use the store your team already knows how to run, and leave it only when you can name the measured limit that forces you out."," For most European B2B teams that means Postgres with pgvector, or the search engine they already operate.",{"type":187,"variant":188,"title":189,"body":190},"callout","tip","My default order",[191,193,195],[192],"First choice: pgvector, when your source data is in Postgres and the vectors fit in memory.",[194],"Second: OpenSearch or Elasticsearch, when keyword search, facets and logs already live there.",[196],"Third: a dedicated vector database, when vectors, filtering or tenancy are the heart of the product.",{"type":162,"content":198},[199,200,205],"The rest of the article explains why, and how to find out which situation you are in. If you are still designing the retrieval layer itself, read ",{"tag":201,"to":202,"children":203},"link","\u002Fblog\u002Frag-pipeline-chunking-hybrid-search-reranking",[204],"my RAG pipeline guide on chunking, hybrid search and reranking"," first: storage is the smaller decision.",{"type":177,"level":178,"id":128,"text":129},{"type":162,"content":208},[209,210,214],"pgvector has changed a lot since the early IVFFlat-only days. The ",{"tag":211,"href":41,"children":212},"a",[213],"changelog"," shows the milestones that matter for RAG, and the latest release I saw was 0.8.7 on 1 October 2026.",{"type":216,"ordered":217,"items":218},"list",false,[219,224,229,234],[220,223],{"tag":169,"children":221},[222],"HNSW indexes"," arrived in 0.5.0 (August 2023), together with parallel IVFFlat builds.",[225,228],{"tag":169,"children":226},[227],"halfvec and sparsevec"," arrived in 0.7.0 (April 2024), along with binary quantization functions and indexing for the bit type.",[230,233],{"tag":169,"children":231},[232],"Iterative index scans"," arrived in 0.8.0 (October 2024), which is the feature that makes filtered queries behave.",[235,238],{"tag":169,"children":236},[237],"Patch releases matter."," 0.8.3 (June 2026) fixed possible index corruption with HNSW vacuuming and 0.8.4 fixed an HNSW repair error, so stay on the latest 0.8.x patch.",{"type":162,"content":240},[241,242,245],"The ",{"tag":211,"href":38,"children":243},[244],"README"," documents the limits you will design around. A vector column can be indexed up to 2,000 dimensions, halfvec up to 4,000, bit up to 64,000, and sparsevec up to 1,000 non-zero elements. HNSW defaults are m = 16, ef_construction = 64 and ef_search = 40. The index builds fastest when the graph fits in `maintenance_work_mem`, and halfvec lets you index the same embeddings at half the storage by indexing an expression such as `embedding::halfvec(1536)` and querying with the same cast.",{"type":162,"content":247},[248],"What you get beyond the index is the real argument for pgvector: embeddings sit next to the rows they describe. Joins, transactions, row-level security, backups and point-in-time recovery are the ones you already run. A deleted customer is deleted in one place, which matters for GDPR. What you give up is isolation: vector queries and index builds compete with your OLTP traffic for memory and CPU unless you use a replica.",{"type":162,"content":250},[251,252,255],"If you outgrow plain pgvector but want to stay in Postgres, ",{"tag":211,"href":44,"children":253},[254],"pgvectorscale"," adds a StreamingDiskANN index, statistical binary quantization and label-based filtered search under the PostgreSQL licence. At the time of writing its managed offering on Timescale Cloud was a private beta, and the vendor benchmarks it publishes are exactly the kind of numbers I would reproduce on your data before believing.",{"type":177,"level":178,"id":131,"text":132},{"type":162,"content":258},[259],"Real RAG queries are never \"nearest neighbours of this vector\". They are \"nearest neighbours among documents this user may see, in this language, from this year\". An approximate index walks a graph, and a filter that is applied afterwards throws results away.",{"type":162,"content":261},[262],"pgvector is explicit about it. With the default hnsw.ef_search of 40, a query with a WHERE clause that matches about 10 percent of rows will typically keep only around four of the 40 candidates. Iterative scans fix this: set `hnsw.iterative_scan` to `strict_order` or `relaxed_order` and the index keeps scanning until enough rows match or `hnsw.max_scan_tuples` (default 20,000) is reached. With `relaxed_order` the results can be slightly out of order, and the README shows how to restore the order with a materialised CTE.",{"type":216,"ordered":217,"items":264},[265,270,275],[266,269],{"tag":169,"children":267},[268],"Few distinct filter values:"," the README suggests partial indexes.",[271,274],{"tag":169,"children":272},[273],"Many distinct values, such as one tenant per customer:"," partition the table.",[276,279],{"tag":169,"children":277},[278],"Low match rate:"," an ordinary B-tree index on the filter column, so Postgres can choose an exact scan.",{"type":162,"content":281},[282,283,286,287,290],"Dedicated systems treat filtering as a design goal. Qdrant ",{"tag":211,"href":47,"children":284},[285],"recommends payload indexes"," on every field you filter by, and supports nested must, should and must_not conditions plus range, geo and full-text conditions. OpenSearch documents ",{"tag":211,"href":77,"children":288},[289],"efficient filtering"," in which both Faiss and Lucene engines apply the filter during graph traversal, so restrictive filters still return an accurate top-k. Elasticsearch describes its kNN filter as a pre-filter applied during the approximate search. Pinecone filters on record metadata inside its query executors.",{"type":162,"content":292},[293,294,298],"The practical test is simple and I run it on every project: take the five most selective real filters, run 200 real queries each at the recall you need, and look at how many results come back and how latency behaves. Pair it with a labelled evaluation set, as I describe in ",{"tag":201,"to":295,"children":296},"\u002Fblog\u002Fllm-evals-for-product-features",[297],"LLM evals for product features",", so you measure answer quality and not only speed.",{"type":177,"level":178,"id":134,"text":135},{"type":162,"content":301},[302,303,307],"Dense vectors miss exact tokens such as part numbers, error codes and names, and B2B catalogues are full of them. Almost every serious RAG system ends up combining keyword and vector retrieval, as I argue in ",{"tag":201,"to":304,"children":305},"\u002Fblog\u002Frag-2026-hybrid-agentic-long-context",[306],"RAG in 2026: hybrid, agentic and long-context",".",{"type":162,"content":309},[310,311,314],"The engines differ in how much of that you get for free. Qdrant ",{"tag":211,"href":50,"children":312},[313],"supports sparse and dense vectors"," in one query, with Reciprocal Rank Fusion and Distribution-Based Score Fusion, nested prefetch stages and multi-vector support for ColBERT-style re-scoring. Weaviate runs BM25 and vector search in parallel and fuses them, with relative score fusion as the default and an alpha parameter that defaults to 0.75. Milvus can search several vector fields, dense and sparse, in one collection. Pinecone supports sparse vectors, hybrid queries and BM25-based full-text search. OpenSearch and Elasticsearch are search engines first, so BM25 and vectors live in one index.",{"type":162,"content":316},[317],"Postgres gives you the pieces: full-text search with tsvector and ranked vector results from pgvector, fused with Reciprocal Rank Fusion in SQL. That is about thirty lines you own and can test, and it is a good trade when you value one transactional store. If your team would rather not own the fusion code and the relevance tuning around it, that is a legitimate reason to choose Weaviate or Qdrant, or to stay in a search engine.",{"type":177,"level":178,"id":137,"text":138},{"type":162,"content":320},[321],"This is the order in which I ask the questions. It is deliberately biased towards the systems you already run.",{"type":323,"attrs":324,"inner":328,"caption":329},"diagram",{"viewBox":325,"role":326,"aria-labelledby":327},"0 0 720 410","img","d1-pgv-t d1-pgv-d","\u003Ctitle id=\"d1-pgv-t\">Choosing vector storage\u003C\u002Ftitle>\u003Cdesc id=\"d1-pgv-d\">Three questions asked in order. If your data is already in Postgres and the vectors fit in memory, use pgvector. Otherwise, if you already run OpenSearch or Elasticsearch, add vector fields there. Otherwise, if the corpus is very large, tenancy is heavy or a team owns search, use a dedicated vector database. If none apply, start with pgvector and measure.\u003C\u002Fdesc>\u003Ctext x=\"20\" y=\"28\" class=\"d-title\">Choosing vector storage\u003C\u002Ftext>\u003Ctext x=\"700\" y=\"28\" text-anchor=\"end\" class=\"d-label\">author's heuristic, Oct 2026\u003C\u002Ftext>\u003Crect x=\"20\" y=\"50\" width=\"320\" height=\"66\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"180\" y=\"79\" text-anchor=\"middle\" class=\"d-text\">Data already in Postgres?\u003C\u002Ftext>\u003Ctext x=\"180\" y=\"100\" text-anchor=\"middle\" class=\"d-small\">and vectors fit in RAM\u003C\u002Ftext>\u003Cpath d=\"M340 83 H422\" class=\"d-line\" \u002F>\u003Cpath d=\"M430 83 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Ctext x=\"385\" y=\"75\" text-anchor=\"middle\" class=\"d-label\">yes\u003C\u002Ftext>\u003Crect x=\"430\" y=\"50\" width=\"270\" height=\"66\" rx=\"10\" class=\"d-mint\" \u002F>\u003Ctext x=\"565\" y=\"79\" text-anchor=\"middle\" class=\"d-text\">pgvector\u003C\u002Ftext>\u003Ctext x=\"565\" y=\"100\" text-anchor=\"middle\" class=\"d-small\">HNSW, halfvec, iterative scans\u003C\u002Ftext>\u003Cpath d=\"M180 116 V142\" class=\"d-line\" \u002F>\u003Cpath d=\"M180 150 l-5 -9 h10 z\" class=\"d-head\" \u002F>\u003Ctext x=\"190\" y=\"136\" class=\"d-label\">no\u003C\u002Ftext>\u003Crect x=\"20\" y=\"150\" width=\"320\" height=\"66\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"180\" y=\"179\" text-anchor=\"middle\" class=\"d-text\">Already run OpenSearch or Elastic?\u003C\u002Ftext>\u003Ctext x=\"180\" y=\"200\" text-anchor=\"middle\" class=\"d-small\">keyword search, facets\u003C\u002Ftext>\u003Cpath d=\"M340 183 H422\" class=\"d-line\" \u002F>\u003Cpath d=\"M430 183 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Ctext x=\"385\" y=\"175\" text-anchor=\"middle\" class=\"d-label\">yes\u003C\u002Ftext>\u003Crect x=\"430\" y=\"150\" width=\"270\" height=\"66\" rx=\"10\" class=\"d-sky\" \u002F>\u003Ctext x=\"565\" y=\"179\" text-anchor=\"middle\" class=\"d-text\">Stay in the engine\u003C\u002Ftext>\u003Ctext x=\"565\" y=\"200\" text-anchor=\"middle\" class=\"d-small\">add vector fields\u003C\u002Ftext>\u003Cpath d=\"M180 216 V242\" class=\"d-line\" \u002F>\u003Cpath d=\"M180 250 l-5 -9 h10 z\" class=\"d-head\" \u002F>\u003Ctext x=\"190\" y=\"236\" class=\"d-label\">no\u003C\u002Ftext>\u003Crect x=\"20\" y=\"250\" width=\"320\" height=\"66\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"180\" y=\"279\" text-anchor=\"middle\" class=\"d-text\">Huge corpus, many tenants?\u003C\u002Ftext>\u003Ctext x=\"180\" y=\"300\" text-anchor=\"middle\" class=\"d-small\">or a team that owns search\u003C\u002Ftext>\u003Cpath d=\"M340 283 H422\" class=\"d-line\" \u002F>\u003Cpath d=\"M430 283 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Ctext x=\"385\" y=\"275\" text-anchor=\"middle\" class=\"d-label\">yes\u003C\u002Ftext>\u003Crect x=\"430\" y=\"250\" width=\"270\" height=\"66\" rx=\"10\" class=\"d-accent\" \u002F>\u003Ctext x=\"565\" y=\"279\" text-anchor=\"middle\" class=\"d-text\">Vector database\u003C\u002Ftext>\u003Ctext x=\"565\" y=\"300\" text-anchor=\"middle\" class=\"d-small\">Qdrant, Weaviate, Milvus, Pinecone\u003C\u002Ftext>\u003Cpath d=\"M180 316 V332\" class=\"d-line\" \u002F>\u003Cpath d=\"M180 340 l-5 -9 h10 z\" class=\"d-head\" \u002F>\u003Ctext x=\"190\" y=\"336\" class=\"d-label\">no\u003C\u002Ftext>\u003Crect x=\"20\" y=\"340\" width=\"680\" height=\"50\" rx=\"10\" class=\"d-gold\" \u002F>\u003Ctext x=\"360\" y=\"370\" text-anchor=\"middle\" class=\"d-text\">None of these apply: start with pgvector, measure, then move\u003C\u002Ftext>",[330],"Ask in order and stop at the first yes. The point is to avoid adding a system, not to avoid choosing one.",{"type":162,"content":332},[333],"Two caveats. The first \"yes\" does not end the conversation if the corpus is enormous: do the memory arithmetic in the scale section. And \"start with pgvector\" for the fallback is my bias for small teams, because it is the cheapest option to reverse: an export of vectors and metadata is all a migration needs.",{"type":177,"level":178,"id":140,"text":141},{"type":162,"content":336},[337],"This table compresses what I verified in the documentation. Versions are the latest releases I saw on GitHub at the start of October 2026: Qdrant 1.19.1, Weaviate 1.39.8, Milvus 3.0.2, OpenSearch 3.9.0 and pgvector 0.8.7.",{"type":339,"head":340,"rows":350},"table",[341,343,345,347,348],[342],"Option",[344],"Strongest at",[346],"Filtering",[12],[349],"Watch out for",[351,363,376,389,402,415,428],[352,355,357,359,361],[353],{"tag":169,"children":354},[9],[356],"Vectors next to relational data, one system",[358],"WHERE plus iterative scans, partial indexes, partitions",[360],"Full-text search plus your own fusion in SQL",[362],"Index memory, competes with OLTP, indexed dimension limits",[364,368,370,372,374],[365],{"tag":169,"children":366},[367],"Qdrant",[369],"Filter-heavy retrieval, flexible hybrid queries",[371],"Payload indexes, nested boolean conditions",[373],"Sparse and dense, RRF, DBSF, prefetch",[375],"A second system to sync and secure",[377,381,383,385,387],[378],{"tag":169,"children":379},[380],"Weaviate",[382],"Built-in hybrid search, multi-tenancy",[384],"Filtered vector search; I did not verify details, test yours",[386],"BM25 plus vector, relative score fusion by default",[388],"Pricing scales with vector dimensions in the cloud",[390,394,396,398,400],[391],{"tag":169,"children":392},[393],"Milvus",[395],"Very large, distributed workloads",[397],"Metadata filters; I did not verify details, test yours",[399],"Multiple dense and sparse vector fields",[401],"Distributed deployment on Kubernetes is real operational work",[403,407,409,411,413],[404],{"tag":169,"children":405},[406],"Pinecone",[408],"Managed, serverless, no servers to run",[410],"Metadata filters inside query executors",[412],"Sparse vectors, hybrid, BM25 full text",[414],"Serverless only in the docs I read, region fixed at creation",[416,420,422,424,426],[417],{"tag":169,"children":418},[419],"OpenSearch",[421],"One engine for keywords, facets and vectors",[423],"Filtering during Faiss or Lucene graph traversal",[425],"BM25 and vectors in one index",[427],"Cluster tuning, JVM and shard planning",[429,433,435,437,439],[430],{"tag":169,"children":431},[432],"Elasticsearch",[434],"Same, with Elastic tooling and BBQ quantization",[436],"Pre-filter during approximate kNN",[438],"BM25 and kNN in one index",[440],"Cluster sizing and operations",{"type":162,"content":442},[443],"A few details behind the cells. Weaviate offers HNSW, flat, dynamic and HFresh index types, where dynamic switches from flat to HNSW above a threshold (default 10,000 objects) and suits many small tenants. Milvus documents HNSW, IVF, DiskANN, ScaNN and GPU indexes, and a stateless, decoupled architecture. Elasticsearch documents that new indices with float vectors of 384 dimensions or more default to BBQ HNSW. OpenSearch supports Lucene (HNSW) and Faiss (HNSW and IVF) engines.",{"type":177,"level":178,"id":143,"text":144},{"type":162,"content":446},[447],"The one threshold I can give you without a benchmark is arithmetic. A float32 vector takes 4 bytes per dimension. At 1,536 dimensions that is 6,144 bytes, so 1 million vectors need about 6.1 GB and 10 million about 61 GB before graph links, metadata and replicas. halfvec halves it to roughly 3 GB and 31 GB, and binary quantization shrinks it far more at the cost of recall that you must measure.",{"type":162,"content":449},[450],"HNSW wants its graph in memory, so the practical question is whether that memory fits on the box you would rent anyway. My rule of thumb, and it is judgement and not measurement: up to a few million chunks, a well-sized Postgres instance is rarely the bottleneck. Somewhere in the tens of millions, or when index builds start hurting your primary, I begin comparing disk-based or quantized options in dedicated systems.",{"type":216,"ordered":217,"items":452},[453,458,463,468],[454,457],{"tag":169,"children":455},[456],"Self-hosted Postgres or managed Postgres:"," cost is the instance you already pay for plus extra RAM. No new vendor.",[459,462],{"tag":169,"children":460},[461],"Search engine cluster:"," you pay per node for memory and disk, but share it with keyword search and logs.",[464,467],{"tag":169,"children":465},[466],"Dedicated cloud service:"," you pay for capacity or usage. Weaviate bills mainly by vector dimensions, so lower-dimension or compressed embeddings cut the bill directly.",[469,472],{"tag":169,"children":470},[471],"Hidden cost:"," a second system means a second sync pipeline, a second access model and a second incident channel.",{"type":162,"content":474},[475,476,307],"Cost also depends on the embeddings you choose. A smaller model, or a model that supports shortened vectors, reduces storage everywhere. I cover the wider levers in ",{"tag":201,"to":477,"children":478},"\u002Fblog\u002Fllm-cost-latency-prompt-caching-routing",[479],"LLM cost, latency, prompt caching and routing",{"type":177,"level":178,"id":146,"text":147},{"type":162,"content":482},[483],"For European clients, \"where do the vectors live\" is a procurement question. Embeddings are derived from your documents and can leak information about them, so treat them as personal data when the source is.",{"type":216,"ordered":217,"items":485},[486,491,496,501,506],[487,490],{"tag":169,"children":488},[489],"Postgres:"," any EU region of a managed provider, or your own servers. Simplest to audit.",[492,495],{"tag":169,"children":493},[494],"Pinecone:"," documents AWS eu-west-1 (Ireland) and eu-central-1 (Frankfurt) and GCP europe-west4 (Netherlands) on Builder plans and above. The Starter plan is limited to AWS us-east-1, and the region cannot be changed after creation.",[497,500],{"tag":169,"children":498},[499],"Weaviate Cloud:"," EU regions on shared and dedicated deployments, with a bring-your-own-cloud option listed as coming soon.",[502,505],{"tag":169,"children":503},[504],"Qdrant:"," managed clusters on AWS, GCP and Azure, plus Hybrid Cloud, a self-managed option.",[507,510],{"tag":169,"children":508},[509],"Self-hosted Milvus, Qdrant, Weaviate, OpenSearch:"," you decide the data centre, and you carry the operations.",{"type":162,"content":512},[513,514,518],"A region in Frankfurt is necessary but not always sufficient: a US-headquartered provider can still raise questions about foreign access, and the embedding model call is a second data flow. My GDPR notes in ",{"tag":201,"to":515,"children":516},"\u002Fblog\u002Fgdpr-llm-api-eu-data-residency",[517],"GDPR and LLM APIs: EU data residency"," cover that part. Check the current region list on the vendor's page at signing time, since these change.",{"type":177,"level":178,"id":149,"text":150},{"type":162,"content":521},[522],"This is the cost that benchmarks never show. Ask four questions of any option.",{"type":216,"ordered":217,"items":524},[525,530,535,540],[526,529],{"tag":169,"children":527},[528],"Backups and restore:"," can you restore vectors and metadata to a consistent point? Postgres has this solved. For others, test the restore, not only the backup.",[531,534],{"tag":169,"children":532},[533],"Re-indexing:"," changing the embedding model means re-embedding the corpus. Can you run the new index beside the old one and switch over?",[536,539],{"tag":169,"children":537},[538],"Upgrades:"," pgvector patches such as 0.8.3 fixed index corruption cases. Who watches the release notes?",[541,544],{"tag":169,"children":542},[543],"On-call:"," a distributed system on Kubernetes is a different commitment from an extension in a database you already run.",{"type":162,"content":546},[547],"In my experience the cheapest operations come from running one fewer system. That is the whole argument for pgvector and for search engines, and it is why a managed dedicated service is the right answer when your team has no appetite for running any of them.",{"type":177,"level":178,"id":152,"text":153},{"type":216,"ordered":550,"items":551},true,[552,554,556,558,560,562,564],[553],"Write down your top five filters, including tenant and permission checks.",[555],"Count chunks, dimensions and growth, then do the memory arithmetic for float32 and halfvec.",[557],"Build a labelled evaluation set of at least a few dozen real questions.",[559],"Run the same queries against your top two candidates with your real filters and compare recall and latency.",[561],"Decide who owns fusion and relevance tuning for hybrid search.",[563],"Confirm EU region, backup, restore and upgrade procedure in writing.",[565],"Plan the re-embedding path before you load the first million vectors.",{"type":162,"content":567},[568],"If two options tie, pick the one with fewer moving parts. You can always move later: vectors and metadata export cleanly.",{"type":177,"level":178,"id":155,"text":156},{"type":162,"content":571},[572],"For a typical European B2B project with a catalogue or knowledge base in the low millions of chunks, I would start with pgvector, halfvec, an HNSW index and iterative scans, add Postgres full-text search for hybrid retrieval, and write the evaluation set on day one. If the client already runs OpenSearch, I would put the vectors there.",{"type":162,"content":574},[575,576,307],"I would move to a dedicated vector database when a measured limit appears: filtered recall that iterative scans cannot rescue, index builds that hurt production, or tenancy and scale that Postgres should not carry. And I would choose it for its operating model, managed or self-hosted in the EU, as much as for its features. If you want help making that call on your own data, see ",{"tag":201,"to":577,"children":578},"\u002Fexpertise\u002Fai-engineer",[579],"my AI engineering work",{"type":177,"level":178,"id":158,"text":159},{"type":216,"ordered":550,"items":582},[583,586,589,592,595,598,601,604,607,610,613,616,619,622,625,628,631],[584],{"tag":211,"href":38,"children":585},[37],[587],{"tag":211,"href":41,"children":588},[40],[590],{"tag":211,"href":44,"children":591},[43],[593],{"tag":211,"href":47,"children":594},[46],[596],{"tag":211,"href":50,"children":597},[49],[599],{"tag":211,"href":53,"children":600},[52],[602],{"tag":211,"href":56,"children":603},[55],[605],{"tag":211,"href":59,"children":606},[58],[608],{"tag":211,"href":62,"children":609},[61],[611],{"tag":211,"href":65,"children":612},[64],[614],{"tag":211,"href":68,"children":615},[67],[617],{"tag":211,"href":71,"children":618},[70],[620],{"tag":211,"href":74,"children":621},[73],[623],{"tag":211,"href":77,"children":624},[76],[626],{"tag":211,"href":80,"children":627},[79],[629],{"tag":211,"href":83,"children":630},[82],[632],{"tag":211,"href":86,"children":633},[85],[635,707,774,841],{"slug":636,"published":5,"minutes":6,"category":7,"tags":637,"keywords":642,"about":653,"sources":661,"cover":701,"og":702,"expertise":89,"locales":703,"lang":91,"title":704,"description":705,"coverAlt":706},"llm-hallucination-grounding-citations",[638,11,639,640,641],"Hallucinations","Citations","Grounding","Faithfulness",[643,644,645,646,647,648,649,650,651,652],"reduce LLM hallucinations in production","how to reduce hallucinations in RAG","LLM citations API","Anthropic citations API","RAG faithfulness metric","LLM abstention I don't know","claim-level verification LLM","grounding LLM answers in sources","check grounding API","show sources in AI chatbot UI",[654,657,658],{"name":655,"url":656},"Hallucination (artificial intelligence)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FHallucination_(artificial_intelligence)",{"name":33,"url":34},{"name":659,"url":660},"Large language model","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FLarge_language_model",[662,665,668,671,674,677,680,683,686,689,692,695,698],{"title":663,"url":664},"Anthropic: Citations (Claude API documentation)","https:\u002F\u002Fplatform.claude.com\u002Fdocs\u002Fen\u002Fbuild-with-claude\u002Fcitations",{"title":666,"url":667},"Anthropic: Search results (Claude API documentation)","https:\u002F\u002Fplatform.claude.com\u002Fdocs\u002Fen\u002Fbuild-with-claude\u002Fsearch-results",{"title":669,"url":670},"Anthropic: Reduce hallucinations (Claude API documentation)","https:\u002F\u002Fplatform.claude.com\u002Fdocs\u002Fen\u002Ftest-and-evaluate\u002Fstrengthen-guardrails\u002Freduce-hallucinations",{"title":672,"url":673},"Anthropic: Introducing Citations on the Anthropic API","https:\u002F\u002Fclaude.com\u002Fblog\u002Fintroducing-citations-api",{"title":675,"url":676},"Simon Willison: Anthropic's new Citations API (24 January 2025)","https:\u002F\u002Fsimonwillison.net\u002F2025\u002FJan\u002F24\u002Fanthropics-new-citations-api\u002F",{"title":678,"url":679},"OpenAI: Web search guide (url_citation annotations and display requirement)","https:\u002F\u002Fdevelopers.openai.com\u002Fapi\u002Fdocs\u002Fguides\u002Ftools-web-search",{"title":681,"url":682},"Cohere: Documents and citations","https:\u002F\u002Fdocs.cohere.com\u002Fdocs\u002Fdocuments-and-citations",{"title":684,"url":685},"Google Cloud: Check grounding API","https:\u002F\u002Fdocs.cloud.google.com\u002Fgenerative-ai-app-builder\u002Fdocs\u002Fcheck-grounding",{"title":687,"url":688},"AWS: Amazon Bedrock Guardrails contextual grounding check","https:\u002F\u002Fdocs.aws.amazon.com\u002Fbedrock\u002Flatest\u002Fuserguide\u002Fguardrails-contextual-grounding-check.html",{"title":690,"url":691},"Ragas: Faithfulness metric","https:\u002F\u002Fdocs.ragas.io\u002Fen\u002Fstable\u002Fconcepts\u002Fmetrics\u002Favailable_metrics\u002Ffaithfulness\u002F",{"title":693,"url":694},"Kalai, Nachum, Vempala, Zhang: Why Language Models Hallucinate (arXiv 2509.04664)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2509.04664",{"title":696,"url":697},"Magesh et al.: Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools (arXiv 2405.20362)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2405.20362",{"title":699,"url":700},"Wallat, Heuss, de Rijke, Anand: Correctness is not Faithfulness in RAG Attributions (arXiv 2412.18004)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2412.18004","\u002Fimages\u002Fblog\u002Fllm-hallucination-grounding-citations\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fllm-hallucination-grounding-citations\u002Fog.jpg",[91,92,93],"Reducing LLM hallucinations in production: grounding, citations and knowing when to say no","Cut hallucinations in production RAG: citation APIs, abstention, claim-level checks, faithfulness metrics, source UI, and the failures that still slip through.","Diagram: a retrieval step feeds an evidence gate, a cited answer and a claim verifier, ending in an answer with sources, with abstain and flag paths branching off.",{"slug":708,"published":5,"minutes":6,"category":7,"tags":709,"keywords":713,"about":723,"sources":731,"cover":768,"og":769,"expertise":89,"locales":770,"lang":91,"title":771,"description":772,"coverAlt":773},"graphrag-knowledge-graph-rag",[710,711,712,11],"GraphRAG","Knowledge graphs","LightRAG",[710,714,715,716,717,718,719,720,721,722],"knowledge graph RAG","GraphRAG vs vector RAG","Microsoft GraphRAG explained","LightRAG vs GraphRAG","GraphRAG global vs local search","GraphRAG indexing cost","when to use GraphRAG","multi-hop RAG","LazyGraphRAG",[724,727,728],{"name":725,"url":726},"Knowledge graph","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FKnowledge_graph",{"name":33,"url":34},{"name":729,"url":730},"Leiden algorithm","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FLeiden_algorithm",[732,735,738,741,744,747,750,753,756,759,762,765],{"title":733,"url":734},"Edge et al.: From Local to Global: A Graph RAG Approach to Query-Focused Summarization (arXiv:2404.16130)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2404.16130",{"title":736,"url":737},"Microsoft GraphRAG documentation: overview","https:\u002F\u002Fmicrosoft.github.io\u002Fgraphrag\u002F",{"title":739,"url":740},"Microsoft GraphRAG documentation: default dataflow","https:\u002F\u002Fmicrosoft.github.io\u002Fgraphrag\u002Findex\u002Fdefault_dataflow\u002F",{"title":742,"url":743},"Microsoft GraphRAG documentation: global search","https:\u002F\u002Fmicrosoft.github.io\u002Fgraphrag\u002Fquery\u002Fglobal_search\u002F",{"title":745,"url":746},"Microsoft GraphRAG documentation: local search","https:\u002F\u002Fmicrosoft.github.io\u002Fgraphrag\u002Fquery\u002Flocal_search\u002F",{"title":748,"url":749},"Microsoft GraphRAG documentation: DRIFT search","https:\u002F\u002Fmicrosoft.github.io\u002Fgraphrag\u002Fquery\u002Fdrift_search\u002F",{"title":751,"url":752},"Microsoft GraphRAG documentation: getting started","https:\u002F\u002Fmicrosoft.github.io\u002Fgraphrag\u002Fget_started\u002F",{"title":754,"url":755},"Microsoft Research: LazyGraphRAG, setting a new standard for quality and cost (25 November 2024)","https:\u002F\u002Fwww.microsoft.com\u002Fen-us\u002Fresearch\u002Fblog\u002Flazygraphrag-setting-a-new-standard-for-quality-and-cost\u002F",{"title":757,"url":758},"Guo et al.: LightRAG: Simple and Fast Retrieval-Augmented Generation (arXiv:2410.05779)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2410.05779",{"title":760,"url":761},"HKUDS\u002FLightRAG on GitHub","https:\u002F\u002Fgithub.com\u002FHKUDS\u002FLightRAG",{"title":763,"url":764},"Gutiérrez et al.: HippoRAG: Neurobiologically Inspired Long-Term Memory for Large Language Models (arXiv:2405.14831)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2405.14831",{"title":766,"url":767},"Xiang et al.: When to use Graphs in RAG: A Comprehensive Analysis for Graph Retrieval-Augmented Generation (arXiv:2506.05690)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2506.05690","\u002Fimages\u002Fblog\u002Fgraphrag-knowledge-graph-rag\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fgraphrag-knowledge-graph-rag\u002Fog.jpg",[91,92,93],"GraphRAG and knowledge-graph RAG: when a graph beats vector search","What Microsoft GraphRAG and LightRAG really do, what indexing costs, and when a knowledge graph beats vector RAG: multi-hop, global questions, product catalogues.","Diagram: a knowledge graph hub linked to entities, communities, local search, global search and product parts.",{"slug":775,"published":5,"minutes":776,"category":7,"tags":777,"keywords":783,"about":793,"sources":801,"cover":835,"og":836,"expertise":89,"locales":837,"lang":91,"title":838,"description":839,"coverAlt":840},"rag-evaluation-metrics",12,[778,779,780,781,782],"RAG evaluation","Retrieval metrics","LLM-as-judge","Golden set","Ragas",[778,784,785,786,787,788,789,790,791,792],"how to evaluate RAG","RAG evaluation metrics","recall@k MRR nDCG","faithfulness vs answer relevance","golden dataset for RAG","LLM as a judge calibration","Ragas vs DeepEval vs TruLens","RAG evals in CI","retrieval vs generation failure",[794,795,798],{"name":33,"url":34},{"name":796,"url":797},"Discounted cumulative gain","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FDiscounted_cumulative_gain",{"name":799,"url":800},"Mean reciprocal rank","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FMean_reciprocal_rank",[802,805,808,811,813,816,819,822,825,828,830,832],{"title":803,"url":804},"Es et al.: RAGAS, Automated Evaluation of Retrieval Augmented Generation (arXiv 2309.15217)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2309.15217",{"title":806,"url":807},"Zheng et al.: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (arXiv 2306.05685)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2306.05685",{"title":809,"url":810},"Ragas documentation: available metrics","https:\u002F\u002Fdocs.ragas.io\u002Fen\u002Fstable\u002Fconcepts\u002Fmetrics\u002Favailable_metrics\u002F",{"title":812,"url":691},"Ragas documentation: faithfulness",{"title":814,"url":815},"Ragas documentation: context precision","https:\u002F\u002Fdocs.ragas.io\u002Fen\u002Fstable\u002Fconcepts\u002Fmetrics\u002Favailable_metrics\u002Fcontext_precision\u002F",{"title":817,"url":818},"Ragas documentation: context recall","https:\u002F\u002Fdocs.ragas.io\u002Fen\u002Fstable\u002Fconcepts\u002Fmetrics\u002Favailable_metrics\u002Fcontext_recall\u002F",{"title":820,"url":821},"DeepEval documentation: metrics introduction","https:\u002F\u002Fdeepeval.com\u002Fdocs\u002Fmetrics-introduction",{"title":823,"url":824},"TruLens","https:\u002F\u002Fwww.trulens.org\u002F",{"title":826,"url":827},"Arize Phoenix documentation","https:\u002F\u002Farize.com\u002Fdocs\u002Fphoenix",{"title":829,"url":797},"Wikipedia: Discounted cumulative gain",{"title":831,"url":800},"Wikipedia: Mean reciprocal rank",{"title":833,"url":834},"Wikipedia: Cohen's kappa","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FCohen%27s_kappa","\u002Fimages\u002Fblog\u002Frag-evaluation-metrics\u002Fcover.webp","\u002Fimages\u002Fblog\u002Frag-evaluation-metrics\u002Fog.jpg",[91,92,93],"Evaluating RAG: retrieval metrics, faithfulness and how to tell which half failed","How to evaluate a RAG system: recall at k, MRR and nDCG vs faithfulness and answer relevance, a golden set from real queries, a calibrated LLM judge and evals in CI.","Diagram: a RAG answer is scored on two sides, retrieval metrics such as recall at k, MRR and nDCG, and generation metrics such as faithfulness and answer relevance, feeding a diagnosis.",{"slug":842,"published":5,"minutes":6,"category":7,"tags":843,"keywords":847,"about":858,"sources":865,"cover":904,"og":905,"expertise":906,"locales":907,"lang":91,"title":908,"description":909,"coverAlt":910},"semantic-product-search-b2b",[844,12,845,846,419],"B2B search","Semantic search","Spryker",[848,849,850,851,852,853,854,855,856,857],"semantic product search B2B","B2B ecommerce search","hybrid search BM25 vector","part number search ecommerce","Spryker search Elasticsearch","OpenSearch hybrid search RRF","multilingual product search German English Hungarian","zero results rate site search","LLM query understanding ecommerce","AI product search for B2B shops",[859,861,863],{"name":845,"url":860},"https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FSemantic_search",{"name":432,"url":862},"https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FElasticsearch",{"name":419,"url":864},"https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FOpenSearch",[866,869,872,874,877,880,883,886,889,892,895,898,901],{"title":867,"url":868},"Baymard Institute: E-commerce search query types","https:\u002F\u002Fbaymard.com\u002Fblog\u002Fecommerce-search-query-types",{"title":870,"url":871},"Elastic Search Labs: Hybrid search in Elasticsearch","https:\u002F\u002Fwww.elastic.co\u002Fsearch-labs\u002Fblog\u002Fhybrid-search-elasticsearch",{"title":873,"url":83},"Elasticsearch documentation: kNN query (pre-filters and post-filters)",{"title":875,"url":876},"Elasticsearch documentation: Semantic reranking","https:\u002F\u002Fwww.elastic.co\u002Fdocs\u002Fsolutions\u002Fsearch\u002Franking\u002Fsemantic-reranking",{"title":878,"url":879},"Elasticsearch documentation: Word delimiter graph token filter","https:\u002F\u002Fwww.elastic.co\u002Fdocs\u002Freference\u002Ftext-analysis\u002Fanalysis-word-delimiter-graph-tokenfilter",{"title":881,"url":882},"Elasticsearch documentation: Synonym graph token filter","https:\u002F\u002Fwww.elastic.co\u002Fdocs\u002Freference\u002Ftext-analysis\u002Fanalysis-synonym-graph-tokenfilter",{"title":884,"url":885},"OpenSearch documentation: Score ranker processor (RRF)","https:\u002F\u002Fdocs.opensearch.org\u002Flatest\u002Fsearch-plugins\u002Fsearch-pipelines\u002Fscore-ranker-processor\u002F",{"title":887,"url":888},"OpenSearch documentation: Normalization processor","https:\u002F\u002Fdocs.opensearch.org\u002Flatest\u002Fsearch-plugins\u002Fsearch-pipelines\u002Fnormalization-processor\u002F",{"title":890,"url":891},"Spryker documentation: Search feature overview","https:\u002F\u002Fdocs.spryker.com\u002Fdocs\u002Fpbc\u002Fall\u002Fsearch\u002Flatest\u002Fbase-shop\u002Fsearch-feature-overview\u002Fsearch-feature-overview",{"title":893,"url":894},"Spryker documentation: Migrate from OpenSearch 1.3 to 3.5","https:\u002F\u002Fdocs.spryker.com\u002Fdocs\u002Fpbc\u002Fall\u002Fsearch\u002Flatest\u002Fbase-shop\u002Finstall-and-upgrade\u002Fmigrate-from-opensearch-1.3-to-3.5.html",{"title":896,"url":897},"Instacart via ZenML: Rebuilding query understanding for e-commerce search with LLMs","https:\u002F\u002Fwww.zenml.io\u002Fllmops-database\u002Frebuilding-query-understanding-for-e-commerce-search-with-llms",{"title":899,"url":900},"arXiv: M3-Embedding, multilingual, multi-functionality, multi-granularity text embeddings","https:\u002F\u002Farxiv.org\u002Fabs\u002F2402.03216",{"title":902,"url":903},"Algolia documentation: Search analytics metrics","https:\u002F\u002Fwww.algolia.com\u002Fdoc\u002Fguides\u002Fsearch-analytics\u002Fconcepts\u002Fmetrics\u002F","\u002Fimages\u002Fblog\u002Fsemantic-product-search-b2b\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fsemantic-product-search-b2b\u002Fog.jpg","b2b-ecommerce-developer",[91,92,93],"Semantic product search for B2B shops: part numbers, hybrid retrieval and what to measure","How to add semantic search to a B2B shop without breaking part-number search: hybrid BM25 and vectors, filters, DE\u002FEN\u002FHU, LLM query parsing, reranking and metrics.","Diagram: a search query is split into an identifier lane, lexical BM25 and vector kNN, fused, reranked and returned as results.",1791009037094]