[{"data":1,"prerenderedAt":894},["ShallowReactive",2],{"blog-semantic-product-search-b2b-en":3},{"slug":4,"published":5,"minutes":6,"category":7,"tags":8,"keywords":14,"about":25,"sources":33,"cover":73,"og":74,"expertise":75,"locales":76,"lang":77,"title":80,"description":81,"coverAlt":82,"metaTitle":83,"takeaways":84,"faq":90,"toc":109,"blocks":149,"others":601},"semantic-product-search-b2b","2026-10-02",13,"rag",[9,10,11,12,13],"B2B search","Hybrid search","Semantic search","Spryker","OpenSearch",[15,16,17,18,19,20,21,22,23,24],"semantic product search B2B","B2B ecommerce search","hybrid search BM25 vector","part number search ecommerce","Spryker search Elasticsearch","OpenSearch hybrid search RRF","multilingual product search German English Hungarian","zero results rate site search","LLM query understanding ecommerce","AI product search for B2B shops",[26,28,31],{"name":11,"url":27},"https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FSemantic_search",{"name":29,"url":30},"Elasticsearch","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FElasticsearch",{"name":13,"url":32},"https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FOpenSearch",[34,37,40,43,46,49,52,55,58,61,64,67,70],{"title":35,"url":36},"Baymard Institute: E-commerce search query types","https:\u002F\u002Fbaymard.com\u002Fblog\u002Fecommerce-search-query-types",{"title":38,"url":39},"Elastic Search Labs: Hybrid search in Elasticsearch","https:\u002F\u002Fwww.elastic.co\u002Fsearch-labs\u002Fblog\u002Fhybrid-search-elasticsearch",{"title":41,"url":42},"Elasticsearch documentation: kNN query (pre-filters and post-filters)","https:\u002F\u002Fwww.elastic.co\u002Fdocs\u002Freference\u002Fquery-languages\u002Fquery-dsl\u002Fquery-dsl-knn-query",{"title":44,"url":45},"Elasticsearch documentation: Semantic reranking","https:\u002F\u002Fwww.elastic.co\u002Fdocs\u002Fsolutions\u002Fsearch\u002Franking\u002Fsemantic-reranking",{"title":47,"url":48},"Elasticsearch documentation: Word delimiter graph token filter","https:\u002F\u002Fwww.elastic.co\u002Fdocs\u002Freference\u002Ftext-analysis\u002Fanalysis-word-delimiter-graph-tokenfilter",{"title":50,"url":51},"Elasticsearch documentation: Synonym graph token filter","https:\u002F\u002Fwww.elastic.co\u002Fdocs\u002Freference\u002Ftext-analysis\u002Fanalysis-synonym-graph-tokenfilter",{"title":53,"url":54},"OpenSearch documentation: Score ranker processor (RRF)","https:\u002F\u002Fdocs.opensearch.org\u002Flatest\u002Fsearch-plugins\u002Fsearch-pipelines\u002Fscore-ranker-processor\u002F",{"title":56,"url":57},"OpenSearch documentation: Normalization processor","https:\u002F\u002Fdocs.opensearch.org\u002Flatest\u002Fsearch-plugins\u002Fsearch-pipelines\u002Fnormalization-processor\u002F",{"title":59,"url":60},"Spryker documentation: Search feature overview","https:\u002F\u002Fdocs.spryker.com\u002Fdocs\u002Fpbc\u002Fall\u002Fsearch\u002Flatest\u002Fbase-shop\u002Fsearch-feature-overview\u002Fsearch-feature-overview",{"title":62,"url":63},"Spryker documentation: Migrate from OpenSearch 1.3 to 3.5","https:\u002F\u002Fdocs.spryker.com\u002Fdocs\u002Fpbc\u002Fall\u002Fsearch\u002Flatest\u002Fbase-shop\u002Finstall-and-upgrade\u002Fmigrate-from-opensearch-1.3-to-3.5.html",{"title":65,"url":66},"Instacart via ZenML: Rebuilding query understanding for e-commerce search with LLMs","https:\u002F\u002Fwww.zenml.io\u002Fllmops-database\u002Frebuilding-query-understanding-for-e-commerce-search-with-llms",{"title":68,"url":69},"arXiv: M3-Embedding, multilingual, multi-functionality, multi-granularity text embeddings","https:\u002F\u002Farxiv.org\u002Fabs\u002F2402.03216",{"title":71,"url":72},"Algolia documentation: Search analytics metrics","https:\u002F\u002Fwww.algolia.com\u002Fdoc\u002Fguides\u002Fsearch-analytics\u002Fconcepts\u002Fmetrics\u002F","\u002Fimages\u002Fblog\u002Fsemantic-product-search-b2b\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fsemantic-product-search-b2b\u002Fog.jpg","b2b-ecommerce-developer",[77,78,79],"en","de","hu","Semantic product search for B2B shops: part numbers, hybrid retrieval and what to measure","How to add semantic search to a B2B shop without breaking part-number search: hybrid BM25 and vectors, filters, DE\u002FEN\u002FHU, LLM query parsing, reranking and metrics.","Diagram: a search query is split into an identifier lane, lexical BM25 and vector kNN, fused, reranked and returned as results.","B2B product search: hybrid and semantic · Balázs Csorba",[85,86,87,88,89],"In B2B, the part number is the most important query. Give identifiers their own lane with normalised exact and prefix matching, and let semantic retrieval help only when that lane has no strong answer.","Hybrid search (BM25 plus vectors, fused with reciprocal rank fusion) beats either method alone, because buyers type both \"10-32-4711\" and \"screw for outdoor wood\" into the same box.","Assortments, price lists and stock must be pre-filters inside the vector search, not post-filters, otherwise results disappear or leak across customers.","Use an LLM to parse queries into product type, attributes and units, validate its output against the catalogue, and cache it offline for frequent queries instead of calling it on every keystroke.","Judge the system by query type: zero-result rate, search-exit rate and click-through, plus a small judged query set where exact identifiers must always rank first.",[91,94,97,100,103,106],{"q":92,"a":93},"What is semantic search for B2B e-commerce?","Semantic search turns product texts and queries into vectors so that a request such as \"screw for outdoor wood\" finds matching products even without shared words. In a B2B shop it should complement, not replace, lexical search, because article numbers, EANs and exact specifications still need precise matching.",{"q":95,"a":96},"What is hybrid search and why do B2B shops need it?","Hybrid search runs a lexical query (BM25) and a vector query in parallel and fuses the two ranked lists, often with reciprocal rank fusion. B2B shops need it because their buyers mix exact identifiers with natural-language descriptions, and neither method alone handles both well.",{"q":98,"a":99},"How do I keep part-number search exact with vector search?","Index identifiers in dedicated fields with a normaliser that strips case, dashes, dots and spaces, match them exactly and by prefix first, and rank those hits above anything the vector arm returns. Do not rely on embeddings for identifiers, because they treat similar-looking numbers as similar.",{"q":101,"a":102},"Does semantic search work for German, English and Hungarian catalogues?","Yes, with a multilingual embedding model and per-language text analysis. Multilingual models such as BGE-M3 cover more than 100 languages, but you still need to test German compound words and Hungarian inflection on your own queries and measure results per language.",{"q":104,"a":105},"How do I measure whether product search improved?","Segment by query type and track zero-result rate, search-exit rate, click-through rate and add-to-cart from search. Add an offline set of judged queries, where every exact identifier must rank first. A falling zero-result rate alone can hide irrelevant results, so always read it next to click-through.",{"q":107,"a":108},"Can I add semantic search to Spryker?","Yes. Spryker ships with Elasticsearch as its default search and documents an OpenSearch upgrade path, so you can add vector fields to the product search documents at publish time and extend the search query with a vector clause. I would add this as a pilot behind a feature flag and compare it with the existing search.",[110,113,116,119,122,125,128,131,134,137,140,143,146],{"id":111,"title":112},"why-b2b-search-differs","Why B2B search is not consumer search",{"id":114,"title":115},"architecture","The architecture: lanes, fusion, rerank",{"id":117,"title":118},"exact-match-first","Article numbers and exact match come first",{"id":120,"title":121},"hybrid-retrieval","Hybrid retrieval: BM25 plus vectors, fused by rank",{"id":123,"title":124},"filters-and-attributes","Attribute-aware filtering, and why numbers need structure",{"id":126,"title":127},"multilingual","German, English and Hungarian in one catalogue",{"id":129,"title":130},"synonyms","Synonyms and part numbers still matter",{"id":132,"title":133},"llm-query-understanding","Query understanding with an LLM",{"id":135,"title":136},"reranking","Reranking the top of the list",{"id":138,"title":139},"evaluation","Evaluation: zero results, exits and a judged set",{"id":141,"title":142},"spryker-integration","Integrating with Spryker, Elasticsearch and OpenSearch",{"id":144,"title":145},"rollout-checklist","A rollout checklist",{"id":147,"title":148},"sources","Sources",[150,154,157,160,168,171,254,257,258,261,270,273,276,277,280,283,308,315,316,323,334,337,344,347,348,355,358,361,362,365,372,375,376,383,386,387,401,408,435,438,439,446,454,455,458,506,509,512,513,524,527,530,533,534,537,556,559,560],{"type":151,"content":152},"paragraph",[153],"A buyer at a plumbing wholesaler types \"4711-32\". A second one types \"Edelstahlschraube für Holz außen\". A third types \"hex bolt M8x40 A2\". All three use the same search box, and in most B2B shops I have seen, at least one of them gets a page of nothing.",{"type":151,"content":155},[156],"Semantic search promises to fix the second and third case. The risk is that it breaks the first. This article is how I would add semantic retrieval to a B2B catalogue without losing exact part-number search: the architecture, the query types, hybrid fusion, filters, three languages, synonyms, LLM query understanding, reranking, measurement and what it means for a Spryker shop on Elasticsearch or OpenSearch.",{"type":158,"level":159,"id":111,"text":112},"heading",2,{"type":151,"content":161},[162,163,167],"B2B queries are more heterogeneous than consumer queries. Buyers paste an article number from a drawing, retype a supplier number from an old order, abbreviate trade jargon, or describe a use case. Baymard, which benchmarks consumer shops, distinguishes ",{"tag":164,"href":36,"children":165},"a",[166],"eight kinds of search query"," and found that 56% of the sites it tested do not adequately support users' search needs. B2B catalogues add identifier-heavy, specification-heavy data on top.",{"type":151,"content":169},[170],"The practical consequence is that no single retrieval method is right. Here is the taxonomy I use when I start a project. The examples are invented, but the patterns are the ones that show up in query logs.",{"type":172,"head":173,"rows":182},"table",[174,176,178,180],[175],"Query type",[177],"Example",[179],"Best retrieval",[181],"Typical failure",[183,192,201,210,219,228,237,245],[184,186,188,190],[185],"Article or part number",[187],"\"4711-32\", \"4711 32\"",[189],"Identifier lane: normalised exact, then prefix",[191],"Dashes and spaces break exact match; a vector arm returns look-alikes",[193,195,197,199],[194],"Supplier number, EAN",[196],"\"4006381333931\"",[198],"Identifier lane on its own field",[200],"Stored as a number, leading zeros lost",[202,204,206,208],[203],"Product type",[205],"\"Sechskantschraube\"",[207],"Lexical plus vector, category boost",[209],"Compounds and plurals miss in lexical search",[211,213,215,217],[212],"Specification",[214],"\"M8x40 A2 DIN 933\"",[216],"Parsed into filters plus lexical",[218],"Embeddings blur numbers and units",[220,222,224,226],[221],"Use case",[223],"\"screw for outdoor wood\"",[225],"Vector plus attribute filters",[227],"Lexical search returns zero results",[229,231,233,235],[230],"Abbreviation or jargon",[232],"\"VA Schraube\"",[234],"Synonyms, then vector",[236],"Abbreviation unknown to the analyzer",[238,240,242,244],[239],"Cross-language",[241],"\"hex bolt\" in a German catalogue",[243],"Multilingual vector plus synonyms",[227],[246,248,250,252],[247],"Non-product",[249],"\"delivery time\", \"datasheet\"",[251],"Route to help or CMS content",[253],"Product index returns random products",{"type":151,"content":255},[256],"Count how many of your real queries fall into each row before you choose anything. In a catalogue of technical parts, the first two rows can be a large share, and they are exactly the rows where semantic search adds nothing and can do harm.",{"type":158,"level":159,"id":114,"text":115},{"type":151,"content":259},[260],"My reference design has one entry point and three retrieval lanes that run in parallel. A cheap understanding step normalises the query and detects whether it looks like an identifier. The identifier lane, a lexical BM25 query and a vector kNN query all run under the same filters. Their ranked lists are fused, a reranker reorders only the top of the list, and the response carries the facets the shop needs.",{"type":262,"attrs":263,"inner":267,"caption":268},"diagram",{"viewBox":264,"role":265,"aria-labelledby":266},"0 0 720 436","img","d1-sps-t d1-sps-d","\u003Ctitle id=\"d1-sps-t\">Hybrid product search in lanes\u003C\u002Ftitle>\u003Cdesc id=\"d1-sps-d\">A query goes through an understanding step into three parallel lanes: identifier, lexical BM25 and vector kNN, all under shared filters. Results are fused with RRF, reranked and returned as results with facets. Search logs feed back into synonyms and a judged query set.\u003C\u002Fdesc>\u003Ctext x=\"20\" y=\"26\" class=\"d-title\">Hybrid product search in lanes\u003C\u002Ftext>\u003Ctext x=\"700\" y=\"26\" text-anchor=\"end\" class=\"d-label\">architecture sketch\u003C\u002Ftext>\u003Crect x=\"20\" y=\"120\" width=\"100\" height=\"56\" rx=\"10\" class=\"d-sky\" \u002F>\u003Ctext x=\"70\" y=\"144\" text-anchor=\"middle\" class=\"d-text\">Query\u003C\u002Ftext>\u003Ctext x=\"70\" y=\"163\" text-anchor=\"middle\" class=\"d-small\">typed text\u003C\u002Ftext>\u003Cpath d=\"M120 148 H132\" class=\"d-line\" \u002F>\u003Cpath d=\"M140 148 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Crect x=\"140\" y=\"120\" width=\"140\" height=\"56\" rx=\"10\" class=\"d-accent\" \u002F>\u003Ctext x=\"210\" y=\"144\" text-anchor=\"middle\" class=\"d-text\">Understand\u003C\u002Ftext>\u003Ctext x=\"210\" y=\"163\" text-anchor=\"middle\" class=\"d-small\">normalise, parse\u003C\u002Ftext>\u003Cpath d=\"M280 148 C298 148 292 68 312 68\" class=\"d-line\" \u002F>\u003Cpath d=\"M320 68 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Crect x=\"320\" y=\"40\" width=\"180\" height=\"56\" rx=\"10\" class=\"d-mint\" \u002F>\u003Ctext x=\"410\" y=\"64\" text-anchor=\"middle\" class=\"d-text\">Identifier lane\u003C\u002Ftext>\u003Ctext x=\"410\" y=\"83\" text-anchor=\"middle\" class=\"d-small\">exact, then prefix\u003C\u002Ftext>\u003Cpath d=\"M500 68 C515 68 512 148 528 148\" class=\"d-line\" \u002F>\u003Cpath d=\"M280 148 C298 148 292 148 312 148\" class=\"d-line\" \u002F>\u003Cpath d=\"M320 148 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Crect x=\"320\" y=\"120\" width=\"180\" height=\"56\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"410\" y=\"144\" text-anchor=\"middle\" class=\"d-text\">Lexical BM25\u003C\u002Ftext>\u003Ctext x=\"410\" y=\"163\" text-anchor=\"middle\" class=\"d-small\">text, synonyms\u003C\u002Ftext>\u003Cpath d=\"M500 148 C515 148 512 148 528 148\" class=\"d-line\" \u002F>\u003Cpath d=\"M280 148 C298 148 292 228 312 228\" class=\"d-line\" \u002F>\u003Cpath d=\"M320 228 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Crect x=\"320\" y=\"200\" width=\"180\" height=\"56\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"410\" y=\"224\" text-anchor=\"middle\" class=\"d-text\">Vector kNN\u003C\u002Ftext>\u003Ctext x=\"410\" y=\"243\" text-anchor=\"middle\" class=\"d-small\">multilingual\u003C\u002Ftext>\u003Cpath d=\"M500 228 C515 228 512 148 528 148\" class=\"d-line\" \u002F>\u003Cpath d=\"M536 148 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Crect x=\"536\" y=\"120\" width=\"164\" height=\"56\" rx=\"10\" class=\"d-accent\" \u002F>\u003Ctext x=\"618\" y=\"144\" text-anchor=\"middle\" class=\"d-text\">Fusion (RRF)\u003C\u002Ftext>\u003Ctext x=\"618\" y=\"163\" text-anchor=\"middle\" class=\"d-small\">ranks, not scores\u003C\u002Ftext>\u003Cpath d=\"M618 176 V198\" class=\"d-line\" \u002F>\u003Cpath d=\"M618 206 l-5 -9 h10 z\" class=\"d-head\" \u002F>\u003Crect x=\"536\" y=\"206\" width=\"164\" height=\"56\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"618\" y=\"230\" text-anchor=\"middle\" class=\"d-text\">Rerank top N\u003C\u002Ftext>\u003Ctext x=\"618\" y=\"249\" text-anchor=\"middle\" class=\"d-small\">cross-encoder\u003C\u002Ftext>\u003Cpath d=\"M618 262 V284\" class=\"d-line\" \u002F>\u003Cpath d=\"M618 292 l-5 -9 h10 z\" class=\"d-head\" \u002F>\u003Crect x=\"536\" y=\"292\" width=\"164\" height=\"56\" rx=\"10\" class=\"d-mint\" \u002F>\u003Ctext x=\"618\" y=\"316\" text-anchor=\"middle\" class=\"d-text\">Results + facets\u003C\u002Ftext>\u003Ctext x=\"618\" y=\"335\" text-anchor=\"middle\" class=\"d-small\">back to the shop\u003C\u002Ftext>\u003Crect x=\"320\" y=\"292\" width=\"180\" height=\"56\" rx=\"10\" class=\"d-gold\" \u002F>\u003Ctext x=\"410\" y=\"316\" text-anchor=\"middle\" class=\"d-text\">Pre-filters\u003C\u002Ftext>\u003Ctext x=\"410\" y=\"335\" text-anchor=\"middle\" class=\"d-small\">assortment, stock, price list\u003C\u002Ftext>\u003Crect x=\"20\" y=\"376\" width=\"680\" height=\"40\" rx=\"10\" class=\"d-box d-dash\" \u002F>\u003Ctext x=\"360\" y=\"401\" text-anchor=\"middle\" class=\"d-small\">Search logs: zero results, exits, clicks, feed synonyms and the judged query set\u003C\u002Ftext>",[269],"Identifier hits win outright, the other lanes compete through fusion, and every lane obeys the same filters.",{"type":151,"content":271},[272],"Two design decisions carry most of the weight. First, the identifier lane is not a feature of the lexical lane: it is its own query against its own fields, and its hits can short-circuit the rest. Second, the filters sit in front of all lanes, because in B2B they are not merely facets, they are entitlements.",{"type":151,"content":274},[275],"For the engine, either Elasticsearch or OpenSearch works. Both support BM25, approximate kNN and rank fusion. I would pick whichever your platform already runs, and avoid adding a second search engine until you have outgrown the first.",{"type":158,"level":159,"id":117,"text":118},{"type":151,"content":278},[279],"The cheapest way to ruin a B2B search is to let a semantic model decide how similar two part numbers are. To an embedding, \"4711-32\" and \"4711-33\" are nearly identical, and to a buyer they are two different parts. So identifiers get their own treatment.",{"type":151,"content":281},[282],"What I put in the identifier lane:",{"type":284,"ordered":285,"items":286},"list",false,[287,293,298,303],[288,292],{"tag":289,"children":290},"strong",[291],"Dedicated keyword fields"," for the article number, the manufacturer number, the supplier number, the customer-specific number and the EAN, always stored as strings.",[294,297],{"tag":289,"children":295},[296],"A normaliser"," that lowercases and strips dashes, dots, slashes and spaces at index and query time, so \"4711-32\", \"4711 32\" and \"471132\" meet in the same form.",[299,302],{"tag":289,"children":300},[301],"Exact first, prefix second."," A full match ranks above a prefix match, so a buyer typing the first digits still sees candidates while a complete number lands on one product.",[304,307],{"tag":289,"children":305},[306],"A short circuit."," If the normalised query is an exact identifier hit, return it without waiting for the vector arm or the reranker.",{"type":151,"content":309},[310,311,314],"Be careful with analyzers that split tokens for you. Elasticsearch's ",{"tag":164,"href":48,"children":312},[313],"word delimiter graph filter"," can split at letter-number transitions, so \"XL500\" becomes \"XL\" and \"500\". That helps free-text matching on model names, and it is harmful for identifiers, which is another reason to keep them in separate fields with their own analysis.",{"type":158,"level":159,"id":120,"text":121},{"type":151,"content":317},[318,319,322],"Lexical search is precise on words it knows. Vector search finds meaning across words it has never seen together. Elastic describes ",{"tag":164,"href":39,"children":320},[321],"hybrid search"," as often far better than the sum of the two, and names two fusion methods: a convex combination of normalised scores, and reciprocal rank fusion (RRF), which uses the position in each list and so needs no score normalisation.",{"type":151,"content":324},[325,326,329,330,333],"I start with RRF. BM25 scores and vector similarities live on different scales, and a weighted sum needs normalisation that is easy to get wrong and drifts when the catalogue changes. RRF adds 1\u002F(k + rank) per list, so the scales never need to match. OpenSearch offers the same idea through its ",{"tag":164,"href":54,"children":327},[328],"score ranker processor",", introduced in 2.19, with a rank constant between 1 and 10,000: a larger constant flattens the influence of top ranks, a smaller one favours them. It also offers a ",{"tag":164,"href":57,"children":331},[332],"normalization processor"," with min-max, L2 and z-score techniques if you prefer score-based fusion.",{"type":151,"content":335},[336],"Tune two things after the first version works: the rank window (how many candidates each lane contributes) and the per-lane weight. Both are cheap experiments against a judged query set, which I come back to in the evaluation section.",{"type":338,"variant":339,"title":340,"body":341},"callout","warn","Check licensing and version before you commit",[342],[343],"When Elastic published its hybrid search article, RRF ranking required a commercial (Enterprise) license, with a trial available. Licensing and features change, so check the current terms for your Elasticsearch version. On OpenSearch the RRF processor needs 2.19 or later, which matters if your cluster is old.",{"type":151,"content":345},[346],"One failure mode deserves its own warning. On an exact-identifier query the vector arm still returns something, and fusion can push a near-miss above the right part. This is why the identifier lane short-circuits, and why the judged set must contain identifier queries that assert rank one.",{"type":158,"level":159,"id":123,"text":124},{"type":151,"content":349},[350,351,354],"Filters in B2B are not cosmetics. A buyer may only see their negotiated assortment, their price list and what ships to their address. If you filter after the vector search, you ask for the ten nearest products and then discard the ones the customer may not buy, which leaves a short or empty list. Elasticsearch's ",{"tag":164,"href":42,"children":352},[353],"kNN query"," documents the difference: a pre-filter is applied during the approximate search so that k matching documents are returned, while a post-filter runs afterwards and can return fewer than k results even when enough matches exist.",{"type":151,"content":356},[357],"So assortment, availability, language and visibility go in as pre-filters on every lane. I treat this as a security property, not a relevance detail: a vector lane without the entitlement filter can surface products a customer is not allowed to see.",{"type":151,"content":359},[360],"Embeddings are weak on numbers and units. \"M8x40\" and \"M8x50\" embed almost the same, and \"1.5 inch\" and \"38 mm\" share little. For specifications I parse the query into structured attributes (thread size, length, material, standard) and apply them as filters or boosts against the attribute fields your PIM already maintains. The vector arm then handles the fuzzy part of the sentence, the use case and the product type, and the attributes do the exact part.",{"type":158,"level":159,"id":126,"text":127},{"type":151,"content":363},[364],"A multilingual catalogue gives you three separate problems. German forms long compounds, so \"Sechskantschraube\" may never match \"Schraube\" lexically. Hungarian is highly inflected, so a stem the buyer types may not equal the form in your product text. English queries against a German catalogue return nothing lexically, because there is no shared word.",{"type":151,"content":366},[367,368,371],"My approach is to combine both worlds. Lexical analysis runs per language, with the stemming and compound handling each language needs, on separate fields per locale. The vector arm uses one multilingual model so a query in one language can retrieve a product described in another. The ",{"tag":164,"href":69,"children":369},[370],"BGE-M3 paper"," describes one such model, with semantic retrieval in more than 100 working languages and inputs up to 8,192 tokens. I would still benchmark two or three candidates on my own queries, since catalogue vocabulary is far from general text.",{"type":151,"content":373},[374],"Two practical rules. Embed the text a buyer would recognise, such as title, key attributes and a short description, rather than the whole datasheet. And report every metric per language, because an average over three languages hides the one where search is broken.",{"type":158,"level":159,"id":129,"text":130},{"type":151,"content":377},[378,379,382],"Vectors do not remove the need for synonyms. They reduce it. Trade abbreviations, brand shorthand, old and new product names and supplier numbers are still best handled by explicit rules, because you can read, test and reverse them. Elasticsearch's ",{"tag":164,"href":51,"children":380},[381],"synonym graph filter"," is designed for search analyzers only, can be reloaded without reindexing when marked updateable, and takes rules from managed synonym sets (up to 100,000 rules per set by default).",{"type":151,"content":384},[385],"The best source of synonyms is your own zero-result log. Review the top failing queries weekly, decide whether each is a missing synonym, a missing product or a query for something you do not sell, and record the decision. That small ritual is worth more than any model upgrade, and it gives the vector arm a clean baseline to beat.",{"type":158,"level":159,"id":132,"text":133},{"type":151,"content":388},[389,390,395,396,400],"An LLM is good at the step in the middle: turning \"stainless hex bolt 8 by 40 for outdoors\" into a structured query with a product type, material, thread size, length and a language. It is poor at being the search engine. I use it as a parser with a strict output schema (see my posts on ",{"tag":391,"to":392,"children":393},"link","\u002Fblog\u002Fjev-typed-decisions-llm-routing",[394],"typed decisions"," and ",{"tag":391,"to":397,"children":398},"\u002Fblog\u002Fllm-evals-for-product-features",[399],"evaluating LLM features",") and nothing else.",{"type":151,"content":402},[403,404,407],"Instacart's account of ",{"tag":164,"href":66,"children":405},[406],"rebuilding query understanding with LLMs"," is a useful reference for the pattern. They injected catalogue taxonomy into prompts, added guardrails that check outputs by semantic similarity, and distilled the result into a smaller fine-tuned model. They served frequent queries from an offline cache and routed only the rare tail to a real-time model, reaching a 300 ms latency target. It is a consumer grocery case, but the shape transfers to B2B.",{"type":284,"ordered":285,"items":409},[410,415,420,430],[411,414],{"tag":289,"children":412},[413],"Validate against the catalogue."," If the LLM returns a material, an attribute value or a part number that does not exist in your data, drop it. Never let it invent identifiers.",[416,419],{"tag":289,"children":417},[418],"Cache the frequent queries."," B2B query distributions are short-headed, so a nightly batch over the top queries removes most real-time calls.",[421,424,425,429],{"tag":289,"children":422},[423],"Set a latency budget and a fallback."," If the parser is late or fails, run the plain hybrid query. Search must never be down because a model is slow. See ",{"tag":391,"to":426,"children":427},"\u002Fblog\u002Fllm-cost-latency-prompt-caching-routing",[428],"cost and latency routing",".",[431,434],{"tag":289,"children":432},[433],"Keep it away from identifiers."," If the identifier lane has an exact hit, the parser is not even called.",{"type":151,"content":436},[437],"Treat the parsed fields as hints, not truth. A boost on a parsed attribute is forgiving; a hard filter on a wrongly parsed attribute produces a zero-result page, so I start with boosts and promote an attribute to a filter only when its parse accuracy is proven.",{"type":158,"level":159,"id":135,"text":136},{"type":151,"content":440},[441,442,445],"Fusion gives a decent list, a reranker makes the first ten better. Elastic's ",{"tag":164,"href":45,"children":443},[444],"semantic reranking"," docs explain the trade-off: a cross-encoder reads query and document together and judges relevance better, at the price of larger models, higher latency and more compute. That is why it runs on a window of candidates (the rank window size) rather than on the whole result set, and why the documentation offers ways to limit the tokens sent, since long documents can be truncated before they reach the model.",{"type":151,"content":447},[448,449,453],"I would apply it only to non-identifier queries, rerank maybe the top 50 to 100 candidates, and send a short product text, not the datasheet. Then add business signals afterwards: availability, a customer's previous orders, preferred suppliers. For the general retrieval-and-rerank pattern, my ",{"tag":391,"to":450,"children":451},"\u002Fblog\u002Frag-pipeline-chunking-hybrid-search-reranking",[452],"RAG pipeline article"," goes deeper.",{"type":158,"level":159,"id":138,"text":139},{"type":151,"content":456},[457],"Without measurement, semantic search is a demo. I track a small set of search metrics, always split by query type and language.",{"type":172,"head":459,"rows":466},[460,462,464],[461],"Metric",[463],"What it tells you",[465],"Trap",[467,478,485,492,499],[468,470,476],[469],"Zero-result rate",[471,472,475],"Share of searches that return nothing; analytics tools such as ",{"tag":164,"href":72,"children":473},[474],"Algolia"," report it as the no results rate",[477],"Semantic search drives it down by returning junk; read it with click-through",[479,481,483],[480],"Search-exit rate",[482],"Share of searches after which the visitor leaves (my definition: no click, no refinement, session ends)",[484],"Needs your own event tracking; bots and bookmarks add noise",[486,488,490],[487],"Click-through rate",[489],"Share of searches with at least one click on a result",[491],"Position bias: a better top result raises it, a bad one hides below the fold",[493,495,497],[494],"Reformulation rate",[496],"Share of searches followed by another query in the same session",[498],"Some reformulation is healthy refinement",[500,502,504],[501],"Rank-one accuracy for identifiers",[503],"Judged set: does the exact part come first",[505],"Must stay at 100%; any drop is a regression",{"type":151,"content":507},[508],"The judged set is the part most teams skip. I take a few hundred real queries from the logs, stratified by the query types in the first table, and record which products are right. Identifier queries assert an exact rank one; descriptive queries assert that a relevant product is in the top ten. Run it on every change to analyzers, synonyms, embeddings or fusion settings, in CI if you can.",{"type":151,"content":510},[511],"Then run an A\u002FB test on live traffic and compare click-through, add-to-cart from search and search-exit rate per query type. A zero-result query is sometimes correct, because the part is not in the range, so do not chase the rate to zero; chase the number of zero-result queries that should have found something.",{"type":158,"level":159,"id":141,"text":142},{"type":151,"content":514},[515,516,519,520,523],"Spryker is ",{"tag":164,"href":60,"children":517},[518],"shipped with Elasticsearch as its default search",", indexing product name, description and SKU, product attributes, reviews and CMS pages. The documentation also describes third-party search integrations and a tutorial for integrating any search engine, and ",{"tag":164,"href":63,"children":521},[522],"a migration path for OpenSearch"," from 1.3 via 2.19 to 3.5. That upgrade matters here, because hybrid fusion in OpenSearch needs a recent version.",{"type":151,"content":525},[526],"My suggested integration, which is a design proposal and not a Spryker feature, has four steps. Compute an embedding for each abstract product and locale in the publish-and-sync flow that builds the search documents, and store it in a vector field next to the text fields. Extend the search query so that it issues the identifier, lexical and vector clauses under the shop's existing filters, and fuse them with RRF. Put the LLM parser in front, behind a timeout. Wrap everything in a feature flag per store and locale.",{"type":151,"content":528},[529],"Two caveats from experience with these platforms. Vectors increase index size and publish time, so size the cluster and re-embed only when the embedded text changes. And keep the existing search as a fallback path: if the vector lane is unavailable, the shop should still answer through the lexical and identifier lanes.",{"type":151,"content":531},[532],"If you are building from scratch, a hosted search product is an alternative, and Spryker documents integrations for that route. I would still insist on the same four things from any vendor: identifier handling, entitlement pre-filters, per-language metrics and a way to run your judged set.",{"type":158,"level":159,"id":144,"text":145},{"type":151,"content":535},[536],"This is the order I would work in.",{"type":284,"ordered":538,"items":539},true,[540,542,544,546,548,550,552,554],[541],"Export three months of search logs and classify queries into the types of the first table. Count them.",[543],"Build the judged set, with identifier queries that assert rank one, and measure the current search as the baseline.",[545],"Fix the lexical basics: identifier fields with a normaliser, per-language analyzers, a maintained synonym set.",[547],"Add the vector lane with a multilingual model, entitlement pre-filters and RRF fusion, behind a feature flag.",[549],"Short-circuit identifier hits so the vector arm and the reranker never touch them.",[551],"Add the LLM parser with schema validation, an offline cache for frequent queries, a timeout and a boost-first policy.",[553],"Add a reranker on a small window for non-identifier queries, and check latency at p95.",[555],"Run an A\u002FB test, compare metrics per query type and language, and review the zero-result log weekly.",{"type":151,"content":557},[558],"Most of the gain in these projects comes from steps 3 and 4, not from the most fashionable model, and the part that earns trust is step 2. Make the identifier lane boringly reliable first, and semantic search will feel like an upgrade instead of a risk.",{"type":158,"level":159,"id":147,"text":148},{"type":284,"ordered":538,"items":561},[562,565,568,571,574,577,580,583,586,589,592,595,598],[563],{"tag":164,"href":36,"children":564},[35],[566],{"tag":164,"href":39,"children":567},[38],[569],{"tag":164,"href":42,"children":570},[41],[572],{"tag":164,"href":45,"children":573},[44],[575],{"tag":164,"href":48,"children":576},[47],[578],{"tag":164,"href":51,"children":579},[50],[581],{"tag":164,"href":54,"children":582},[53],[584],{"tag":164,"href":57,"children":585},[56],[587],{"tag":164,"href":60,"children":588},[59],[590],{"tag":164,"href":63,"children":591},[62],[593],{"tag":164,"href":66,"children":594},[65],[596],{"tag":164,"href":69,"children":597},[68],[599],{"tag":164,"href":72,"children":600},[71],[602,678,760,827],{"slug":603,"published":5,"minutes":6,"category":7,"tags":604,"keywords":610,"about":621,"sources":631,"cover":671,"og":672,"expertise":673,"locales":674,"lang":77,"title":675,"description":676,"coverAlt":677},"llm-hallucination-grounding-citations",[605,606,607,608,609],"Hallucinations","RAG","Citations","Grounding","Faithfulness",[611,612,613,614,615,616,617,618,619,620],"reduce LLM hallucinations in production","how to reduce hallucinations in RAG","LLM citations API","Anthropic citations API","RAG faithfulness metric","LLM abstention I don't know","claim-level verification LLM","grounding LLM answers in sources","check grounding API","show sources in AI chatbot UI",[622,625,628],{"name":623,"url":624},"Hallucination (artificial intelligence)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FHallucination_(artificial_intelligence)",{"name":626,"url":627},"Retrieval-augmented generation","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FRetrieval-augmented_generation",{"name":629,"url":630},"Large language model","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FLarge_language_model",[632,635,638,641,644,647,650,653,656,659,662,665,668],{"title":633,"url":634},"Anthropic: Citations (Claude API documentation)","https:\u002F\u002Fplatform.claude.com\u002Fdocs\u002Fen\u002Fbuild-with-claude\u002Fcitations",{"title":636,"url":637},"Anthropic: Search results (Claude API documentation)","https:\u002F\u002Fplatform.claude.com\u002Fdocs\u002Fen\u002Fbuild-with-claude\u002Fsearch-results",{"title":639,"url":640},"Anthropic: Reduce hallucinations (Claude API documentation)","https:\u002F\u002Fplatform.claude.com\u002Fdocs\u002Fen\u002Ftest-and-evaluate\u002Fstrengthen-guardrails\u002Freduce-hallucinations",{"title":642,"url":643},"Anthropic: Introducing Citations on the Anthropic API","https:\u002F\u002Fclaude.com\u002Fblog\u002Fintroducing-citations-api",{"title":645,"url":646},"Simon Willison: Anthropic's new Citations API (24 January 2025)","https:\u002F\u002Fsimonwillison.net\u002F2025\u002FJan\u002F24\u002Fanthropics-new-citations-api\u002F",{"title":648,"url":649},"OpenAI: Web search guide (url_citation annotations and display requirement)","https:\u002F\u002Fdevelopers.openai.com\u002Fapi\u002Fdocs\u002Fguides\u002Ftools-web-search",{"title":651,"url":652},"Cohere: Documents and citations","https:\u002F\u002Fdocs.cohere.com\u002Fdocs\u002Fdocuments-and-citations",{"title":654,"url":655},"Google Cloud: Check grounding API","https:\u002F\u002Fdocs.cloud.google.com\u002Fgenerative-ai-app-builder\u002Fdocs\u002Fcheck-grounding",{"title":657,"url":658},"AWS: Amazon Bedrock Guardrails contextual grounding check","https:\u002F\u002Fdocs.aws.amazon.com\u002Fbedrock\u002Flatest\u002Fuserguide\u002Fguardrails-contextual-grounding-check.html",{"title":660,"url":661},"Ragas: Faithfulness metric","https:\u002F\u002Fdocs.ragas.io\u002Fen\u002Fstable\u002Fconcepts\u002Fmetrics\u002Favailable_metrics\u002Ffaithfulness\u002F",{"title":663,"url":664},"Kalai, Nachum, Vempala, Zhang: Why Language Models Hallucinate (arXiv 2509.04664)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2509.04664",{"title":666,"url":667},"Magesh et al.: Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools (arXiv 2405.20362)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2405.20362",{"title":669,"url":670},"Wallat, Heuss, de Rijke, Anand: Correctness is not Faithfulness in RAG Attributions (arXiv 2412.18004)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2412.18004","\u002Fimages\u002Fblog\u002Fllm-hallucination-grounding-citations\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fllm-hallucination-grounding-citations\u002Fog.jpg","ai-engineer",[77,78,79],"Reducing LLM hallucinations in production: grounding, citations and knowing when to say no","Cut hallucinations in production RAG: citation APIs, abstention, claim-level checks, faithfulness metrics, source UI, and the failures that still slip through.","Diagram: a retrieval step feeds an evidence gate, a cited answer and a claim verifier, ending in an answer with sources, with abstain and flag paths branching off.",{"slug":679,"published":5,"minutes":6,"category":7,"tags":680,"keywords":684,"about":695,"sources":703,"cover":754,"og":755,"expertise":673,"locales":756,"lang":77,"title":757,"description":758,"coverAlt":759},"pgvector-vs-vector-databases",[681,682,606,10,683],"pgvector","Vector databases","EU hosting",[685,686,687,688,689,690,691,692,693,694],"pgvector vs vector database","pgvector vs Qdrant","pgvector vs Pinecone","best vector database 2026","pgvector HNSW iterative scan","pgvector halfvec","OpenSearch vs Elasticsearch vector search","vector database EU hosting","hybrid search Postgres","Weaviate vs Milvus",[696,699,702],{"name":697,"url":698},"Vector database","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FVector_database",{"name":700,"url":701},"PostgreSQL","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FPostgreSQL",{"name":626,"url":627},[704,707,710,713,716,719,722,725,728,731,734,737,740,743,746,749,751],{"title":705,"url":706},"pgvector README (index limits, HNSW defaults, iterative scans, filtering, halfvec)","https:\u002F\u002Fgithub.com\u002Fpgvector\u002Fpgvector\u002Fblob\u002Fmaster\u002FREADME.md",{"title":708,"url":709},"pgvector CHANGELOG (0.4.0 to 0.8.7)","https:\u002F\u002Fgithub.com\u002Fpgvector\u002Fpgvector\u002Fblob\u002Fmaster\u002FCHANGELOG.md",{"title":711,"url":712},"pgvectorscale: StreamingDiskANN, statistical binary quantization, filtered search","https:\u002F\u002Fgithub.com\u002Ftimescale\u002Fpgvectorscale",{"title":714,"url":715},"Qdrant documentation: Filtering","https:\u002F\u002Fqdrant.tech\u002Fdocumentation\u002Fconcepts\u002Ffiltering\u002F",{"title":717,"url":718},"Qdrant documentation: Hybrid queries","https:\u002F\u002Fqdrant.tech\u002Fdocumentation\u002Fconcepts\u002Fhybrid-queries\u002F",{"title":720,"url":721},"Qdrant documentation: Create a cluster (providers, free tier, Hybrid Cloud)","https:\u002F\u002Fqdrant.tech\u002Fdocumentation\u002Fcloud\u002Fcreate-cluster\u002F",{"title":723,"url":724},"Weaviate documentation: Hybrid search","https:\u002F\u002Fdocs.weaviate.io\u002Fweaviate\u002Fconcepts\u002Fsearch\u002Fhybrid-search",{"title":726,"url":727},"Weaviate documentation: Vector index types","https:\u002F\u002Fdocs.weaviate.io\u002Fweaviate\u002Fconcepts\u002Fvector-index",{"title":729,"url":730},"Weaviate Cloud pricing and deployment options","https:\u002F\u002Fweaviate.io\u002Fpricing",{"title":732,"url":733},"Milvus documentation: Overview","https:\u002F\u002Fmilvus.io\u002Fdocs\u002Foverview.md",{"title":735,"url":736},"Pinecone documentation: Database architecture","https:\u002F\u002Fdocs.pinecone.io\u002Fguides\u002Fget-started\u002Fdatabase-architecture",{"title":738,"url":739},"Pinecone documentation: Create an index (clouds, regions, sparse and hybrid)","https:\u002F\u002Fdocs.pinecone.io\u002Fguides\u002Findex-data\u002Fcreate-an-index",{"title":741,"url":742},"OpenSearch documentation: Methods and engines","https:\u002F\u002Fdocs.opensearch.org\u002Flatest\u002Fmappings\u002Fsupported-field-types\u002Fknn-methods-engines\u002F",{"title":744,"url":745},"OpenSearch documentation: Efficient k-NN filtering","https:\u002F\u002Fdocs.opensearch.org\u002Flatest\u002Fvector-search\u002Ffilter-search-knn\u002Fefficient-knn-filtering\u002F",{"title":747,"url":748},"Elasticsearch documentation: Dense vector search","https:\u002F\u002Fwww.elastic.co\u002Fdocs\u002Fsolutions\u002Fsearch\u002Fvector\u002Fdense-vector",{"title":750,"url":42},"Elasticsearch documentation: kNN query (filter as pre-filter)",{"title":752,"url":753},"GitHub releases: Qdrant, Weaviate, Milvus, OpenSearch, pgvectorscale (versions as of 1 October 2026)","https:\u002F\u002Fgithub.com\u002Fqdrant\u002Fqdrant\u002Freleases","\u002Fimages\u002Fblog\u002Fpgvector-vs-vector-databases\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fpgvector-vs-vector-databases\u002Fog.jpg",[77,78,79],"pgvector or a vector database? How to choose vector storage in 2026","pgvector, Qdrant, Weaviate, Milvus, Pinecone, OpenSearch or Elasticsearch? A practical 2026 guide to filtering, hybrid search, scale, cost and EU hosting.","Diagram: a decision path from your data to pgvector in Postgres, a search engine with vector fields, or a dedicated vector database.",{"slug":761,"published":5,"minutes":6,"category":7,"tags":762,"keywords":766,"about":776,"sources":784,"cover":821,"og":822,"expertise":673,"locales":823,"lang":77,"title":824,"description":825,"coverAlt":826},"graphrag-knowledge-graph-rag",[763,764,765,606],"GraphRAG","Knowledge graphs","LightRAG",[763,767,768,769,770,771,772,773,774,775],"knowledge graph RAG","GraphRAG vs vector RAG","Microsoft GraphRAG explained","LightRAG vs GraphRAG","GraphRAG global vs local search","GraphRAG indexing cost","when to use GraphRAG","multi-hop RAG","LazyGraphRAG",[777,780,781],{"name":778,"url":779},"Knowledge graph","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FKnowledge_graph",{"name":626,"url":627},{"name":782,"url":783},"Leiden algorithm","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FLeiden_algorithm",[785,788,791,794,797,800,803,806,809,812,815,818],{"title":786,"url":787},"Edge et al.: From Local to Global: A Graph RAG Approach to Query-Focused Summarization (arXiv:2404.16130)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2404.16130",{"title":789,"url":790},"Microsoft GraphRAG documentation: overview","https:\u002F\u002Fmicrosoft.github.io\u002Fgraphrag\u002F",{"title":792,"url":793},"Microsoft GraphRAG documentation: default dataflow","https:\u002F\u002Fmicrosoft.github.io\u002Fgraphrag\u002Findex\u002Fdefault_dataflow\u002F",{"title":795,"url":796},"Microsoft GraphRAG documentation: global search","https:\u002F\u002Fmicrosoft.github.io\u002Fgraphrag\u002Fquery\u002Fglobal_search\u002F",{"title":798,"url":799},"Microsoft GraphRAG documentation: local search","https:\u002F\u002Fmicrosoft.github.io\u002Fgraphrag\u002Fquery\u002Flocal_search\u002F",{"title":801,"url":802},"Microsoft GraphRAG documentation: DRIFT search","https:\u002F\u002Fmicrosoft.github.io\u002Fgraphrag\u002Fquery\u002Fdrift_search\u002F",{"title":804,"url":805},"Microsoft GraphRAG documentation: getting started","https:\u002F\u002Fmicrosoft.github.io\u002Fgraphrag\u002Fget_started\u002F",{"title":807,"url":808},"Microsoft Research: LazyGraphRAG, setting a new standard for quality and cost (25 November 2024)","https:\u002F\u002Fwww.microsoft.com\u002Fen-us\u002Fresearch\u002Fblog\u002Flazygraphrag-setting-a-new-standard-for-quality-and-cost\u002F",{"title":810,"url":811},"Guo et al.: LightRAG: Simple and Fast Retrieval-Augmented Generation (arXiv:2410.05779)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2410.05779",{"title":813,"url":814},"HKUDS\u002FLightRAG on GitHub","https:\u002F\u002Fgithub.com\u002FHKUDS\u002FLightRAG",{"title":816,"url":817},"Gutiérrez et al.: HippoRAG: Neurobiologically Inspired Long-Term Memory for Large Language Models (arXiv:2405.14831)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2405.14831",{"title":819,"url":820},"Xiang et al.: When to use Graphs in RAG: A Comprehensive Analysis for Graph Retrieval-Augmented Generation (arXiv:2506.05690)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2506.05690","\u002Fimages\u002Fblog\u002Fgraphrag-knowledge-graph-rag\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fgraphrag-knowledge-graph-rag\u002Fog.jpg",[77,78,79],"GraphRAG and knowledge-graph RAG: when a graph beats vector search","What Microsoft GraphRAG and LightRAG really do, what indexing costs, and when a knowledge graph beats vector RAG: multi-hop, global questions, product catalogues.","Diagram: a knowledge graph hub linked to entities, communities, local search, global search and product parts.",{"slug":828,"published":5,"minutes":829,"category":7,"tags":830,"keywords":836,"about":846,"sources":854,"cover":888,"og":889,"expertise":673,"locales":890,"lang":77,"title":891,"description":892,"coverAlt":893},"rag-evaluation-metrics",12,[831,832,833,834,835],"RAG evaluation","Retrieval metrics","LLM-as-judge","Golden set","Ragas",[831,837,838,839,840,841,842,843,844,845],"how to evaluate RAG","RAG evaluation metrics","recall@k MRR nDCG","faithfulness vs answer relevance","golden dataset for RAG","LLM as a judge calibration","Ragas vs DeepEval vs TruLens","RAG evals in CI","retrieval vs generation failure",[847,848,851],{"name":626,"url":627},{"name":849,"url":850},"Discounted cumulative gain","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FDiscounted_cumulative_gain",{"name":852,"url":853},"Mean reciprocal rank","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FMean_reciprocal_rank",[855,858,861,864,866,869,872,875,878,881,883,885],{"title":856,"url":857},"Es et al.: RAGAS, Automated Evaluation of Retrieval Augmented Generation (arXiv 2309.15217)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2309.15217",{"title":859,"url":860},"Zheng et al.: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (arXiv 2306.05685)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2306.05685",{"title":862,"url":863},"Ragas documentation: available metrics","https:\u002F\u002Fdocs.ragas.io\u002Fen\u002Fstable\u002Fconcepts\u002Fmetrics\u002Favailable_metrics\u002F",{"title":865,"url":661},"Ragas documentation: faithfulness",{"title":867,"url":868},"Ragas documentation: context precision","https:\u002F\u002Fdocs.ragas.io\u002Fen\u002Fstable\u002Fconcepts\u002Fmetrics\u002Favailable_metrics\u002Fcontext_precision\u002F",{"title":870,"url":871},"Ragas documentation: context recall","https:\u002F\u002Fdocs.ragas.io\u002Fen\u002Fstable\u002Fconcepts\u002Fmetrics\u002Favailable_metrics\u002Fcontext_recall\u002F",{"title":873,"url":874},"DeepEval documentation: metrics introduction","https:\u002F\u002Fdeepeval.com\u002Fdocs\u002Fmetrics-introduction",{"title":876,"url":877},"TruLens","https:\u002F\u002Fwww.trulens.org\u002F",{"title":879,"url":880},"Arize Phoenix documentation","https:\u002F\u002Farize.com\u002Fdocs\u002Fphoenix",{"title":882,"url":850},"Wikipedia: Discounted cumulative gain",{"title":884,"url":853},"Wikipedia: Mean reciprocal rank",{"title":886,"url":887},"Wikipedia: Cohen's kappa","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FCohen%27s_kappa","\u002Fimages\u002Fblog\u002Frag-evaluation-metrics\u002Fcover.webp","\u002Fimages\u002Fblog\u002Frag-evaluation-metrics\u002Fog.jpg",[77,78,79],"Evaluating RAG: retrieval metrics, faithfulness and how to tell which half failed","How to evaluate a RAG system: recall at k, MRR and nDCG vs faithfulness and answer relevance, a golden set from real queries, a calibrated LLM judge and evals in CI.","Diagram: a RAG answer is scored on two sides, retrieval metrics such as recall at k, MRR and nDCG, and generation metrics such as faithfulness and answer relevance, feeding a diagnosis.",1791009037158]