> How to add semantic search to a B2B shop without breaking part-number search: hybrid BM25 and vectors, filters, DE/EN/HU, LLM query parsing, reranking and metrics.
>
> Web page: https://balazscsorba.com/blog/semantic-product-search-b2b · Language: English · Also available in: [Deutsch](https://balazscsorba.com/de/blog/semantic-product-search-b2b.md) · [Magyar](https://balazscsorba.com/hu/blog/semantic-product-search-b2b.md)
> Author: Balázs Csorba · Published: 2026-10-02 · Keywords: semantic product search B2B, B2B ecommerce search, hybrid search BM25 vector, part number search ecommerce, Spryker search Elasticsearch, OpenSearch hybrid search RRF, multilingual product search German English Hungarian, zero results rate site search, LLM query understanding ecommerce, AI product search for B2B shops

[Blog](https://balazscsorba.com/blog)/RAG & retrieval

# Semantic product search for B2B shops: part numbers, hybrid retrieval and what to measure

How to add semantic search to a B2B shop without breaking part-number search: hybrid BM25 and vectors, filters, DE/EN/HU, LLM query parsing, reranking and metrics.

[Balázs Csorba](https://balazscsorba.com/about)·October 2, 2026·13 min read

-   B2B search
-   Hybrid search
-   Semantic search
-   Spryker
-   OpenSearch

![Diagram: a search query is split into an identifier lane, lexical BM25 and vector kNN, fused, reranked and returned as results.](https://balazscsorba.com/images/blog/semantic-product-search-b2b/cover.webp?v=a90485ce98)

## Key takeaways

-   In B2B, the part number is the most important query. Give identifiers their own lane with normalised exact and prefix matching, and let semantic retrieval help only when that lane has no strong answer.
-   Hybrid search (BM25 plus vectors, fused with reciprocal rank fusion) beats either method alone, because buyers type both "10-32-4711" and "screw for outdoor wood" into the same box.
-   Assortments, price lists and stock must be pre-filters inside the vector search, not post-filters, otherwise results disappear or leak across customers.
-   Use an LLM to parse queries into product type, attributes and units, validate its output against the catalogue, and cache it offline for frequent queries instead of calling it on every keystroke.
-   Judge the system by query type: zero-result rate, search-exit rate and click-through, plus a small judged query set where exact identifiers must always rank first.

On this page

1.  [Why B2B search is not consumer search](https://balazscsorba.com/#why-b2b-search-differs)
2.  [The architecture: lanes, fusion, rerank](https://balazscsorba.com/#architecture)
3.  [Article numbers and exact match come first](https://balazscsorba.com/#exact-match-first)
4.  [Hybrid retrieval: BM25 plus vectors, fused by rank](https://balazscsorba.com/#hybrid-retrieval)
5.  [Attribute-aware filtering, and why numbers need structure](https://balazscsorba.com/#filters-and-attributes)
6.  [German, English and Hungarian in one catalogue](https://balazscsorba.com/#multilingual)
7.  [Synonyms and part numbers still matter](https://balazscsorba.com/#synonyms)
8.  [Query understanding with an LLM](https://balazscsorba.com/#llm-query-understanding)
9.  [Reranking the top of the list](https://balazscsorba.com/#reranking)
10.  [Evaluation: zero results, exits and a judged set](https://balazscsorba.com/#evaluation)
11.  [Integrating with Spryker, Elasticsearch and OpenSearch](https://balazscsorba.com/#spryker-integration)
12.  [A rollout checklist](https://balazscsorba.com/#rollout-checklist)
13.  [Sources](https://balazscsorba.com/#sources)

A buyer at a plumbing wholesaler types "4711-32". A second one types "Edelstahlschraube für Holz außen". A third types "hex bolt M8x40 A2". All three use the same search box, and in most B2B shops I have seen, at least one of them gets a page of nothing.

Semantic search promises to fix the second and third case. The risk is that it breaks the first. This article is how I would add semantic retrieval to a B2B catalogue without losing exact part-number search: the architecture, the query types, hybrid fusion, filters, three languages, synonyms, LLM query understanding, reranking, measurement and what it means for a Spryker shop on Elasticsearch or OpenSearch.

## Why B2B search is not consumer search

B2B queries are more heterogeneous than consumer queries. Buyers paste an article number from a drawing, retype a supplier number from an old order, abbreviate trade jargon, or describe a use case. Baymard, which benchmarks consumer shops, distinguishes [eight kinds of search query](https://baymard.com/blog/ecommerce-search-query-types) and found that 56% of the sites it tested do not adequately support users' search needs. B2B catalogues add identifier-heavy, specification-heavy data on top.

The practical consequence is that no single retrieval method is right. Here is the taxonomy I use when I start a project. The examples are invented, but the patterns are the ones that show up in query logs.

Query type

Example

Best retrieval

Typical failure

Article or part number

"4711-32", "4711 32"

Identifier lane: normalised exact, then prefix

Dashes and spaces break exact match; a vector arm returns look-alikes

Supplier number, EAN

"4006381333931"

Identifier lane on its own field

Stored as a number, leading zeros lost

Product type

"Sechskantschraube"

Lexical plus vector, category boost

Compounds and plurals miss in lexical search

Specification

"M8x40 A2 DIN 933"

Parsed into filters plus lexical

Embeddings blur numbers and units

Use case

"screw for outdoor wood"

Vector plus attribute filters

Lexical search returns zero results

Abbreviation or jargon

"VA Schraube"

Synonyms, then vector

Abbreviation unknown to the analyzer

Cross-language

"hex bolt" in a German catalogue

Multilingual vector plus synonyms

Lexical search returns zero results

Non-product

"delivery time", "datasheet"

Route to help or CMS content

Product index returns random products

Count how many of your real queries fall into each row before you choose anything. In a catalogue of technical parts, the first two rows can be a large share, and they are exactly the rows where semantic search adds nothing and can do harm.

## The architecture: lanes, fusion, rerank

My reference design has one entry point and three retrieval lanes that run in parallel. A cheap understanding step normalises the query and detects whether it looks like an identifier. The identifier lane, a lexical BM25 query and a vector kNN query all run under the same filters. Their ranked lists are fused, a reranker reorders only the top of the list, and the response carries the facets the shop needs.

Identifier hits win outright, the other lanes compete through fusion, and every lane obeys the same filters.

Two design decisions carry most of the weight. First, the identifier lane is not a feature of the lexical lane: it is its own query against its own fields, and its hits can short-circuit the rest. Second, the filters sit in front of all lanes, because in B2B they are not merely facets, they are entitlements.

For the engine, either Elasticsearch or OpenSearch works. Both support BM25, approximate kNN and rank fusion. I would pick whichever your platform already runs, and avoid adding a second search engine until you have outgrown the first.

## Article numbers and exact match come first

The cheapest way to ruin a B2B search is to let a semantic model decide how similar two part numbers are. To an embedding, "4711-32" and "4711-33" are nearly identical, and to a buyer they are two different parts. So identifiers get their own treatment.

What I put in the identifier lane:

-   **Dedicated keyword fields** for the article number, the manufacturer number, the supplier number, the customer-specific number and the EAN, always stored as strings.
-   **A normaliser** that lowercases and strips dashes, dots, slashes and spaces at index and query time, so "4711-32", "4711 32" and "471132" meet in the same form.
-   **Exact first, prefix second.** A full match ranks above a prefix match, so a buyer typing the first digits still sees candidates while a complete number lands on one product.
-   **A short circuit.** If the normalised query is an exact identifier hit, return it without waiting for the vector arm or the reranker.

Be careful with analyzers that split tokens for you. Elasticsearch's [word delimiter graph filter](https://www.elastic.co/docs/reference/text-analysis/analysis-word-delimiter-graph-tokenfilter) can split at letter-number transitions, so "XL500" becomes "XL" and "500". That helps free-text matching on model names, and it is harmful for identifiers, which is another reason to keep them in separate fields with their own analysis.

## Hybrid retrieval: BM25 plus vectors, fused by rank

Lexical search is precise on words it knows. Vector search finds meaning across words it has never seen together. Elastic describes [hybrid search](https://www.elastic.co/search-labs/blog/hybrid-search-elasticsearch) as often far better than the sum of the two, and names two fusion methods: a convex combination of normalised scores, and reciprocal rank fusion (RRF), which uses the position in each list and so needs no score normalisation.

I start with RRF. BM25 scores and vector similarities live on different scales, and a weighted sum needs normalisation that is easy to get wrong and drifts when the catalogue changes. RRF adds 1/(k + rank) per list, so the scales never need to match. OpenSearch offers the same idea through its [score ranker processor](https://docs.opensearch.org/latest/search-plugins/search-pipelines/score-ranker-processor/), introduced in 2.19, with a rank constant between 1 and 10,000: a larger constant flattens the influence of top ranks, a smaller one favours them. It also offers a [normalization processor](https://docs.opensearch.org/latest/search-plugins/search-pipelines/normalization-processor/) with min-max, L2 and z-score techniques if you prefer score-based fusion.

Tune two things after the first version works: the rank window (how many candidates each lane contributes) and the per-lane weight. Both are cheap experiments against a judged query set, which I come back to in the evaluation section.

**Check licensing and version before you commit**

When Elastic published its hybrid search article, RRF ranking required a commercial (Enterprise) license, with a trial available. Licensing and features change, so check the current terms for your Elasticsearch version. On OpenSearch the RRF processor needs 2.19 or later, which matters if your cluster is old.

One failure mode deserves its own warning. On an exact-identifier query the vector arm still returns something, and fusion can push a near-miss above the right part. This is why the identifier lane short-circuits, and why the judged set must contain identifier queries that assert rank one.

## Attribute-aware filtering, and why numbers need structure

Filters in B2B are not cosmetics. A buyer may only see their negotiated assortment, their price list and what ships to their address. If you filter after the vector search, you ask for the ten nearest products and then discard the ones the customer may not buy, which leaves a short or empty list. Elasticsearch's [kNN query](https://www.elastic.co/docs/reference/query-languages/query-dsl/query-dsl-knn-query) documents the difference: a pre-filter is applied during the approximate search so that k matching documents are returned, while a post-filter runs afterwards and can return fewer than k results even when enough matches exist.

So assortment, availability, language and visibility go in as pre-filters on every lane. I treat this as a security property, not a relevance detail: a vector lane without the entitlement filter can surface products a customer is not allowed to see.

Embeddings are weak on numbers and units. "M8x40" and "M8x50" embed almost the same, and "1.5 inch" and "38 mm" share little. For specifications I parse the query into structured attributes (thread size, length, material, standard) and apply them as filters or boosts against the attribute fields your PIM already maintains. The vector arm then handles the fuzzy part of the sentence, the use case and the product type, and the attributes do the exact part.

## German, English and Hungarian in one catalogue

A multilingual catalogue gives you three separate problems. German forms long compounds, so "Sechskantschraube" may never match "Schraube" lexically. Hungarian is highly inflected, so a stem the buyer types may not equal the form in your product text. English queries against a German catalogue return nothing lexically, because there is no shared word.

My approach is to combine both worlds. Lexical analysis runs per language, with the stemming and compound handling each language needs, on separate fields per locale. The vector arm uses one multilingual model so a query in one language can retrieve a product described in another. The [BGE-M3 paper](https://arxiv.org/abs/2402.03216) describes one such model, with semantic retrieval in more than 100 working languages and inputs up to 8,192 tokens. I would still benchmark two or three candidates on my own queries, since catalogue vocabulary is far from general text.

Two practical rules. Embed the text a buyer would recognise, such as title, key attributes and a short description, rather than the whole datasheet. And report every metric per language, because an average over three languages hides the one where search is broken.

## Synonyms and part numbers still matter

Vectors do not remove the need for synonyms. They reduce it. Trade abbreviations, brand shorthand, old and new product names and supplier numbers are still best handled by explicit rules, because you can read, test and reverse them. Elasticsearch's [synonym graph filter](https://www.elastic.co/docs/reference/text-analysis/analysis-synonym-graph-tokenfilter) is designed for search analyzers only, can be reloaded without reindexing when marked updateable, and takes rules from managed synonym sets (up to 100,000 rules per set by default).

The best source of synonyms is your own zero-result log. Review the top failing queries weekly, decide whether each is a missing synonym, a missing product or a query for something you do not sell, and record the decision. That small ritual is worth more than any model upgrade, and it gives the vector arm a clean baseline to beat.

## Query understanding with an LLM

An LLM is good at the step in the middle: turning "stainless hex bolt 8 by 40 for outdoors" into a structured query with a product type, material, thread size, length and a language. It is poor at being the search engine. I use it as a parser with a strict output schema (see my posts on [typed decisions](https://balazscsorba.com/blog/jev-typed-decisions-llm-routing) and [evaluating LLM features](https://balazscsorba.com/blog/llm-evals-for-product-features)) and nothing else.

Instacart's account of [rebuilding query understanding with LLMs](https://www.zenml.io/llmops-database/rebuilding-query-understanding-for-e-commerce-search-with-llms) is a useful reference for the pattern. They injected catalogue taxonomy into prompts, added guardrails that check outputs by semantic similarity, and distilled the result into a smaller fine-tuned model. They served frequent queries from an offline cache and routed only the rare tail to a real-time model, reaching a 300 ms latency target. It is a consumer grocery case, but the shape transfers to B2B.

-   **Validate against the catalogue.** If the LLM returns a material, an attribute value or a part number that does not exist in your data, drop it. Never let it invent identifiers.
-   **Cache the frequent queries.** B2B query distributions are short-headed, so a nightly batch over the top queries removes most real-time calls.
-   **Set a latency budget and a fallback.** If the parser is late or fails, run the plain hybrid query. Search must never be down because a model is slow. See [cost and latency routing](https://balazscsorba.com/blog/llm-cost-latency-prompt-caching-routing).
-   **Keep it away from identifiers.** If the identifier lane has an exact hit, the parser is not even called.

Treat the parsed fields as hints, not truth. A boost on a parsed attribute is forgiving; a hard filter on a wrongly parsed attribute produces a zero-result page, so I start with boosts and promote an attribute to a filter only when its parse accuracy is proven.

## Reranking the top of the list

Fusion gives a decent list, a reranker makes the first ten better. Elastic's [semantic reranking](https://www.elastic.co/docs/solutions/search/ranking/semantic-reranking) docs explain the trade-off: a cross-encoder reads query and document together and judges relevance better, at the price of larger models, higher latency and more compute. That is why it runs on a window of candidates (the rank window size) rather than on the whole result set, and why the documentation offers ways to limit the tokens sent, since long documents can be truncated before they reach the model.

I would apply it only to non-identifier queries, rerank maybe the top 50 to 100 candidates, and send a short product text, not the datasheet. Then add business signals afterwards: availability, a customer's previous orders, preferred suppliers. For the general retrieval-and-rerank pattern, my [RAG pipeline article](https://balazscsorba.com/blog/rag-pipeline-chunking-hybrid-search-reranking) goes deeper.

## Evaluation: zero results, exits and a judged set

Without measurement, semantic search is a demo. I track a small set of search metrics, always split by query type and language.

Metric

What it tells you

Trap

Zero-result rate

Share of searches that return nothing; analytics tools such as [Algolia](https://www.algolia.com/doc/guides/search-analytics/concepts/metrics/) report it as the no results rate

Semantic search drives it down by returning junk; read it with click-through

Search-exit rate

Share of searches after which the visitor leaves (my definition: no click, no refinement, session ends)

Needs your own event tracking; bots and bookmarks add noise

Click-through rate

Share of searches with at least one click on a result

Position bias: a better top result raises it, a bad one hides below the fold

Reformulation rate

Share of searches followed by another query in the same session

Some reformulation is healthy refinement

Rank-one accuracy for identifiers

Judged set: does the exact part come first

Must stay at 100%; any drop is a regression

The judged set is the part most teams skip. I take a few hundred real queries from the logs, stratified by the query types in the first table, and record which products are right. Identifier queries assert an exact rank one; descriptive queries assert that a relevant product is in the top ten. Run it on every change to analyzers, synonyms, embeddings or fusion settings, in CI if you can.

Then run an A/B test on live traffic and compare click-through, add-to-cart from search and search-exit rate per query type. A zero-result query is sometimes correct, because the part is not in the range, so do not chase the rate to zero; chase the number of zero-result queries that should have found something.

## Integrating with Spryker, Elasticsearch and OpenSearch

Spryker is [shipped with Elasticsearch as its default search](https://docs.spryker.com/docs/pbc/all/search/latest/base-shop/search-feature-overview/search-feature-overview), indexing product name, description and SKU, product attributes, reviews and CMS pages. The documentation also describes third-party search integrations and a tutorial for integrating any search engine, and [a migration path for OpenSearch](https://docs.spryker.com/docs/pbc/all/search/latest/base-shop/install-and-upgrade/migrate-from-opensearch-1.3-to-3.5.html) from 1.3 via 2.19 to 3.5. That upgrade matters here, because hybrid fusion in OpenSearch needs a recent version.

My suggested integration, which is a design proposal and not a Spryker feature, has four steps. Compute an embedding for each abstract product and locale in the publish-and-sync flow that builds the search documents, and store it in a vector field next to the text fields. Extend the search query so that it issues the identifier, lexical and vector clauses under the shop's existing filters, and fuse them with RRF. Put the LLM parser in front, behind a timeout. Wrap everything in a feature flag per store and locale.

Two caveats from experience with these platforms. Vectors increase index size and publish time, so size the cluster and re-embed only when the embedded text changes. And keep the existing search as a fallback path: if the vector lane is unavailable, the shop should still answer through the lexical and identifier lanes.

If you are building from scratch, a hosted search product is an alternative, and Spryker documents integrations for that route. I would still insist on the same four things from any vendor: identifier handling, entitlement pre-filters, per-language metrics and a way to run your judged set.

## A rollout checklist

This is the order I would work in.

1.  Export three months of search logs and classify queries into the types of the first table. Count them.
2.  Build the judged set, with identifier queries that assert rank one, and measure the current search as the baseline.
3.  Fix the lexical basics: identifier fields with a normaliser, per-language analyzers, a maintained synonym set.
4.  Add the vector lane with a multilingual model, entitlement pre-filters and RRF fusion, behind a feature flag.
5.  Short-circuit identifier hits so the vector arm and the reranker never touch them.
6.  Add the LLM parser with schema validation, an offline cache for frequent queries, a timeout and a boost-first policy.
7.  Add a reranker on a small window for non-identifier queries, and check latency at p95.
8.  Run an A/B test, compare metrics per query type and language, and review the zero-result log weekly.

Most of the gain in these projects comes from steps 3 and 4, not from the most fashionable model, and the part that earns trust is step 2. Make the identifier lane boringly reliable first, and semantic search will feel like an upgrade instead of a risk.

## Sources

1.  [Baymard Institute: E-commerce search query types](https://baymard.com/blog/ecommerce-search-query-types)
2.  [Elastic Search Labs: Hybrid search in Elasticsearch](https://www.elastic.co/search-labs/blog/hybrid-search-elasticsearch)
3.  [Elasticsearch documentation: kNN query (pre-filters and post-filters)](https://www.elastic.co/docs/reference/query-languages/query-dsl/query-dsl-knn-query)
4.  [Elasticsearch documentation: Semantic reranking](https://www.elastic.co/docs/solutions/search/ranking/semantic-reranking)
5.  [Elasticsearch documentation: Word delimiter graph token filter](https://www.elastic.co/docs/reference/text-analysis/analysis-word-delimiter-graph-tokenfilter)
6.  [Elasticsearch documentation: Synonym graph token filter](https://www.elastic.co/docs/reference/text-analysis/analysis-synonym-graph-tokenfilter)
7.  [OpenSearch documentation: Score ranker processor (RRF)](https://docs.opensearch.org/latest/search-plugins/search-pipelines/score-ranker-processor/)
8.  [OpenSearch documentation: Normalization processor](https://docs.opensearch.org/latest/search-plugins/search-pipelines/normalization-processor/)
9.  [Spryker documentation: Search feature overview](https://docs.spryker.com/docs/pbc/all/search/latest/base-shop/search-feature-overview/search-feature-overview)
10.  [Spryker documentation: Migrate from OpenSearch 1.3 to 3.5](https://docs.spryker.com/docs/pbc/all/search/latest/base-shop/install-and-upgrade/migrate-from-opensearch-1.3-to-3.5.html)
11.  [Instacart via ZenML: Rebuilding query understanding for e-commerce search with LLMs](https://www.zenml.io/llmops-database/rebuilding-query-understanding-for-e-commerce-search-with-llms)
12.  [arXiv: M3-Embedding, multilingual, multi-functionality, multi-granularity text embeddings](https://arxiv.org/abs/2402.03216)
13.  [Algolia documentation: Search analytics metrics](https://www.algolia.com/doc/guides/search-analytics/concepts/metrics/)

## Frequently asked questions

What is semantic search for B2B e-commerce?

Semantic search turns product texts and queries into vectors so that a request such as "screw for outdoor wood" finds matching products even without shared words. In a B2B shop it should complement, not replace, lexical search, because article numbers, EANs and exact specifications still need precise matching.

What is hybrid search and why do B2B shops need it?

Hybrid search runs a lexical query (BM25) and a vector query in parallel and fuses the two ranked lists, often with reciprocal rank fusion. B2B shops need it because their buyers mix exact identifiers with natural-language descriptions, and neither method alone handles both well.

How do I keep part-number search exact with vector search?

Index identifiers in dedicated fields with a normaliser that strips case, dashes, dots and spaces, match them exactly and by prefix first, and rank those hits above anything the vector arm returns. Do not rely on embeddings for identifiers, because they treat similar-looking numbers as similar.

Does semantic search work for German, English and Hungarian catalogues?

Yes, with a multilingual embedding model and per-language text analysis. Multilingual models such as BGE-M3 cover more than 100 languages, but you still need to test German compound words and Hungarian inflection on your own queries and measure results per language.

How do I measure whether product search improved?

Segment by query type and track zero-result rate, search-exit rate, click-through rate and add-to-cart from search. Add an offline set of judged queries, where every exact identifier must rank first. A falling zero-result rate alone can hide irrelevant results, so always read it next to click-through.

Can I add semantic search to Spryker?

Yes. Spryker ships with Elasticsearch as its default search and documents an OpenSearch upgrade path, so you can add vector fields to the product search documents at publish time and extend the search query with a vector clause. I would add this as a pilot behind a feature flag and compare it with the existing search.

Written by Balázs Csorba

Senior fullstack & AI engineer in Styria, Austria – 10+ years of Vue, Nuxt, Node.js and PHP, now building tooling for AI agents.

[B2B e-commerce & PIM →](https://balazscsorba.com/expertise/b2b-ecommerce-developer)[About me →](https://balazscsorba.com/about)

## More articles

-   [Reducing LLM hallucinations in production: grounding, citations and knowing when to say no](https://balazscsorba.com/blog/llm-hallucination-grounding-citations)
-   [pgvector or a vector database? How to choose vector storage in 2026](https://balazscsorba.com/blog/pgvector-vs-vector-databases)
-   [GraphRAG and knowledge-graph RAG: when a graph beats vector search](https://balazscsorba.com/blog/graphrag-knowledge-graph-rag)
-   [Evaluating RAG: retrieval metrics, faithfulness and how to tell which half failed](https://balazscsorba.com/blog/rag-evaluation-metrics)

## Sounds like what you need?

Tell me about your project or role – I’d love to hear from you.

[Book a call](mailto:contact@balazscsorba.com) [Connect on LinkedIn](https://www.linkedin.com/in/balazs-csorba)
