[{"data":1,"prerenderedAt":1027},["ShallowReactive",2],{"blog-llm-hallucination-grounding-citations-en":3},{"slug":4,"published":5,"minutes":6,"category":7,"tags":8,"keywords":14,"about":25,"sources":35,"cover":75,"og":76,"expertise":77,"locales":78,"lang":79,"title":82,"description":83,"coverAlt":84,"metaTitle":85,"takeaways":86,"faq":92,"toc":111,"blocks":148,"others":736},"llm-hallucination-grounding-citations","2026-10-02",13,"rag",[9,10,11,12,13],"Hallucinations","RAG","Citations","Grounding","Faithfulness",[15,16,17,18,19,20,21,22,23,24],"reduce LLM hallucinations in production","how to reduce hallucinations in RAG","LLM citations API","Anthropic citations API","RAG faithfulness metric","LLM abstention I don't know","claim-level verification LLM","grounding LLM answers in sources","check grounding API","show sources in AI chatbot UI",[26,29,32],{"name":27,"url":28},"Hallucination (artificial intelligence)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FHallucination_(artificial_intelligence)",{"name":30,"url":31},"Retrieval-augmented generation","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FRetrieval-augmented_generation",{"name":33,"url":34},"Large language model","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FLarge_language_model",[36,39,42,45,48,51,54,57,60,63,66,69,72],{"title":37,"url":38},"Anthropic: Citations (Claude API documentation)","https:\u002F\u002Fplatform.claude.com\u002Fdocs\u002Fen\u002Fbuild-with-claude\u002Fcitations",{"title":40,"url":41},"Anthropic: Search results (Claude API documentation)","https:\u002F\u002Fplatform.claude.com\u002Fdocs\u002Fen\u002Fbuild-with-claude\u002Fsearch-results",{"title":43,"url":44},"Anthropic: Reduce hallucinations (Claude API documentation)","https:\u002F\u002Fplatform.claude.com\u002Fdocs\u002Fen\u002Ftest-and-evaluate\u002Fstrengthen-guardrails\u002Freduce-hallucinations",{"title":46,"url":47},"Anthropic: Introducing Citations on the Anthropic API","https:\u002F\u002Fclaude.com\u002Fblog\u002Fintroducing-citations-api",{"title":49,"url":50},"Simon Willison: Anthropic's new Citations API (24 January 2025)","https:\u002F\u002Fsimonwillison.net\u002F2025\u002FJan\u002F24\u002Fanthropics-new-citations-api\u002F",{"title":52,"url":53},"OpenAI: Web search guide (url_citation annotations and display requirement)","https:\u002F\u002Fdevelopers.openai.com\u002Fapi\u002Fdocs\u002Fguides\u002Ftools-web-search",{"title":55,"url":56},"Cohere: Documents and citations","https:\u002F\u002Fdocs.cohere.com\u002Fdocs\u002Fdocuments-and-citations",{"title":58,"url":59},"Google Cloud: Check grounding API","https:\u002F\u002Fdocs.cloud.google.com\u002Fgenerative-ai-app-builder\u002Fdocs\u002Fcheck-grounding",{"title":61,"url":62},"AWS: Amazon Bedrock Guardrails contextual grounding check","https:\u002F\u002Fdocs.aws.amazon.com\u002Fbedrock\u002Flatest\u002Fuserguide\u002Fguardrails-contextual-grounding-check.html",{"title":64,"url":65},"Ragas: Faithfulness metric","https:\u002F\u002Fdocs.ragas.io\u002Fen\u002Fstable\u002Fconcepts\u002Fmetrics\u002Favailable_metrics\u002Ffaithfulness\u002F",{"title":67,"url":68},"Kalai, Nachum, Vempala, Zhang: Why Language Models Hallucinate (arXiv 2509.04664)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2509.04664",{"title":70,"url":71},"Magesh et al.: Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools (arXiv 2405.20362)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2405.20362",{"title":73,"url":74},"Wallat, Heuss, de Rijke, Anand: Correctness is not Faithfulness in RAG Attributions (arXiv 2412.18004)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2412.18004","\u002Fimages\u002Fblog\u002Fllm-hallucination-grounding-citations\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fllm-hallucination-grounding-citations\u002Fog.jpg","ai-engineer",[79,80,81],"en","de","hu","Reducing LLM hallucinations in production: grounding, citations and knowing when to say no","Cut hallucinations in production RAG: citation APIs, abstention, claim-level checks, faithfulness metrics, source UI, and the failures that still slip through.","Diagram: a retrieval step feeds an evidence gate, a cited answer and a claim verifier, ending in an answer with sources, with abstain and flag paths branching off.","LLM hallucinations: grounding and citations · Balázs Csorba",[87,88,89,90,91],"Retrieval does not remove hallucinations. A Stanford study found leading RAG legal tools still hallucinated between 17% and 33% of the time, often by citing a real source that does not support the claim.","Treat grounding as a pipeline, not a prompt: an evidence gate before generation, native citations during generation, claim-level verification after, and a UI that shows the evidence.","Abstention is a product state, not an error message. Design the \"I cannot answer this from the documents\" path, measure it, and give it a next step.","Citation APIs make pointers valid, not conclusions true. Anthropic guarantees valid document pointers; whether the passage supports the claim still has to be checked, because up to 57% of citations in one study were post-rationalised.","Measure faithfulness (claims supported by retrieved context) separately from answer correctness, on a test set that includes questions the documents cannot answer.",[93,96,99,102,105,108],{"q":94,"a":95},"How do you reduce hallucinations in a RAG application?","Stack several layers instead of relying on one. Improve retrieval first, then restrict the model to the provided documents, allow it to say it does not know, make it cite the passages it used, verify each claim against those passages, and show the sources in the UI. No single layer eliminates hallucinations; Anthropic's own guidance says these techniques reduce them significantly but do not remove them.",{"q":97,"a":98},"Does RAG eliminate hallucinations?","No. In the Stanford and Yale study of commercial legal research tools, products that marketed RAG as a fix still produced hallucinated answers between 17% and 33% of the time. The study counts an answer as hallucinated when it is incorrect or misgrounded, meaning it claims a source supports something it does not.",{"q":100,"a":101},"What is the difference between faithfulness and correctness?","Faithfulness asks whether every claim in the answer is supported by the retrieved context. Correctness asks whether the claim is true in the world. A faithful answer built on an outdated document is wrong but faithful; a correct answer from the model's memory that the documents do not support is correct but unfaithful. In RAG you usually want both, and you measure them separately.",{"q":103,"a":104},"How do the Anthropic citations work?","You pass documents or search_result blocks with citations enabled, and the response text blocks carry citation objects pointing to character ranges, page numbers or content blocks in your sources. The cited_text field does not count toward output tokens, and the API guarantees the pointers are valid. Citations cannot be combined with structured outputs and currently cover text only.",{"q":106,"a":107},"When should an LLM say \"I don't know\"?","When the retrieved evidence does not contain the answer. Implement it in two places: a retrieval gate that stops generation when the best passages are weak, and a prompt that explicitly allows the model to state that the documents lack the information. Then test it with questions your corpus cannot answer, otherwise the behaviour is never verified.",{"q":109,"a":110},"How should a chatbot show its sources?","Put numbered markers next to the claims they support, show the cited passage on hover or tap, link to the original document at the right location, and mark answers or sentences that could not be verified. Never show a citation that has not been checked to point at real text, because a confident-looking footnote makes a wrong answer more believable.",[112,115,118,121,124,127,130,133,136,139,142,145],{"id":113,"title":114},"what-hallucination-means","What counts as a hallucination in a RAG system",{"id":116,"title":117},"pipeline","The pipeline: four gates, not one prompt",{"id":119,"title":120},"grounding-prompts","Step one: constrain the model to what you gave it",{"id":122,"title":123},"citation-apis","Step two: use a citation API instead of asking nicely",{"id":125,"title":126},"abstention","Step three: design the \"I cannot answer that\" path",{"id":128,"title":129},"claim-verification","Step four: verify claims after generation",{"id":131,"title":132},"techniques-table","Technique versus effect",{"id":134,"title":135},"measuring","Measuring faithfulness",{"id":137,"title":138},"ui-patterns","Showing sources in the interface",{"id":140,"title":141},"what-still-slips","Where hallucination still slips through",{"id":143,"title":144},"checklist","A checklist you can start with this week",{"id":146,"title":147},"sources","Sources",[149,153,167,170,173,188,191,194,195,198,207,230,233,234,237,259,262,265,266,269,272,275,305,308,311,343,346,348,351,354,355,358,364,370,373,383,384,387,390,420,423,445,446,449,532,533,539,561,564,571,572,575,607,610,613,614,617,659,667,668,686,693,694],{"type":150,"content":151},"paragraph",[152],"Every team that ships a retrieval-augmented assistant goes through the same stages. First comes the demo, where it answers beautifully. Then comes the first real user, who finds a confident, fluent, wrong answer on day two. Then someone says \"we already use RAG, why does it still make things up?\"",{"type":150,"content":154},[155,156,161,162,166],"The honest answer is that retrieval changes the odds, not the nature of the system. A language model still produces plausible text; retrieval only gives it better material and a chance to show its work. This article is how I would build the part around the model so that wrong answers become rarer, visible and recoverable. It builds on my posts about ",{"tag":157,"to":158,"children":159},"link","\u002Fblog\u002Frag-pipeline-chunking-hybrid-search-reranking",[160],"chunking, hybrid search and reranking"," and ",{"tag":157,"to":163,"children":164},"\u002Fblog\u002Fllm-evals-for-product-features",[165],"evals for product features","; here the focus is what happens after the passages are retrieved.",{"type":168,"level":169,"id":113,"text":114},"heading",2,{"type":150,"content":171},[172],"In a closed-book chatbot, a hallucination is a false statement. In a RAG system there are two distinct failures, and they need different fixes:",{"type":174,"ordered":175,"items":176},"list",false,[177,183],[178,182],{"tag":179,"children":180},"strong",[181],"Wrong content:"," the answer states something false, either because retrieval returned the wrong or outdated passage, or because the model ignored the context and answered from memory.",[184,187],{"tag":179,"children":185},[186],"Misgrounded content:"," the answer cites a source, but the source does not say that. The Stanford and Yale team that evaluated commercial legal research tools defines it precisely: a response is hallucinated if it is incorrect or misgrounded, meaning the answer \"falsely asserts that a source supports a statement\".",{"type":150,"content":189},[190],"The second type is the dangerous one. The same study tested tools from LexisNexis and Thomson Reuters that were marketed with claims like \"hallucination-free\" citations, and found they hallucinated between 17% and 33% of the time. The authors note that checking such errors means clicking through, reading the source and comparing it to the claim, which is exactly the work users expect the tool to have done for them.",{"type":150,"content":192},[193],"Why do models guess at all? The paper \"Why Language Models Hallucinate\" argues that training and evaluation procedures reward guessing over acknowledging uncertainty, because a model that guesses scores better on benchmarks than one that abstains. That matters for you in a practical way: the default behaviour of the model is to answer, so abstention has to be designed into the system, not hoped for.",{"type":168,"level":169,"id":116,"text":117},{"type":150,"content":196},[197],"I think of grounding as a sequence of gates. Each one catches something the previous one cannot, and each one has a cost in latency and complexity.",{"type":199,"attrs":200,"inner":204,"caption":205},"diagram",{"viewBox":201,"role":202,"aria-labelledby":203},"0 0 752 370","img","d1-hall-t d1-hall-d","\u003Ctitle id=\"d1-hall-t\">A grounded answer pipeline\u003C\u002Ftitle>\u003Cdesc id=\"d1-hall-d\">Five steps in a row: retrieve, gate, generate with citations, verify claims, show the answer with sources. From the gate, a dashed path leads to abstaining when the evidence is too weak. From the verifier, a dashed path leads to dropping or flagging unsupported claims. A bar below says every step is logged and feeds evaluations.\u003C\u002Fdesc>\u003Ctext x=\"20\" y=\"28\" class=\"d-title\">A grounded answer pipeline\u003C\u002Ftext>\u003Crect x=\"20\" y=\"64\" width=\"122\" height=\"76\" rx=\"10\" class=\"d-sky\" \u002F>\u003Ctext x=\"81\" y=\"98\" text-anchor=\"middle\" class=\"d-text\">Retrieve\u003C\u002Ftext>\u003Ctext x=\"81\" y=\"119\" text-anchor=\"middle\" class=\"d-small\">hybrid + rerank\u003C\u002Ftext>\u003Cpath d=\"M142 102 H154\" class=\"d-line\" \u002F>\u003Cpath d=\"M162 102 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Crect x=\"162\" y=\"64\" width=\"122\" height=\"76\" rx=\"10\" class=\"d-accent\" \u002F>\u003Ctext x=\"223\" y=\"98\" text-anchor=\"middle\" class=\"d-text\">Gate\u003C\u002Ftext>\u003Ctext x=\"223\" y=\"119\" text-anchor=\"middle\" class=\"d-small\">enough evidence?\u003C\u002Ftext>\u003Cpath d=\"M284 102 H296\" class=\"d-line\" \u002F>\u003Cpath d=\"M304 102 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Crect x=\"304\" y=\"64\" width=\"122\" height=\"76\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"365\" y=\"98\" text-anchor=\"middle\" class=\"d-text\">Generate\u003C\u002Ftext>\u003Ctext x=\"365\" y=\"119\" text-anchor=\"middle\" class=\"d-small\">with citations\u003C\u002Ftext>\u003Cpath d=\"M426 102 H438\" class=\"d-line\" \u002F>\u003Cpath d=\"M446 102 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Crect x=\"446\" y=\"64\" width=\"122\" height=\"76\" rx=\"10\" class=\"d-accent\" \u002F>\u003Ctext x=\"507\" y=\"98\" text-anchor=\"middle\" class=\"d-text\">Verify\u003C\u002Ftext>\u003Ctext x=\"507\" y=\"119\" text-anchor=\"middle\" class=\"d-small\">claim by claim\u003C\u002Ftext>\u003Cpath d=\"M568 102 H580\" class=\"d-line\" \u002F>\u003Cpath d=\"M588 102 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Crect x=\"588\" y=\"64\" width=\"122\" height=\"76\" rx=\"10\" class=\"d-mint\" \u002F>\u003Ctext x=\"649\" y=\"98\" text-anchor=\"middle\" class=\"d-text\">Show\u003C\u002Ftext>\u003Ctext x=\"649\" y=\"119\" text-anchor=\"middle\" class=\"d-small\">answer + sources\u003C\u002Ftext>\u003Cpath d=\"M223 140 V182\" class=\"d-line d-dash\" \u002F>\u003Cpath d=\"M223 190 l-5 -9 h10 z\" class=\"d-head\" \u002F>\u003Crect x=\"154\" y=\"190\" width=\"138\" height=\"56\" rx=\"10\" class=\"d-gold\" \u002F>\u003Ctext x=\"223\" y=\"214\" text-anchor=\"middle\" class=\"d-text\">Abstain\u003C\u002Ftext>\u003Ctext x=\"223\" y=\"235\" text-anchor=\"middle\" class=\"d-small\">say what is missing\u003C\u002Ftext>\u003Cpath d=\"M507 140 V182\" class=\"d-line d-dash\" \u002F>\u003Cpath d=\"M507 190 l-5 -9 h10 z\" class=\"d-head\" \u002F>\u003Crect x=\"438\" y=\"190\" width=\"138\" height=\"56\" rx=\"10\" class=\"d-gold\" \u002F>\u003Ctext x=\"507\" y=\"214\" text-anchor=\"middle\" class=\"d-text\">Drop or flag\u003C\u002Ftext>\u003Ctext x=\"507\" y=\"235\" text-anchor=\"middle\" class=\"d-small\">unsupported claims\u003C\u002Ftext>\u003Crect x=\"20\" y=\"276\" width=\"712\" height=\"46\" rx=\"10\" class=\"d-box d-dash\" \u002F>\u003Ctext x=\"376\" y=\"304\" text-anchor=\"middle\" class=\"d-small\">Log query, chunks, answer and verdicts: they become your eval set\u003C\u002Ftext>\u003Ctext x=\"376\" y=\"350\" text-anchor=\"middle\" class=\"d-label\">Each gate removes a different failure mode.\u003C\u002Ftext>",[206],"Grounding is a chain of checks. Each stage can stop or downgrade the answer, and each stage is measured.",{"type":174,"ordered":208,"items":209},true,[210,215,220,225],[211,214],{"tag":179,"children":212},[213],"Retrieve well."," Hybrid search and reranking decide whether the right passage even reaches the model. Most \"hallucinations\" I debug are retrieval misses where the model did what it could with the wrong material.",[216,219],{"tag":179,"children":217},[218],"Gate on evidence."," If the best passages are weak, do not generate; abstain and say what is missing.",[221,224],{"tag":179,"children":222},[223],"Generate with citations."," Use a native citation mechanism so every claim carries a pointer into a document you supplied.",[226,229],{"tag":179,"children":227},[228],"Verify claim by claim."," Check that each cited passage supports the sentence it is attached to, and drop or flag what fails.",{"type":150,"content":231},[232],"Then comes the display step, which is also a control: the interface decides whether a reader can check the answer in five seconds or has to trust it.",{"type":168,"level":169,"id":119,"text":120},{"type":150,"content":235},[236],"Anthropic's guide to reducing hallucinations lists a handful of basic techniques, and they are cheap enough that I would use all of them by default:",{"type":174,"ordered":175,"items":238},[239,244,249,254],[240,243],{"tag":179,"children":241},[242],"Allow \"I don't know\"."," Explicitly give the model permission to admit uncertainty. Anthropic says this simple technique can drastically reduce false information.",[245,248],{"tag":179,"children":246},[247],"Quote first for long documents."," For documents over roughly 20,000 tokens, ask the model to extract word-for-word quotes first and base its answer on those quotes only.",[250,253],{"tag":179,"children":251},[252],"Restrict external knowledge."," Instruct the model to use only the provided documents and not its general knowledge.",[255,258],{"tag":179,"children":256},[257],"Verify after drafting."," Ask the model to find a supporting quote for each claim and to retract any claim it cannot support.",{"type":150,"content":260},[261],"The same guide lists best-of-N comparison (run the prompt several times and treat disagreement as a warning) and iterative refinement as advanced options, and it ends with a caveat I want to repeat: these techniques significantly reduce hallucinations but do not eliminate them, and critical information still needs validation.",{"type":150,"content":263},[264],"My practical addition: treat prompt rules as a weak layer. A prompt asks the model to behave; a gate or a verifier checks that it did. Use the prompt to raise the baseline and the later stages to catch the remainder.",{"type":168,"level":169,"id":122,"text":123},{"type":150,"content":267},[268],"You can prompt a model to write \"[1]\" after sentences, and for prototypes that works. In production, the pointer is the problem: models invent plausible source names, mis-number them, or quote text that is not in the document. Native citation features move that work out of free text and into the API response.",{"type":168,"level":270,"text":271},3,"Anthropic: citations and search results",{"type":150,"content":273},[274],"Anthropic's citations feature works on three document types, and the citation format follows the type:",{"type":276,"head":277,"rows":284},"table",[278,280,282],[279],"Document type",[281],"Chunking",[283],"Citation points to",[285,292,298],[286,288,290],[287],"Plain text",[289],"Sentences",[291],"Character indices (0-indexed)",[293,295,296],[294],"PDF",[289],[297],"Page numbers (1-indexed)",[299,301,303],[300],"Custom content",[302],"None added: your blocks are used as-is",[304],"Block indices (0-indexed)",{"type":150,"content":306},[307],"For RAG, the documentation's own advice is to put each retrieved chunk into a plain text document, or to use `search_result` content blocks, which carry a source and a title and can be returned from your own search tools or placed directly in the user message. Citations then appear on the text blocks that draw on your content, without special prompting. In my experience this chunk-per-document approach is also the cleanest way to keep your own chunk IDs traceable in the answer.",{"type":150,"content":309},[310],"The details that matter in production, all from the documentation:",{"type":174,"ordered":175,"items":312},[313,318,323,333,338],[314,317],{"tag":179,"children":315},[316],"Valid pointers."," Because the API parses citations and extracts `cited_text` itself, citations are guaranteed to contain valid pointers to the documents you provided. That removes fabricated references, not misread ones.",[319,322],{"tag":179,"children":320},[321],"Cost."," Enabling citations slightly increases input tokens, but `cited_text` does not count toward output tokens, so it can be cheaper than prompting the model to quote.",[324,327,328,332],{"tag":179,"children":325},[326],"Caching."," The source documents can be cached with `cache_control`; the citation blocks in responses cannot. See my post on ",{"tag":157,"to":329,"children":330},"\u002Fblog\u002Fllm-cost-latency-prompt-caching-routing",[331],"prompt caching and routing"," for when this pays off.",[334,337],{"tag":179,"children":335},[336],"Streaming."," Citations arrive as `citations_delta` events, one citation per event, attached to the current text block.",[339,342],{"tag":179,"children":340},[341],"Limits."," Citations must be enabled on all or none of the documents in a request, only text is citable (scanned PDFs without extractable text are not), and combining citations with structured outputs returns a 400 error.",{"type":150,"content":344},[345],"When the feature launched, Anthropic reported that its internal evaluations showed built-in citations outperforming most custom implementations by up to 15% in recall accuracy, and a customer, Endex, said source hallucinations and formatting issues fell from 10% to 0%. Those are vendor-reported figures for specific setups; I would use them as a reason to test the feature, not as a number to expect.",{"type":168,"level":270,"text":347},"Other providers",{"type":150,"content":349},[350],"Anthropic is not alone. OpenAI's web search tool returns `url_citation` annotations with a URL, title and location in the text, and its documentation requires that inline citations be clearly visible and clickable in the user interface when you show web results. Cohere's chat API returns citation objects with start and end positions, the cited text and the source documents behind it. The shared idea is the same: the model's claim and its evidence travel together as structured data, and your UI has to respect that.",{"type":150,"content":352},[353],"One structural caveat applies to all of them: the model still decides which passage to point to. A citation tells you what the model attached to a sentence, not that the sentence follows from it.",{"type":168,"level":169,"id":125,"text":126},{"type":150,"content":356},[357],"If the model's default is to answer, you need two places to interrupt that default.",{"type":150,"content":359},[360,363],{"tag":179,"children":361},[362],"A retrieval gate before generation."," Look at the retrieval result: no passages above a similarity or reranker threshold, a large gap between what was asked and what was found, or contradictory top passages. In these cases skip the model call entirely and respond with what is missing and what the user can do. This is cheaper than generating and also the most reliable abstention, because it does not depend on the model's self-assessment. Calibrate the threshold on real queries, not by feel.",{"type":150,"content":365},[366,369],{"tag":179,"children":367},[368],"A permission in the prompt."," Tell the model that stating \"the documents do not contain this\" is a valid and preferred answer when the evidence is missing. Without that permission, the benchmark-style incentive to guess is still operating.",{"type":150,"content":371},[372],"Design the abstention response as part of the product:",{"type":174,"ordered":175,"items":374},[375,377,379,381],[376],"Say what was searched and what was not found, instead of a bare \"I don't know\".",[378],"Offer a next step: rephrase, narrow the scope, search another source, or hand over to a person.",[380],"Allow partial answers: answer the supported part and mark the unsupported part explicitly.",[382],"Count it. The abstention rate is a metric with a healthy range; zero means the system is guessing, and too high means the gate is too strict or retrieval is poor.",{"type":168,"level":169,"id":128,"text":129},{"type":150,"content":385},[386],"The most reliable pattern I know for catching misgrounded answers is to decompose the answer into claims and check each one against its evidence. It is the same idea as the faithfulness metric in Ragas, which splits a response into individual statements, checks whether each can be inferred from the retrieved context, and computes the share of supported claims.",{"type":150,"content":388},[389],"You can build this yourself with a second model call, or use a managed checker:",{"type":276,"head":391,"rows":398},[392,394,396],[393],"Option",[395],"What it does",[397],"Notable limits (per documentation)",[399,406,413],[400,402,404],[401],"Prompted self-check (Anthropic guide)",[403],"Model finds a supporting quote per claim, retracts the rest",[405],"Same model family judging its own draft; extra call",[407,409,411],[408],"Google check grounding API",[410],"Splits the answer into claims, returns a 0 to 1 support score, citations and optional per-claim scores",[412],"Answer up to 4,096 tokens, up to 200 facts; partial truths count as ungrounded; documented as under 500 ms",[414,416,418],[415],"Amazon Bedrock contextual grounding check",[417],"Scores grounding and relevance against a source and query; blocks below your threshold",[419],"Not for conversational QA; source up to 100,000 characters; on streaming, irrelevance may only be flagged after the response is sent",{"type":150,"content":421},[422],"Three design decisions come up every time:",{"type":174,"ordered":208,"items":424},[425,430,440],[426,429],{"tag":179,"children":427},[428],"What happens on failure?"," Options are to regenerate with the failing claim removed, to drop the sentence, to keep it with an \"unverified\" marker, or to abstain on the whole answer. For high-stakes domains I prefer dropping or marking over silent regeneration, because users should see that something was removed.",[431,434,435,439],{"tag":179,"children":432},[433],"Where does it run?"," A blocking verifier adds latency before the first token is shown if you wait for it. For streaming UIs, stream the draft with citations and update each sentence's state as verdicts arrive, or verify before streaming for the few flows where wrong answers are costly. My post on ",{"tag":157,"to":436,"children":437},"\u002Fblog\u002Fnuxt-llm-features-ai-sdk-streaming",[438],"streaming LLM features in Nuxt"," covers the transport side.",[441,444],{"tag":179,"children":442},[443],"Who checks the checker?"," A verifier is a model or a classifier and it makes mistakes. Label a few hundred verdicts by hand and track the verifier's own precision and recall, otherwise you have moved trust from one unmeasured component to another.",{"type":168,"level":169,"id":131,"text":132},{"type":150,"content":447},[448],"This is how I rank the techniques by what they actually address. I deliberately give no percentage per row: the effect depends on your corpus, your queries and your model, and the only numbers I would trust are the ones you measure on your own test set.",{"type":276,"head":450,"rows":459},[451,453,455,457],[452],"Technique",[454],"Failure it targets",[456],"Cost",[458],"Where it still fails",[460,469,478,487,496,505,514,523],[461,463,465,467],[462],"Better retrieval (hybrid, rerank)",[464],"Wrong or missing passages",[466],"Engineering time; some latency",[468],"Corpus gaps, stale documents",[470,472,474,476],[471],"\"Only use the documents\" prompt rules",[473],"Answers from model memory",[475],"Almost none",[477],"A prompt is a request, not a guarantee",[479,481,483,485],[480],"Permission to say \"I don't know\"",[482],"Forced guessing",[484],"None",[486],"Over-abstention if not tested",[488,490,492,494],[489],"Retrieval evidence gate",[491],"Generating from weak evidence",[493],"One threshold to calibrate",[495],"Strong but irrelevant passages pass it",[497,499,501,503],[498],"Quote-first extraction",[500],"Paraphrase drift on long documents",[502],"Extra tokens or a second step",[504],"Quotes can still be misread",[506,508,510,512],[507],"Native citations",[509],"Fabricated or invalid references",[511],"Slightly more input tokens",[513],"Valid pointer, unsupported claim",[515,517,519,521],[516],"Claim-level verification",[518],"Misgrounded claims",[520],"Extra call and latency",[522],"Verifier errors; multi-hop reasoning",[524,526,528,530],[525],"Source-first UI",[527],"Unchecked trust",[529],"Design and frontend work",[531],"Users who never click",{"type":168,"level":169,"id":134,"text":135},{"type":150,"content":534},[535,536,538],"Without measurement, every change in this article is a belief. I would track four numbers on a fixed evaluation set, and run them in CI the way you run unit tests (see ",{"tag":157,"to":163,"children":537},[165],"):",{"type":174,"ordered":175,"items":540},[541,546,551,556],[542,545],{"tag":179,"children":543},[544],"Faithfulness:"," the share of claims supported by the retrieved context, as in the Ragas definition. It is separate from correctness: a faithful answer from a wrong document is wrong.",[547,550],{"tag":179,"children":548},[549],"Citation precision and recall:"," of the citations shown, how many truly support the sentence; of the claims made, how many have a supporting citation.",[552,555],{"tag":179,"children":553},[554],"Abstention quality:"," on questions the corpus cannot answer, how often does the system decline; on answerable questions, how often does it decline wrongly.",[557,560],{"tag":179,"children":558},[559],"Answer correctness"," against a reference answer, so that faithfulness is not your only signal.",{"type":150,"content":562},[563],"Build the set from three groups: answerable questions with known passages, unanswerable questions, and adversarial ones (outdated information, near-duplicate documents, questions that tempt the model to use its own knowledge). Most teams skip the unanswerable group, and that is why their abstention behaviour is never tested.",{"type":565,"variant":566,"title":567,"body":568},"callout","warn","A citation is not proof of faithfulness",[569],[570],"Research on attribution in RAG distinguishes citation correctness from faithfulness. In the Wallat et al. study, up to 57% of citations were post-rationalised: the model had already decided its answer from prior knowledge and attached a document that happened to agree. A citation that matches the claim can still be decoration. Check whether the answer changes when the cited passage is removed, at least on a sample.",{"type":168,"level":169,"id":137,"text":138},{"type":150,"content":573},[574],"The interface is the last line of defence and the only one the user sees. These are the patterns I would use:",{"type":174,"ordered":175,"items":576},[577,582,587,592,597,602],[578,581],{"tag":179,"children":579},[580],"Inline numbered markers"," next to the sentence they support, not one block of links at the end. Per-claim attachment is what makes checking possible.",[583,586],{"tag":179,"children":584},[585],"Preview on hover or tap"," showing the cited passage, with the document title. This is where `cited_text` is useful, and it makes a five-second check realistic.",[588,591],{"tag":179,"children":589},[590],"Deep links"," that open the source at the cited location (page number, anchor or highlighted range), not just at the top of a 60-page PDF.",[593,596],{"tag":179,"children":594},[595],"Visible verification state"," per sentence or per answer: verified, unverified, or removed. If the verifier dropped something, say so.",[598,601],{"tag":179,"children":599},[600],"A designed abstention state"," with next steps, as above, styled as a normal outcome rather than an error.",[603,606],{"tag":179,"children":604},[605],"Citations that are actually clickable and visible."," OpenAI's documentation makes this a requirement for web results, and it is a good rule everywhere.",{"type":150,"content":608},[609],"One warning from the Stanford study applies to design: real, authoritative-looking citations make a wrong answer more convincing. Do not let the presence of a footnote signal more certainty than your verification supports. If your verifier did not run, do not render the \"verified\" badge.",{"type":150,"content":611},[612],"Also mind the engineering side: responses with citations are no longer a plain text stream. Simon Willison pointed out when the feature launched that this forces an abstraction for responses that are annotated chunks rather than text. Plan your streaming protocol and your message storage for structured segments from the start.",{"type":168,"level":169,"id":140,"text":141},{"type":150,"content":615},[616],"After all four gates, these are the failures I would still expect, and what I would do about each:",{"type":174,"ordered":175,"items":618},[619,624,629,634,639,649,654],[620,623],{"tag":179,"children":621},[622],"Retrieval misses that look like answers."," A partially relevant passage passes the gate and the model fills the gap. Mitigation: reranker thresholds and an explicit \"does this passage answer the question\" check.",[625,628],{"tag":179,"children":626},[627],"Wrong or stale sources."," The answer is faithful to an outdated document. Mitigation: document dates and versions in the metadata, shown in the UI, and freshness filters.",[630,633],{"tag":179,"children":631},[632],"Misgrounded citations."," The valid pointer, unsupported claim problem. Mitigation: claim-level verification and sampled human review.",[635,638],{"tag":179,"children":636},[637],"Reasoning across passages."," Totals, comparisons and multi-hop conclusions are not literally in any one passage, so verifiers struggle with them. Mitigation: compute numbers with code and show the inputs.",[640,643,644,648],{"tag":179,"children":641},[642],"Poisoned documents."," If a retrieved document contains instructions, the model may follow them. Grounding on untrusted text is a security question; see ",{"tag":157,"to":645,"children":646},"\u002Fblog\u002Fprompt-injection-lethal-trifecta-patterns",[647],"prompt injection and the lethal trifecta",".",[650,653],{"tag":179,"children":651},[652],"Format constraints."," Structured outputs and citations cannot be combined on the Anthropic API today, so an extraction pipeline needs a different grounding strategy, for example a verification pass over the JSON values.",[655,658],{"tag":179,"children":656},[657],"Verifier blind spots."," A verifier can be wrong in both directions. Measure it.",{"type":150,"content":660},[661,662,666],"The longer-term direction, covered in my post on ",{"tag":157,"to":663,"children":664},"\u002Fblog\u002Frag-2026-hybrid-agentic-long-context",[665],"hybrid, agentic and long-context RAG",", does not remove this problem either. An agent that retrieves several times has more chances to find the right passage and more chances to chain a wrong inference.",{"type":168,"level":169,"id":143,"text":144},{"type":174,"ordered":208,"items":669},[670,672,674,676,678,680,682,684],[671],"Add a \"the documents do not contain this\" instruction and test it with at least 20 unanswerable questions.",[673],"Add a retrieval gate with a threshold calibrated on real queries; log every abstention.",[675],"Turn on native citations; put each chunk in its own document or `search_result` block with your own ID in the source field.",[677],"Store the answer, the cited passages and the retrieval scores for every response.",[679],"Add a claim-level verifier on a sample first, then on the flows where wrong answers cost money or trust.",[681],"Track faithfulness, citation precision and recall, and abstention quality in CI on a fixed evaluation set.",[683],"Show numbered, clickable citations with passage previews and a visible verification state.",[685],"Review a random sample of answers with citations by hand every week, because citations can be post-rationalised.",{"type":150,"content":687},[688,689,648],"If you take only one thing from this: make wrong answers cheap to notice. A system that is occasionally wrong and shows its evidence is a tool; one that is occasionally wrong and sounds certain is a liability. If you need help building or auditing this kind of pipeline, see my ",{"tag":157,"to":690,"children":691},"\u002Fexpertise\u002Fai-engineer",[692],"AI engineering work",{"type":168,"level":169,"id":146,"text":147},{"type":174,"ordered":208,"items":695},[696,700,703,706,709,712,715,718,721,724,727,730,733],[697],{"tag":698,"href":38,"children":699},"a",[37],[701],{"tag":698,"href":41,"children":702},[40],[704],{"tag":698,"href":44,"children":705},[43],[707],{"tag":698,"href":47,"children":708},[46],[710],{"tag":698,"href":50,"children":711},[49],[713],{"tag":698,"href":53,"children":714},[52],[716],{"tag":698,"href":56,"children":717},[55],[719],{"tag":698,"href":59,"children":720},[58],[722],{"tag":698,"href":62,"children":723},[61],[725],{"tag":698,"href":65,"children":726},[64],[728],{"tag":698,"href":68,"children":729},[67],[731],{"tag":698,"href":71,"children":732},[70],[734],{"tag":698,"href":74,"children":735},[73],[737,821,888,955],{"slug":738,"published":5,"minutes":6,"category":7,"tags":739,"keywords":744,"about":755,"sources":763,"cover":815,"og":816,"expertise":77,"locales":817,"lang":79,"title":818,"description":819,"coverAlt":820},"pgvector-vs-vector-databases",[740,741,10,742,743],"pgvector","Vector databases","Hybrid search","EU hosting",[745,746,747,748,749,750,751,752,753,754],"pgvector vs vector database","pgvector vs Qdrant","pgvector vs Pinecone","best vector database 2026","pgvector HNSW iterative scan","pgvector halfvec","OpenSearch vs Elasticsearch vector search","vector database EU hosting","hybrid search Postgres","Weaviate vs Milvus",[756,759,762],{"name":757,"url":758},"Vector database","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FVector_database",{"name":760,"url":761},"PostgreSQL","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FPostgreSQL",{"name":30,"url":31},[764,767,770,773,776,779,782,785,788,791,794,797,800,803,806,809,812],{"title":765,"url":766},"pgvector README (index limits, HNSW defaults, iterative scans, filtering, halfvec)","https:\u002F\u002Fgithub.com\u002Fpgvector\u002Fpgvector\u002Fblob\u002Fmaster\u002FREADME.md",{"title":768,"url":769},"pgvector CHANGELOG (0.4.0 to 0.8.7)","https:\u002F\u002Fgithub.com\u002Fpgvector\u002Fpgvector\u002Fblob\u002Fmaster\u002FCHANGELOG.md",{"title":771,"url":772},"pgvectorscale: StreamingDiskANN, statistical binary quantization, filtered search","https:\u002F\u002Fgithub.com\u002Ftimescale\u002Fpgvectorscale",{"title":774,"url":775},"Qdrant documentation: Filtering","https:\u002F\u002Fqdrant.tech\u002Fdocumentation\u002Fconcepts\u002Ffiltering\u002F",{"title":777,"url":778},"Qdrant documentation: Hybrid queries","https:\u002F\u002Fqdrant.tech\u002Fdocumentation\u002Fconcepts\u002Fhybrid-queries\u002F",{"title":780,"url":781},"Qdrant documentation: Create a cluster (providers, free tier, Hybrid Cloud)","https:\u002F\u002Fqdrant.tech\u002Fdocumentation\u002Fcloud\u002Fcreate-cluster\u002F",{"title":783,"url":784},"Weaviate documentation: Hybrid search","https:\u002F\u002Fdocs.weaviate.io\u002Fweaviate\u002Fconcepts\u002Fsearch\u002Fhybrid-search",{"title":786,"url":787},"Weaviate documentation: Vector index types","https:\u002F\u002Fdocs.weaviate.io\u002Fweaviate\u002Fconcepts\u002Fvector-index",{"title":789,"url":790},"Weaviate Cloud pricing and deployment options","https:\u002F\u002Fweaviate.io\u002Fpricing",{"title":792,"url":793},"Milvus documentation: Overview","https:\u002F\u002Fmilvus.io\u002Fdocs\u002Foverview.md",{"title":795,"url":796},"Pinecone documentation: Database architecture","https:\u002F\u002Fdocs.pinecone.io\u002Fguides\u002Fget-started\u002Fdatabase-architecture",{"title":798,"url":799},"Pinecone documentation: Create an index (clouds, regions, sparse and hybrid)","https:\u002F\u002Fdocs.pinecone.io\u002Fguides\u002Findex-data\u002Fcreate-an-index",{"title":801,"url":802},"OpenSearch documentation: Methods and engines","https:\u002F\u002Fdocs.opensearch.org\u002Flatest\u002Fmappings\u002Fsupported-field-types\u002Fknn-methods-engines\u002F",{"title":804,"url":805},"OpenSearch documentation: Efficient k-NN filtering","https:\u002F\u002Fdocs.opensearch.org\u002Flatest\u002Fvector-search\u002Ffilter-search-knn\u002Fefficient-knn-filtering\u002F",{"title":807,"url":808},"Elasticsearch documentation: Dense vector search","https:\u002F\u002Fwww.elastic.co\u002Fdocs\u002Fsolutions\u002Fsearch\u002Fvector\u002Fdense-vector",{"title":810,"url":811},"Elasticsearch documentation: kNN query (filter as pre-filter)","https:\u002F\u002Fwww.elastic.co\u002Fdocs\u002Freference\u002Fquery-languages\u002Fquery-dsl\u002Fquery-dsl-knn-query",{"title":813,"url":814},"GitHub releases: Qdrant, Weaviate, Milvus, OpenSearch, pgvectorscale (versions as of 1 October 2026)","https:\u002F\u002Fgithub.com\u002Fqdrant\u002Fqdrant\u002Freleases","\u002Fimages\u002Fblog\u002Fpgvector-vs-vector-databases\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fpgvector-vs-vector-databases\u002Fog.jpg",[79,80,81],"pgvector or a vector database? How to choose vector storage in 2026","pgvector, Qdrant, Weaviate, Milvus, Pinecone, OpenSearch or Elasticsearch? A practical 2026 guide to filtering, hybrid search, scale, cost and EU hosting.","Diagram: a decision path from your data to pgvector in Postgres, a search engine with vector fields, or a dedicated vector database.",{"slug":822,"published":5,"minutes":6,"category":7,"tags":823,"keywords":827,"about":837,"sources":845,"cover":882,"og":883,"expertise":77,"locales":884,"lang":79,"title":885,"description":886,"coverAlt":887},"graphrag-knowledge-graph-rag",[824,825,826,10],"GraphRAG","Knowledge graphs","LightRAG",[824,828,829,830,831,832,833,834,835,836],"knowledge graph RAG","GraphRAG vs vector RAG","Microsoft GraphRAG explained","LightRAG vs GraphRAG","GraphRAG global vs local search","GraphRAG indexing cost","when to use GraphRAG","multi-hop RAG","LazyGraphRAG",[838,841,842],{"name":839,"url":840},"Knowledge graph","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FKnowledge_graph",{"name":30,"url":31},{"name":843,"url":844},"Leiden algorithm","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FLeiden_algorithm",[846,849,852,855,858,861,864,867,870,873,876,879],{"title":847,"url":848},"Edge et al.: From Local to Global: A Graph RAG Approach to Query-Focused Summarization (arXiv:2404.16130)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2404.16130",{"title":850,"url":851},"Microsoft GraphRAG documentation: overview","https:\u002F\u002Fmicrosoft.github.io\u002Fgraphrag\u002F",{"title":853,"url":854},"Microsoft GraphRAG documentation: default dataflow","https:\u002F\u002Fmicrosoft.github.io\u002Fgraphrag\u002Findex\u002Fdefault_dataflow\u002F",{"title":856,"url":857},"Microsoft GraphRAG documentation: global search","https:\u002F\u002Fmicrosoft.github.io\u002Fgraphrag\u002Fquery\u002Fglobal_search\u002F",{"title":859,"url":860},"Microsoft GraphRAG documentation: local search","https:\u002F\u002Fmicrosoft.github.io\u002Fgraphrag\u002Fquery\u002Flocal_search\u002F",{"title":862,"url":863},"Microsoft GraphRAG documentation: DRIFT search","https:\u002F\u002Fmicrosoft.github.io\u002Fgraphrag\u002Fquery\u002Fdrift_search\u002F",{"title":865,"url":866},"Microsoft GraphRAG documentation: getting started","https:\u002F\u002Fmicrosoft.github.io\u002Fgraphrag\u002Fget_started\u002F",{"title":868,"url":869},"Microsoft Research: LazyGraphRAG, setting a new standard for quality and cost (25 November 2024)","https:\u002F\u002Fwww.microsoft.com\u002Fen-us\u002Fresearch\u002Fblog\u002Flazygraphrag-setting-a-new-standard-for-quality-and-cost\u002F",{"title":871,"url":872},"Guo et al.: LightRAG: Simple and Fast Retrieval-Augmented Generation (arXiv:2410.05779)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2410.05779",{"title":874,"url":875},"HKUDS\u002FLightRAG on GitHub","https:\u002F\u002Fgithub.com\u002FHKUDS\u002FLightRAG",{"title":877,"url":878},"Gutiérrez et al.: HippoRAG: Neurobiologically Inspired Long-Term Memory for Large Language Models (arXiv:2405.14831)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2405.14831",{"title":880,"url":881},"Xiang et al.: When to use Graphs in RAG: A Comprehensive Analysis for Graph Retrieval-Augmented Generation (arXiv:2506.05690)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2506.05690","\u002Fimages\u002Fblog\u002Fgraphrag-knowledge-graph-rag\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fgraphrag-knowledge-graph-rag\u002Fog.jpg",[79,80,81],"GraphRAG and knowledge-graph RAG: when a graph beats vector search","What Microsoft GraphRAG and LightRAG really do, what indexing costs, and when a knowledge graph beats vector RAG: multi-hop, global questions, product catalogues.","Diagram: a knowledge graph hub linked to entities, communities, local search, global search and product parts.",{"slug":889,"published":5,"minutes":890,"category":7,"tags":891,"keywords":897,"about":907,"sources":915,"cover":949,"og":950,"expertise":77,"locales":951,"lang":79,"title":952,"description":953,"coverAlt":954},"rag-evaluation-metrics",12,[892,893,894,895,896],"RAG evaluation","Retrieval metrics","LLM-as-judge","Golden set","Ragas",[892,898,899,900,901,902,903,904,905,906],"how to evaluate RAG","RAG evaluation metrics","recall@k MRR nDCG","faithfulness vs answer relevance","golden dataset for RAG","LLM as a judge calibration","Ragas vs DeepEval vs TruLens","RAG evals in CI","retrieval vs generation failure",[908,909,912],{"name":30,"url":31},{"name":910,"url":911},"Discounted cumulative gain","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FDiscounted_cumulative_gain",{"name":913,"url":914},"Mean reciprocal rank","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FMean_reciprocal_rank",[916,919,922,925,927,930,933,936,939,942,944,946],{"title":917,"url":918},"Es et al.: RAGAS, Automated Evaluation of Retrieval Augmented Generation (arXiv 2309.15217)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2309.15217",{"title":920,"url":921},"Zheng et al.: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (arXiv 2306.05685)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2306.05685",{"title":923,"url":924},"Ragas documentation: available metrics","https:\u002F\u002Fdocs.ragas.io\u002Fen\u002Fstable\u002Fconcepts\u002Fmetrics\u002Favailable_metrics\u002F",{"title":926,"url":65},"Ragas documentation: faithfulness",{"title":928,"url":929},"Ragas documentation: context precision","https:\u002F\u002Fdocs.ragas.io\u002Fen\u002Fstable\u002Fconcepts\u002Fmetrics\u002Favailable_metrics\u002Fcontext_precision\u002F",{"title":931,"url":932},"Ragas documentation: context recall","https:\u002F\u002Fdocs.ragas.io\u002Fen\u002Fstable\u002Fconcepts\u002Fmetrics\u002Favailable_metrics\u002Fcontext_recall\u002F",{"title":934,"url":935},"DeepEval documentation: metrics introduction","https:\u002F\u002Fdeepeval.com\u002Fdocs\u002Fmetrics-introduction",{"title":937,"url":938},"TruLens","https:\u002F\u002Fwww.trulens.org\u002F",{"title":940,"url":941},"Arize Phoenix documentation","https:\u002F\u002Farize.com\u002Fdocs\u002Fphoenix",{"title":943,"url":911},"Wikipedia: Discounted cumulative gain",{"title":945,"url":914},"Wikipedia: Mean reciprocal rank",{"title":947,"url":948},"Wikipedia: Cohen's kappa","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FCohen%27s_kappa","\u002Fimages\u002Fblog\u002Frag-evaluation-metrics\u002Fcover.webp","\u002Fimages\u002Fblog\u002Frag-evaluation-metrics\u002Fog.jpg",[79,80,81],"Evaluating RAG: retrieval metrics, faithfulness and how to tell which half failed","How to evaluate a RAG system: recall at k, MRR and nDCG vs faithfulness and answer relevance, a golden set from real queries, a calibrated LLM judge and evals in CI.","Diagram: a RAG answer is scored on two sides, retrieval metrics such as recall at k, MRR and nDCG, and generation metrics such as faithfulness and answer relevance, feeding a diagnosis.",{"slug":956,"published":5,"minutes":6,"category":7,"tags":957,"keywords":962,"about":973,"sources":981,"cover":1020,"og":1021,"expertise":1022,"locales":1023,"lang":79,"title":1024,"description":1025,"coverAlt":1026},"semantic-product-search-b2b",[958,742,959,960,961],"B2B search","Semantic search","Spryker","OpenSearch",[963,964,965,966,967,968,969,970,971,972],"semantic product search B2B","B2B ecommerce search","hybrid search BM25 vector","part number search ecommerce","Spryker search Elasticsearch","OpenSearch hybrid search RRF","multilingual product search German English Hungarian","zero results rate site search","LLM query understanding ecommerce","AI product search for B2B shops",[974,976,979],{"name":959,"url":975},"https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FSemantic_search",{"name":977,"url":978},"Elasticsearch","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FElasticsearch",{"name":961,"url":980},"https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FOpenSearch",[982,985,988,990,993,996,999,1002,1005,1008,1011,1014,1017],{"title":983,"url":984},"Baymard Institute: E-commerce search query types","https:\u002F\u002Fbaymard.com\u002Fblog\u002Fecommerce-search-query-types",{"title":986,"url":987},"Elastic Search Labs: Hybrid search in Elasticsearch","https:\u002F\u002Fwww.elastic.co\u002Fsearch-labs\u002Fblog\u002Fhybrid-search-elasticsearch",{"title":989,"url":811},"Elasticsearch documentation: kNN query (pre-filters and post-filters)",{"title":991,"url":992},"Elasticsearch documentation: Semantic reranking","https:\u002F\u002Fwww.elastic.co\u002Fdocs\u002Fsolutions\u002Fsearch\u002Franking\u002Fsemantic-reranking",{"title":994,"url":995},"Elasticsearch documentation: Word delimiter graph token filter","https:\u002F\u002Fwww.elastic.co\u002Fdocs\u002Freference\u002Ftext-analysis\u002Fanalysis-word-delimiter-graph-tokenfilter",{"title":997,"url":998},"Elasticsearch documentation: Synonym graph token filter","https:\u002F\u002Fwww.elastic.co\u002Fdocs\u002Freference\u002Ftext-analysis\u002Fanalysis-synonym-graph-tokenfilter",{"title":1000,"url":1001},"OpenSearch documentation: Score ranker processor (RRF)","https:\u002F\u002Fdocs.opensearch.org\u002Flatest\u002Fsearch-plugins\u002Fsearch-pipelines\u002Fscore-ranker-processor\u002F",{"title":1003,"url":1004},"OpenSearch documentation: Normalization processor","https:\u002F\u002Fdocs.opensearch.org\u002Flatest\u002Fsearch-plugins\u002Fsearch-pipelines\u002Fnormalization-processor\u002F",{"title":1006,"url":1007},"Spryker documentation: Search feature overview","https:\u002F\u002Fdocs.spryker.com\u002Fdocs\u002Fpbc\u002Fall\u002Fsearch\u002Flatest\u002Fbase-shop\u002Fsearch-feature-overview\u002Fsearch-feature-overview",{"title":1009,"url":1010},"Spryker documentation: Migrate from OpenSearch 1.3 to 3.5","https:\u002F\u002Fdocs.spryker.com\u002Fdocs\u002Fpbc\u002Fall\u002Fsearch\u002Flatest\u002Fbase-shop\u002Finstall-and-upgrade\u002Fmigrate-from-opensearch-1.3-to-3.5.html",{"title":1012,"url":1013},"Instacart via ZenML: Rebuilding query understanding for e-commerce search with LLMs","https:\u002F\u002Fwww.zenml.io\u002Fllmops-database\u002Frebuilding-query-understanding-for-e-commerce-search-with-llms",{"title":1015,"url":1016},"arXiv: M3-Embedding, multilingual, multi-functionality, multi-granularity text embeddings","https:\u002F\u002Farxiv.org\u002Fabs\u002F2402.03216",{"title":1018,"url":1019},"Algolia documentation: Search analytics metrics","https:\u002F\u002Fwww.algolia.com\u002Fdoc\u002Fguides\u002Fsearch-analytics\u002Fconcepts\u002Fmetrics\u002F","\u002Fimages\u002Fblog\u002Fsemantic-product-search-b2b\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fsemantic-product-search-b2b\u002Fog.jpg","b2b-ecommerce-developer",[79,80,81],"Semantic product search for B2B shops: part numbers, hybrid retrieval and what to measure","How to add semantic search to a B2B shop without breaking part-number search: hybrid BM25 and vectors, filters, DE\u002FEN\u002FHU, LLM query parsing, reranking and metrics.","Diagram: a search query is split into an identifier lane, lexical BM25 and vector kNN, fused, reranked and returned as results.",1791009037078]