[{"data":1,"prerenderedAt":870},["ShallowReactive",2],{"blog-rag-evaluation-metrics-en":3},{"slug":4,"published":5,"minutes":6,"category":7,"tags":8,"keywords":14,"about":24,"sources":34,"cover":69,"og":70,"expertise":71,"locales":72,"lang":73,"title":76,"description":77,"coverAlt":78,"metaTitle":79,"takeaways":80,"faq":86,"toc":105,"blocks":136,"others":574},"rag-evaluation-metrics","2026-10-02",12,"rag",[9,10,11,12,13],"RAG evaluation","Retrieval metrics","LLM-as-judge","Golden set","Ragas",[9,15,16,17,18,19,20,21,22,23],"how to evaluate RAG","RAG evaluation metrics","recall@k MRR nDCG","faithfulness vs answer relevance","golden dataset for RAG","LLM as a judge calibration","Ragas vs DeepEval vs TruLens","RAG evals in CI","retrieval vs generation failure",[25,28,31],{"name":26,"url":27},"Retrieval-augmented generation","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FRetrieval-augmented_generation",{"name":29,"url":30},"Discounted cumulative gain","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FDiscounted_cumulative_gain",{"name":32,"url":33},"Mean reciprocal rank","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FMean_reciprocal_rank",[35,38,41,44,47,50,53,56,59,62,64,66],{"title":36,"url":37},"Es et al.: RAGAS, Automated Evaluation of Retrieval Augmented Generation (arXiv 2309.15217)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2309.15217",{"title":39,"url":40},"Zheng et al.: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (arXiv 2306.05685)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2306.05685",{"title":42,"url":43},"Ragas documentation: available metrics","https:\u002F\u002Fdocs.ragas.io\u002Fen\u002Fstable\u002Fconcepts\u002Fmetrics\u002Favailable_metrics\u002F",{"title":45,"url":46},"Ragas documentation: faithfulness","https:\u002F\u002Fdocs.ragas.io\u002Fen\u002Fstable\u002Fconcepts\u002Fmetrics\u002Favailable_metrics\u002Ffaithfulness\u002F",{"title":48,"url":49},"Ragas documentation: context precision","https:\u002F\u002Fdocs.ragas.io\u002Fen\u002Fstable\u002Fconcepts\u002Fmetrics\u002Favailable_metrics\u002Fcontext_precision\u002F",{"title":51,"url":52},"Ragas documentation: context recall","https:\u002F\u002Fdocs.ragas.io\u002Fen\u002Fstable\u002Fconcepts\u002Fmetrics\u002Favailable_metrics\u002Fcontext_recall\u002F",{"title":54,"url":55},"DeepEval documentation: metrics introduction","https:\u002F\u002Fdeepeval.com\u002Fdocs\u002Fmetrics-introduction",{"title":57,"url":58},"TruLens","https:\u002F\u002Fwww.trulens.org\u002F",{"title":60,"url":61},"Arize Phoenix documentation","https:\u002F\u002Farize.com\u002Fdocs\u002Fphoenix",{"title":63,"url":30},"Wikipedia: Discounted cumulative gain",{"title":65,"url":33},"Wikipedia: Mean reciprocal rank",{"title":67,"url":68},"Wikipedia: Cohen's kappa","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FCohen%27s_kappa","\u002Fimages\u002Fblog\u002Frag-evaluation-metrics\u002Fcover.webp","\u002Fimages\u002Fblog\u002Frag-evaluation-metrics\u002Fog.jpg","ai-engineer",[73,74,75],"en","de","hu","Evaluating RAG: retrieval metrics, faithfulness and how to tell which half failed","How to evaluate a RAG system: recall at k, MRR and nDCG vs faithfulness and answer relevance, a golden set from real queries, a calibrated LLM judge and evals in CI.","Diagram: a RAG answer is scored on two sides, retrieval metrics such as recall at k, MRR and nDCG, and generation metrics such as faithfulness and answer relevance, feeding a diagnosis.","Evaluating RAG: metrics, golden sets, CI · Balázs Csorba",[81,82,83,84,85],"A RAG system has two halves that fail differently, so score them separately: retrieval with recall@k, MRR and nDCG, generation with faithfulness and answer relevance.","Recall@k is the ceiling. If the needed chunk is not in the top k, no prompt or model can produce a grounded answer, so check retrieval first when an answer is wrong.","A golden set of real user queries beats synthetic questions. Start small, label what a correct answer needs, and grow it from every production failure.","An LLM judge is a measuring instrument, not a ground truth. Calibrate it against human labels on a sample, report chance-corrected agreement and re-check it whenever the judge model changes.","Run cheap retrieval checks on every pull request and the LLM-judged suite on a schedule or when the index, prompt or model changes, and gate on regression against a baseline.",[87,90,93,96,99,102],{"q":88,"a":89},"How do you evaluate a RAG system?","Evaluate retrieval and generation separately on the same golden set of real queries. For retrieval, measure whether the right chunks were found and how high they ranked (recall@k, MRR, nDCG). For generation, measure whether the answer is supported by the retrieved context (faithfulness) and whether it addresses the question (answer relevance). Then use the two sets of scores to find which half failed.",{"q":91,"a":92},"What is the difference between recall@k, MRR and nDCG?","Recall@k asks whether the relevant material appears anywhere in the top k results. MRR is the average of 1 divided by the rank of the first relevant result, so it only cares about the first hit. nDCG scores the whole ranking with graded relevance, discounting results that appear lower, and is normalised to a range of 0 to 1.",{"q":94,"a":95},"What is faithfulness in RAG evaluation?","Faithfulness measures whether the claims in the generated answer are supported by the retrieved context. In Ragas it is the number of supported claims divided by the total number of claims in the response, a score between 0 and 1. A faithful answer can still be wrong if retrieval brought back the wrong context.",{"q":97,"a":98},"Can I trust an LLM as a judge?","Partly. The MT-Bench study found that a strong judge such as GPT-4 reached over 80 percent agreement with human preferences, the same level as agreement between humans, but it also documented position, verbosity and self-enhancement bias. Treat the judge as an instrument: label a sample by hand, measure agreement, and recalibrate when the judge model or the rubric changes.",{"q":100,"a":101},"Which RAG evaluation tool should I use: Ragas, DeepEval, TruLens or Phoenix?","They overlap on metrics and differ in workflow. Ragas is a metrics framework, DeepEval is built around pytest-style tests and a CI command, TruLens centres on the RAG triad with tracing, and Arize Phoenix combines OpenTelemetry-based tracing with evaluations, datasets and experiments. I would pick by where you want the evals to live, not by metric names.",{"q":103,"a":104},"How do I know whether retrieval or generation caused a wrong answer?","Look at the retrieved chunks for that query. If the needed fact was never in the corpus, it is a content gap. If it was in the corpus but not in the top k, retrieval failed. If it was in the context but the answer ignores or contradicts it, generation failed, which shows up as low faithfulness or low answer relevance.",[106,109,112,115,118,121,124,127,130,133],{"id":107,"title":108},"two-halves","Two halves, seven numbers",{"id":110,"title":111},"retrieval-metrics","Retrieval metrics: recall@k, MRR and nDCG",{"id":113,"title":114},"generation-metrics","Generation metrics: faithfulness and answer relevance",{"id":116,"title":117},"golden-set","Build the golden set from real queries",{"id":119,"title":120},"diagnosis","Diagnose: did retrieval or generation fail?",{"id":122,"title":123},"llm-judge","LLM-as-judge: calibrate it against humans",{"id":125,"title":126},"tools","Tools: Ragas, DeepEval, TruLens and Phoenix",{"id":128,"title":129},"ci","Run evals in CI",{"id":131,"title":132},"first-steps","What I would do first",{"id":134,"title":135},"sources","Sources",[137,141,144,158,161,177,180,249,252,253,256,274,281,284,285,291,300,303,310,311,314,351,354,355,358,367,386,393,394,397,404,431,439,440,443,478,481,482,485,512,517,518,532,535,536],{"type":138,"content":139},"paragraph",[140],"Most RAG systems are shipped on the strength of a demo. Someone asks five questions, the answers look right, and the team moves on. Then the corpus grows, the chunking changes, the model is upgraded, and nobody can say whether the system got better or worse. The honest answer to \"is it good?\" is a number you can re-compute, and for RAG that takes more than one number.",{"type":138,"content":142},[143],"The reason is structural. A RAG system is two systems in a row: a retriever that decides what the model sees, and a generator that decides what to say about it. Each fails in its own way, needs its own metrics and gets fixed by different work. If you only score the final answer, a wrong answer tells you that something broke, but not where.",{"type":138,"content":145},[146,147,152,153,157],"This article is how I would set up RAG evaluation for a product team: the metrics and what each one actually measures, a golden set built from real queries, an LLM judge you have checked against humans, the tools worth looking at, evals that run in CI, and a short procedure for diagnosing whether retrieval or generation failed. It builds on my posts about ",{"tag":148,"to":149,"children":150},"link","\u002Fblog\u002Fllm-evals-for-product-features",[151],"evals for LLM product features"," and the ",{"tag":148,"to":154,"children":155},"\u002Fblog\u002Frag-pipeline-chunking-hybrid-search-reranking",[156],"RAG pipeline itself",".",{"type":159,"level":160,"id":107,"text":108},"heading",2,{"type":138,"content":162},[163,164,168,169,173,174],"The ",{"tag":165,"href":37,"children":166},"a",[167],"RAGAS paper"," frames RAG evaluation along three dimensions: the retrieval system's ability to find relevant and focused passages, the generator's ability to use those passages faithfully, and the quality of the output. In practice I split the numbers by the question they answer. Retrieval metrics ask ",{"tag":170,"children":171},"em",[172],"did the right material reach the model, and in a good order?"," Generation metrics ask ",{"tag":170,"children":175},[176],"given what it saw, did the model behave?",{"type":138,"content":178},[179],"The table lists the metrics I actually use. Names differ between tools, so I describe what each one measures rather than what a library calls it.",{"type":181,"head":182,"rows":191},"table",[183,185,187,189],[184],"Metric",[186],"Layer",[188],"The question it answers",[190],"Needs labels?",[192,201,208,216,225,233,242],[193,195,197,199],[194],"recall@k",[196],"Retrieval",[198],"Is the relevant material anywhere in the top k results?",[200],"Relevant chunks per query",[202,204,205,207],[203],"MRR",[196],[206],"How high is the first relevant result? The mean of 1 divided by its rank.",[200],[209,211,212,214],[210],"nDCG",[196],[213],"Is the whole ranking in a good order, with graded relevance and a discount for lower positions?",[215],"Graded relevance",[217,219,221,223],[218],"Context precision",[220],"Retrieval (ranking)",[222],"Are relevant chunks ranked above irrelevant ones in what the model received?",[224],"Optional reference answer in Ragas",[226,228,229,231],[227],"Context recall",[196],[230],"Is everything the reference answer says supported by the retrieved context?",[232],"Reference answer",[234,236,238,240],[235],"Faithfulness",[237],"Generation",[239],"Is each claim in the answer supported by the retrieved context?",[241],"No",[243,245,246,248],[244],"Answer relevance",[237],[247],"Does the answer address the question that was asked?",[241],{"type":138,"content":250},[251],"Note that context precision sits with retrieval, even though it is often listed beside the generation metrics. It is computed over the retrieved chunks, so it tells you about the ranker, not about the model's writing. Keeping the layer column honest is what makes the diagnosis later in this article possible.",{"type":159,"level":160,"id":110,"text":111},{"type":138,"content":254},[255],"Retrieval metrics come from search evaluation and need one thing the generation metrics do not: a notion of which chunks are relevant for each query.",{"type":257,"ordered":258,"items":259},"list",false,[260,266,270],[261,265],{"tag":262,"children":263},"strong",[264],"Recall@k"," asks whether the relevant material is in the top k. It is the most important retrieval number, because it is a ceiling. If the chunk that contains the answer is not among the k chunks you pass to the model, the model cannot ground its answer in it, whatever the prompt says. Choose k to match what you actually send to the generator.",[267,269],{"tag":262,"children":268},[203]," (mean reciprocal rank) averages 1 divided by the rank of the first relevant result: 1 for first place, 0.5 for second. It rewards putting a good chunk at the top. Its documented limitation is that only the first relevant result counts and any further ones are ignored, so it suits questions with one right passage and misleads on questions that need several.",[271,273],{"tag":262,"children":272},[210]," sums graded relevance scores with a logarithmic discount for lower positions, then divides by the ideal ordering so the result lies between 0 and 1. It is the metric to use when relevance is not binary, for example a chunk that fully answers versus one that only mentions the topic, and when ordering matters because the context window or a reranker's cut-off truncates the list.",{"type":138,"content":275},[276,277,280],"I read them together. High recall@k with a poor MRR or nDCG means the right chunk is retrieved but buried, which a reranker usually fixes (see the ",{"tag":148,"to":154,"children":278},[279],"pipeline post"," for chunking, hybrid search and reranking). Low recall@k means the right chunk is never found, which is a problem with chunking, embeddings, query rewriting or the index, and no reranker can repair a candidate list that does not contain the answer.",{"type":138,"content":282},[283],"These metrics need relevance labels, which is the expensive part. The Ragas documentation describes an ID-based and a non-LLM variant of context recall that compare retrieved chunk IDs or text to reference contexts without an LLM call. Those are cheap, deterministic and perfect for CI, if your golden set stores which chunks are relevant.",{"type":159,"level":160,"id":113,"text":114},{"type":138,"content":286},[287,288,290],"Generation metrics judge the answer given what the model was shown. ",{"tag":262,"children":289},[235]," is the central one. In the Ragas definition it is the number of claims in the response that the retrieved context supports, divided by the total number of claims: the response is decomposed into statements, each statement is checked against the context, and the ratio is the score, from 0 to 1. It needs no reference answer, which is why it is so useful on production traffic.",{"type":138,"content":292},[293,295,296,299],{"tag":262,"children":294},[244]," asks whether the answer addresses the question at all. An answer can be perfectly faithful to the context and still dodge the question, for example by summarising a document when the user asked for a number. TruLens calls the equivalent triad ",{"tag":170,"children":297},[298],"context relevance, groundedness and answer relevance","; Ragas calls it response relevancy; DeepEval calls it answer relevancy. Same idea, different labels.",{"type":138,"content":301},[302],"Two properties are easy to miss. First, faithfulness measures agreement with the retrieved context, not with the truth. If retrieval returned a stale policy document, a perfectly faithful answer is wrong. Second, a high score can be earned by saying little: a model that refuses or hedges makes few claims, and few claims can all be supported. Always read faithfulness next to answer relevance and a correctness check against the golden answer.",{"type":304,"variant":305,"title":306,"body":307},"callout","note","Faithful is not correct",[308],[309],"Faithfulness only tells you the answer matches the chunks. Whether the chunks were the right ones is a retrieval question, and whether the final answer matches the golden answer is a separate correctness score. I track all three, because the combinations are what point at the cause.",{"type":159,"level":160,"id":116,"text":117},{"type":138,"content":312},[313],"Everything above depends on a golden set: queries with the expectations you score against. Synthetic questions generated from your documents are a useful bootstrap, but they share the vocabulary of the documents, which real users do not. Real queries have typos, abbreviations, vague phrasing, product codes and questions the corpus cannot answer. I would build the set from them.",{"type":257,"ordered":315,"items":316},true,[317,326,331,336,341,346],[318,321,322,157],{"tag":262,"children":319},[320],"Sample real queries."," Take them from search logs, support tickets or chat transcripts. Remove personal data first, and mind your legal basis: see my notes on ",{"tag":148,"to":323,"children":324},"\u002Fblog\u002Fgdpr-llm-api-eu-data-residency",[325],"GDPR and LLM APIs",[327,330],{"tag":262,"children":328},[329],"Stratify, do not just sample."," Include the common questions, the long tail, multi-document questions, questions with exact identifiers, and questions that should get \"I don't know\" because the corpus has no answer.",[332,335],{"tag":262,"children":333},[334],"Label what a correct answer needs."," For each query, record the relevant chunks or documents (for recall@k, MRR and nDCG), a short reference answer, and where it matters the facts the answer must include.",[337,340],{"tag":262,"children":338},[339],"Have a domain expert label it, and measure whether two people agree."," If two humans disagree about what is relevant, no metric can resolve that for you.",[342,345],{"tag":262,"children":343},[344],"Version it and keep a held-out part."," Treat the set like code. Keep a slice you never tune against, so improvements are not just overfitting to the examples you stared at.",[347,350],{"tag":262,"children":348},[349],"Feed it from production."," Every thumbs-down, escalation or wrong answer found in review becomes a new case, with its label. This is how the set stays representative as the product changes.",{"type":138,"content":352},[353],"How big? I would start with a few dozen well-labelled queries and grow from there; a small clean set that you trust beats a large one nobody checked. The point is not statistical power at the start, it is to have something that makes a regression visible the day it happens. As the set grows, report results per stratum, because an average hides a collapse in one query type.",{"type":159,"level":160,"id":119,"text":120},{"type":138,"content":356},[357],"When a wrong answer is reported, resist the urge to edit the prompt. Walk the chain from the bottom, using the stored retrieval results for that query, and stop at the first question you answer with \"no\".",{"type":359,"attrs":360,"inner":364,"caption":365},"diagram",{"viewBox":361,"role":362,"aria-labelledby":363},"0 0 720 380","img","d1-rageval-t d1-rageval-d","\u003Ctitle id=\"d1-rageval-t\">Diagnosing a bad RAG answer\u003C\u002Ftitle>\u003Cdesc id=\"d1-rageval-d\">A decision chain of four questions. If the fact is not in the corpus, it is a content gap. If it is in the corpus but not in the top k chunks, retrieval failed. If it is in the context but the answer is not supported, generation is unfaithful. If it is supported but does not address the question, the answer is off target. If all checks pass, re-check the golden label or the judge.\u003C\u002Fdesc>\u003Ctext x=\"700\" y=\"22\" text-anchor=\"end\" class=\"d-label\">Start at the top, stop at the first no\u003C\u002Ftext>\u003Crect x=\"20\" y=\"40\" width=\"300\" height=\"52\" rx=\"10\" class=\"d-gold\" \u002F>\u003Ctext x=\"170\" y=\"71\" text-anchor=\"middle\" class=\"d-text\">Is the fact in the corpus?\u003C\u002Ftext>\u003Cpath d=\"M170 92 V102\" class=\"d-line\" \u002F>\u003Cpath d=\"M170 110 l-5 -9 h10 z\" class=\"d-head\" \u002F>\u003Ctext x=\"180\" y=\"105\" class=\"d-label\">yes\u003C\u002Ftext>\u003Cpath d=\"M320 66 H392\" class=\"d-line-accent\" \u002F>\u003Cpath d=\"M400 66 l-9 -5 v10 z\" class=\"d-head-accent\" \u002F>\u003Ctext x=\"332\" y=\"60\" class=\"d-label\">no\u003C\u002Ftext>\u003Crect x=\"400\" y=\"40\" width=\"300\" height=\"52\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"550\" y=\"62\" text-anchor=\"middle\" class=\"d-text\">Content gap\u003C\u002Ftext>\u003Ctext x=\"550\" y=\"83\" text-anchor=\"middle\" class=\"d-small\">Fix the sources, not the model\u003C\u002Ftext>\u003Crect x=\"20\" y=\"110\" width=\"300\" height=\"52\" rx=\"10\" class=\"d-gold\" \u002F>\u003Ctext x=\"170\" y=\"141\" text-anchor=\"middle\" class=\"d-text\">Is it in the top k chunks?\u003C\u002Ftext>\u003Cpath d=\"M170 162 V172\" class=\"d-line\" \u002F>\u003Cpath d=\"M170 180 l-5 -9 h10 z\" class=\"d-head\" \u002F>\u003Ctext x=\"180\" y=\"175\" class=\"d-label\">yes\u003C\u002Ftext>\u003Cpath d=\"M320 136 H392\" class=\"d-line-accent\" \u002F>\u003Cpath d=\"M400 136 l-9 -5 v10 z\" class=\"d-head-accent\" \u002F>\u003Ctext x=\"332\" y=\"130\" class=\"d-label\">no\u003C\u002Ftext>\u003Crect x=\"400\" y=\"110\" width=\"300\" height=\"52\" rx=\"10\" class=\"d-accent\" \u002F>\u003Ctext x=\"550\" y=\"132\" text-anchor=\"middle\" class=\"d-text\">Retrieval failure\u003C\u002Ftext>\u003Ctext x=\"550\" y=\"153\" text-anchor=\"middle\" class=\"d-small\">recall@k, MRR, nDCG are low\u003C\u002Ftext>\u003Crect x=\"20\" y=\"180\" width=\"300\" height=\"52\" rx=\"10\" class=\"d-gold\" \u002F>\u003Ctext x=\"170\" y=\"211\" text-anchor=\"middle\" class=\"d-text\">Is every claim supported?\u003C\u002Ftext>\u003Cpath d=\"M170 232 V242\" class=\"d-line\" \u002F>\u003Cpath d=\"M170 250 l-5 -9 h10 z\" class=\"d-head\" \u002F>\u003Ctext x=\"180\" y=\"245\" class=\"d-label\">yes\u003C\u002Ftext>\u003Cpath d=\"M320 206 H392\" class=\"d-line-accent\" \u002F>\u003Cpath d=\"M400 206 l-9 -5 v10 z\" class=\"d-head-accent\" \u002F>\u003Ctext x=\"332\" y=\"200\" class=\"d-label\">no\u003C\u002Ftext>\u003Crect x=\"400\" y=\"180\" width=\"300\" height=\"52\" rx=\"10\" class=\"d-sky\" \u002F>\u003Ctext x=\"550\" y=\"202\" text-anchor=\"middle\" class=\"d-text\">Generation: unfaithful\u003C\u002Ftext>\u003Ctext x=\"550\" y=\"223\" text-anchor=\"middle\" class=\"d-small\">faithfulness is low\u003C\u002Ftext>\u003Crect x=\"20\" y=\"250\" width=\"300\" height=\"52\" rx=\"10\" class=\"d-gold\" \u002F>\u003Ctext x=\"170\" y=\"281\" text-anchor=\"middle\" class=\"d-text\">Does it answer the question?\u003C\u002Ftext>\u003Cpath d=\"M170 302 V312\" class=\"d-line\" \u002F>\u003Cpath d=\"M170 320 l-5 -9 h10 z\" class=\"d-head\" \u002F>\u003Ctext x=\"180\" y=\"319\" class=\"d-label\">yes\u003C\u002Ftext>\u003Cpath d=\"M320 276 H392\" class=\"d-line-accent\" \u002F>\u003Cpath d=\"M400 276 l-9 -5 v10 z\" class=\"d-head-accent\" \u002F>\u003Ctext x=\"332\" y=\"270\" class=\"d-label\">no\u003C\u002Ftext>\u003Crect x=\"400\" y=\"250\" width=\"300\" height=\"52\" rx=\"10\" class=\"d-sky\" \u002F>\u003Ctext x=\"550\" y=\"272\" text-anchor=\"middle\" class=\"d-text\">Generation: off target\u003C\u002Ftext>\u003Ctext x=\"550\" y=\"293\" text-anchor=\"middle\" class=\"d-small\">answer relevance is low\u003C\u002Ftext>\u003Crect x=\"20\" y=\"320\" width=\"300\" height=\"44\" rx=\"10\" class=\"d-mint\" \u002F>\u003Ctext x=\"170\" y=\"347\" text-anchor=\"middle\" class=\"d-text\">All pass: re-check label or judge\u003C\u002Ftext>",[366],"The order matters: each question assumes the ones above it have passed. Most of the surprises I see are in the first two.",{"type":138,"content":368},[369,370,373,374,377,378,381,382,385],"The branches map to different work. A ",{"tag":262,"children":371},[372],"content gap"," is an editorial or ingestion problem: the document is missing, outdated or was never parsed properly. A ",{"tag":262,"children":375},[376],"retrieval failure"," is fixed in chunking, embeddings, hybrid search, query rewriting or reranking. An ",{"tag":262,"children":379},[380],"unfaithful"," answer is a generation problem: tighten the instruction to answer only from the context, reduce irrelevant context that invites speculation, or change the model. An ",{"tag":262,"children":383},[384],"off-target"," answer is usually the prompt, or a retrieval problem in disguise where the context is on topic but not on the question.",{"type":138,"content":387},[388,389,157],"The last box matters too. If every metric passes and the answer is still wrong, suspect the golden label, the reference answer or the judge itself before the system. For agentic setups, where the model decides what to retrieve and may retrieve several times, the same logic applies per retrieval step; see ",{"tag":148,"to":390,"children":391},"\u002Fblog\u002Frag-2026-hybrid-agentic-long-context",[392],"RAG in 2026: hybrid, agentic and long context",{"type":159,"level":160,"id":122,"text":123},{"type":138,"content":395},[396],"Faithfulness, answer relevance and most context metrics are scored by an LLM. Ragas notes that its LLM-based metrics may use one or more LLM calls to produce a score. That makes the judge part of your measurement setup, and an instrument needs calibration.",{"type":138,"content":398},[399,400,403],"The evidence is encouraging but not unconditional. The ",{"tag":165,"href":40,"children":401},[402],"MT-Bench paper"," found that a strong judge such as GPT-4 reached over 80 percent agreement with human preferences, the same level as agreement between humans. The same paper names position bias, verbosity bias and self-enhancement bias, and notes limited reasoning ability. Those findings are about pairwise preference on chat answers; they do not transfer automatically to your domain, your rubric or your judge model.",{"type":257,"ordered":315,"items":405},[406,411,416,421,426],[407,410],{"tag":262,"children":408},[409],"Label a sample by hand."," Take a few dozen cases that span good, bad and borderline outputs, and have a person score them with the same rubric the judge gets.",[412,415],{"tag":262,"children":413},[414],"Measure agreement beyond chance."," Percent agreement flatters you when one label dominates. Cohen's kappa corrects for chance agreement (kappa equals observed agreement minus expected agreement, divided by one minus expected agreement) and runs from -1 to 1. Its own caveat applies: the value depends on prevalence and bias, so no single threshold is universal. Look at the confusion matrix, not only the number.",[417,420],{"tag":262,"children":418},[419],"Fix the judge's inputs."," Use a narrow rubric, binary or small-scale labels, and ask for the reasoning before the score. Pin the exact judge model version, because a silent update changes your trend line.",[422,425],{"tag":262,"children":423},[424],"Counter known biases."," Randomise answer order in pairwise comparisons, avoid using the same model to generate and to judge when you can, and check whether scores correlate with answer length.",[427,430],{"tag":262,"children":428},[429],"Re-calibrate on change."," New judge model, new rubric, new domain: repeat the comparison. Keep the labelled sample as a permanent calibration set.",{"type":138,"content":432},[433,434,438],"Judge calls cost tokens on every run, so the cost and latency tactics from ",{"tag":148,"to":435,"children":436},"\u002Fblog\u002Fllm-cost-latency-prompt-caching-routing",[437],"LLM cost, latency and prompt caching"," apply to your eval suite too. A smaller judge model is acceptable if, and only if, the calibration says so.",{"type":159,"level":160,"id":125,"text":126},{"type":138,"content":441},[442],"You can write all of this yourself in a few hundred lines, and for the deterministic retrieval metrics I often do. For LLM-judged metrics and for tracing, the established tools save real time. This is what their own documentation says, checked on 2 October 2026:",{"type":181,"head":444,"rows":451},[445,447,449],[446],"Tool",[448],"What it is",[450],"RAG metrics and workflow",[452,458,465,471],[453,454,456],[13],[455],"A metrics framework from the RAGAS paper",[457],"Context precision, context recall, noise sensitivity, response relevancy, faithfulness; LLM-based and non-LLM variants; custom metrics supported",[459,461,463],[460],"DeepEval",[462],"A test-style evaluation framework",[464],"Contextual relevancy, precision and recall for the retriever; answer relevancy and faithfulness for the generator; native pytest integration through the deepeval test run command for CI\u002FCD",[466,467,469],[57],[468],"Open-source evaluation and tracing, OpenTelemetry-native, maintained by Snowflake",[470],"The RAG triad of context relevance, groundedness and answer relevance; works with LangChain, LlamaIndex and LangGraph",[472,474,476],[473],"Arize Phoenix",[475],"Open-source AI observability and evaluation, built on OpenTelemetry and OpenInference",[477],"Tracing, evaluations with LLM, code or human labels, datasets and experiments to compare changes on the same inputs; Arize AX is the managed option",{"type":138,"content":479},[480],"My practical advice: choose by where evals should live. If you want them as tests in the repository, DeepEval or a thin Ragas wrapper fits. If you want to inspect real production traces and attach scores to them, TruLens or Phoenix fit. Whatever you choose, keep the golden set and the labels in your own repository in a plain format, so switching tools costs an afternoon rather than a quarter. And remember that metrics with the same name can be computed differently across tools, so never compare scores between them.",{"type":159,"level":160,"id":128,"text":129},{"type":138,"content":483},[484],"An eval that nobody runs is documentation. The goal is that a change to chunking, the embedding model, the prompt or the generator cannot merge without producing numbers. I use two tiers, because the cost profiles differ.",{"type":257,"ordered":258,"items":486},[487,492,497,502,507],[488,491],{"tag":262,"children":489},[490],"Every pull request: retrieval only."," Run recall@k, MRR and nDCG against a fixed index snapshot or fixture. These can be deterministic and need no LLM call, so they are fast, free and stable.",[493,496],{"tag":262,"children":494},[495],"On a schedule or on relevant changes: the LLM-judged suite."," Run faithfulness, answer relevance and correctness when the index, prompt, model or judge changes, and nightly otherwise. DeepEval's documentation describes a test command for running evals in CI\u002FCD pipelines.",[498,501],{"tag":262,"children":499},[500],"Gate on regression, not on a magic number."," Compare to the last accepted baseline per stratum and fail on a drop beyond the noise you measured by running the suite several times.",[503,506],{"tag":262,"children":504},[505],"Control the noise."," Pin judge and generator versions, set temperature low, cache judge calls for unchanged inputs, and store every run's scores and retrieved chunk IDs so a failure can be diagnosed without re-running.",[508,511],{"tag":262,"children":509},[510],"Close the loop."," A failing production case becomes a golden-set row in the same pull request that fixes it.",{"type":138,"content":513},[514,515,157],"For how this fits the broader practice of testing LLM features, including online monitoring and human review, see ",{"tag":148,"to":149,"children":516},[151],{"type":159,"level":160,"id":131,"text":132},{"type":257,"ordered":315,"items":519},[520,522,524,526,528,530],[521],"Log, for every request, the query, the retrieved chunk IDs with scores and the final answer.",[523],"Collect 30 to 50 real queries across your strata and label relevant chunks and a short reference answer.",[525],"Compute recall@k, MRR and nDCG today. Fix retrieval before touching the prompt.",[527],"Add faithfulness and answer relevance with an LLM judge, and calibrate it on a hand-labelled sample using kappa.",[529],"Wire the cheap metrics into pull requests and the judged suite into a nightly job.",[531],"Add every production failure to the golden set, with its diagnosis from the decision chain.",{"type":138,"content":533},[534],"None of this needs a platform. It needs a labelled set, a few honest numbers and the habit of asking, for each bad answer, which half failed.",{"type":159,"level":160,"id":134,"text":135},{"type":257,"ordered":315,"items":537},[538,541,544,547,550,553,556,559,562,565,568,571],[539],{"tag":165,"href":37,"children":540},[36],[542],{"tag":165,"href":40,"children":543},[39],[545],{"tag":165,"href":43,"children":546},[42],[548],{"tag":165,"href":46,"children":549},[45],[551],{"tag":165,"href":49,"children":552},[48],[554],{"tag":165,"href":52,"children":555},[51],[557],{"tag":165,"href":55,"children":558},[54],[560],{"tag":165,"href":58,"children":561},[57],[563],{"tag":165,"href":61,"children":564},[60],[566],{"tag":165,"href":30,"children":567},[63],[569],{"tag":165,"href":33,"children":570},[65],[572],{"tag":165,"href":68,"children":573},[67],[575,647,731,798],{"slug":576,"published":5,"minutes":577,"category":7,"tags":578,"keywords":583,"about":594,"sources":602,"cover":641,"og":642,"expertise":71,"locales":643,"lang":73,"title":644,"description":645,"coverAlt":646},"llm-hallucination-grounding-citations",13,[579,580,581,582,235],"Hallucinations","RAG","Citations","Grounding",[584,585,586,587,588,589,590,591,592,593],"reduce LLM hallucinations in production","how to reduce hallucinations in RAG","LLM citations API","Anthropic citations API","RAG faithfulness metric","LLM abstention I don't know","claim-level verification LLM","grounding LLM answers in sources","check grounding API","show sources in AI chatbot UI",[595,598,599],{"name":596,"url":597},"Hallucination (artificial intelligence)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FHallucination_(artificial_intelligence)",{"name":26,"url":27},{"name":600,"url":601},"Large language model","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FLarge_language_model",[603,606,609,612,615,618,621,624,627,630,632,635,638],{"title":604,"url":605},"Anthropic: Citations (Claude API documentation)","https:\u002F\u002Fplatform.claude.com\u002Fdocs\u002Fen\u002Fbuild-with-claude\u002Fcitations",{"title":607,"url":608},"Anthropic: Search results (Claude API documentation)","https:\u002F\u002Fplatform.claude.com\u002Fdocs\u002Fen\u002Fbuild-with-claude\u002Fsearch-results",{"title":610,"url":611},"Anthropic: Reduce hallucinations (Claude API documentation)","https:\u002F\u002Fplatform.claude.com\u002Fdocs\u002Fen\u002Ftest-and-evaluate\u002Fstrengthen-guardrails\u002Freduce-hallucinations",{"title":613,"url":614},"Anthropic: Introducing Citations on the Anthropic API","https:\u002F\u002Fclaude.com\u002Fblog\u002Fintroducing-citations-api",{"title":616,"url":617},"Simon Willison: Anthropic's new Citations API (24 January 2025)","https:\u002F\u002Fsimonwillison.net\u002F2025\u002FJan\u002F24\u002Fanthropics-new-citations-api\u002F",{"title":619,"url":620},"OpenAI: Web search guide (url_citation annotations and display requirement)","https:\u002F\u002Fdevelopers.openai.com\u002Fapi\u002Fdocs\u002Fguides\u002Ftools-web-search",{"title":622,"url":623},"Cohere: Documents and citations","https:\u002F\u002Fdocs.cohere.com\u002Fdocs\u002Fdocuments-and-citations",{"title":625,"url":626},"Google Cloud: Check grounding API","https:\u002F\u002Fdocs.cloud.google.com\u002Fgenerative-ai-app-builder\u002Fdocs\u002Fcheck-grounding",{"title":628,"url":629},"AWS: Amazon Bedrock Guardrails contextual grounding check","https:\u002F\u002Fdocs.aws.amazon.com\u002Fbedrock\u002Flatest\u002Fuserguide\u002Fguardrails-contextual-grounding-check.html",{"title":631,"url":46},"Ragas: Faithfulness metric",{"title":633,"url":634},"Kalai, Nachum, Vempala, Zhang: Why Language Models Hallucinate (arXiv 2509.04664)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2509.04664",{"title":636,"url":637},"Magesh et al.: Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools (arXiv 2405.20362)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2405.20362",{"title":639,"url":640},"Wallat, Heuss, de Rijke, Anand: Correctness is not Faithfulness in RAG Attributions (arXiv 2412.18004)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2412.18004","\u002Fimages\u002Fblog\u002Fllm-hallucination-grounding-citations\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fllm-hallucination-grounding-citations\u002Fog.jpg",[73,74,75],"Reducing LLM hallucinations in production: grounding, citations and knowing when to say no","Cut hallucinations in production RAG: citation APIs, abstention, claim-level checks, faithfulness metrics, source UI, and the failures that still slip through.","Diagram: a retrieval step feeds an evidence gate, a cited answer and a claim verifier, ending in an answer with sources, with abstain and flag paths branching off.",{"slug":648,"published":5,"minutes":577,"category":7,"tags":649,"keywords":654,"about":665,"sources":673,"cover":725,"og":726,"expertise":71,"locales":727,"lang":73,"title":728,"description":729,"coverAlt":730},"pgvector-vs-vector-databases",[650,651,580,652,653],"pgvector","Vector databases","Hybrid search","EU hosting",[655,656,657,658,659,660,661,662,663,664],"pgvector vs vector database","pgvector vs Qdrant","pgvector vs Pinecone","best vector database 2026","pgvector HNSW iterative scan","pgvector halfvec","OpenSearch vs Elasticsearch vector search","vector database EU hosting","hybrid search Postgres","Weaviate vs Milvus",[666,669,672],{"name":667,"url":668},"Vector database","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FVector_database",{"name":670,"url":671},"PostgreSQL","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FPostgreSQL",{"name":26,"url":27},[674,677,680,683,686,689,692,695,698,701,704,707,710,713,716,719,722],{"title":675,"url":676},"pgvector README (index limits, HNSW defaults, iterative scans, filtering, halfvec)","https:\u002F\u002Fgithub.com\u002Fpgvector\u002Fpgvector\u002Fblob\u002Fmaster\u002FREADME.md",{"title":678,"url":679},"pgvector CHANGELOG (0.4.0 to 0.8.7)","https:\u002F\u002Fgithub.com\u002Fpgvector\u002Fpgvector\u002Fblob\u002Fmaster\u002FCHANGELOG.md",{"title":681,"url":682},"pgvectorscale: StreamingDiskANN, statistical binary quantization, filtered search","https:\u002F\u002Fgithub.com\u002Ftimescale\u002Fpgvectorscale",{"title":684,"url":685},"Qdrant documentation: Filtering","https:\u002F\u002Fqdrant.tech\u002Fdocumentation\u002Fconcepts\u002Ffiltering\u002F",{"title":687,"url":688},"Qdrant documentation: Hybrid queries","https:\u002F\u002Fqdrant.tech\u002Fdocumentation\u002Fconcepts\u002Fhybrid-queries\u002F",{"title":690,"url":691},"Qdrant documentation: Create a cluster (providers, free tier, Hybrid Cloud)","https:\u002F\u002Fqdrant.tech\u002Fdocumentation\u002Fcloud\u002Fcreate-cluster\u002F",{"title":693,"url":694},"Weaviate documentation: Hybrid search","https:\u002F\u002Fdocs.weaviate.io\u002Fweaviate\u002Fconcepts\u002Fsearch\u002Fhybrid-search",{"title":696,"url":697},"Weaviate documentation: Vector index types","https:\u002F\u002Fdocs.weaviate.io\u002Fweaviate\u002Fconcepts\u002Fvector-index",{"title":699,"url":700},"Weaviate Cloud pricing and deployment options","https:\u002F\u002Fweaviate.io\u002Fpricing",{"title":702,"url":703},"Milvus documentation: Overview","https:\u002F\u002Fmilvus.io\u002Fdocs\u002Foverview.md",{"title":705,"url":706},"Pinecone documentation: Database architecture","https:\u002F\u002Fdocs.pinecone.io\u002Fguides\u002Fget-started\u002Fdatabase-architecture",{"title":708,"url":709},"Pinecone documentation: Create an index (clouds, regions, sparse and hybrid)","https:\u002F\u002Fdocs.pinecone.io\u002Fguides\u002Findex-data\u002Fcreate-an-index",{"title":711,"url":712},"OpenSearch documentation: Methods and engines","https:\u002F\u002Fdocs.opensearch.org\u002Flatest\u002Fmappings\u002Fsupported-field-types\u002Fknn-methods-engines\u002F",{"title":714,"url":715},"OpenSearch documentation: Efficient k-NN filtering","https:\u002F\u002Fdocs.opensearch.org\u002Flatest\u002Fvector-search\u002Ffilter-search-knn\u002Fefficient-knn-filtering\u002F",{"title":717,"url":718},"Elasticsearch documentation: Dense vector search","https:\u002F\u002Fwww.elastic.co\u002Fdocs\u002Fsolutions\u002Fsearch\u002Fvector\u002Fdense-vector",{"title":720,"url":721},"Elasticsearch documentation: kNN query (filter as pre-filter)","https:\u002F\u002Fwww.elastic.co\u002Fdocs\u002Freference\u002Fquery-languages\u002Fquery-dsl\u002Fquery-dsl-knn-query",{"title":723,"url":724},"GitHub releases: Qdrant, Weaviate, Milvus, OpenSearch, pgvectorscale (versions as of 1 October 2026)","https:\u002F\u002Fgithub.com\u002Fqdrant\u002Fqdrant\u002Freleases","\u002Fimages\u002Fblog\u002Fpgvector-vs-vector-databases\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fpgvector-vs-vector-databases\u002Fog.jpg",[73,74,75],"pgvector or a vector database? How to choose vector storage in 2026","pgvector, Qdrant, Weaviate, Milvus, Pinecone, OpenSearch or Elasticsearch? A practical 2026 guide to filtering, hybrid search, scale, cost and EU hosting.","Diagram: a decision path from your data to pgvector in Postgres, a search engine with vector fields, or a dedicated vector database.",{"slug":732,"published":5,"minutes":577,"category":7,"tags":733,"keywords":737,"about":747,"sources":755,"cover":792,"og":793,"expertise":71,"locales":794,"lang":73,"title":795,"description":796,"coverAlt":797},"graphrag-knowledge-graph-rag",[734,735,736,580],"GraphRAG","Knowledge graphs","LightRAG",[734,738,739,740,741,742,743,744,745,746],"knowledge graph RAG","GraphRAG vs vector RAG","Microsoft GraphRAG explained","LightRAG vs GraphRAG","GraphRAG global vs local search","GraphRAG indexing cost","when to use GraphRAG","multi-hop RAG","LazyGraphRAG",[748,751,752],{"name":749,"url":750},"Knowledge graph","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FKnowledge_graph",{"name":26,"url":27},{"name":753,"url":754},"Leiden algorithm","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FLeiden_algorithm",[756,759,762,765,768,771,774,777,780,783,786,789],{"title":757,"url":758},"Edge et al.: From Local to Global: A Graph RAG Approach to Query-Focused Summarization (arXiv:2404.16130)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2404.16130",{"title":760,"url":761},"Microsoft GraphRAG documentation: overview","https:\u002F\u002Fmicrosoft.github.io\u002Fgraphrag\u002F",{"title":763,"url":764},"Microsoft GraphRAG documentation: default dataflow","https:\u002F\u002Fmicrosoft.github.io\u002Fgraphrag\u002Findex\u002Fdefault_dataflow\u002F",{"title":766,"url":767},"Microsoft GraphRAG documentation: global search","https:\u002F\u002Fmicrosoft.github.io\u002Fgraphrag\u002Fquery\u002Fglobal_search\u002F",{"title":769,"url":770},"Microsoft GraphRAG documentation: local search","https:\u002F\u002Fmicrosoft.github.io\u002Fgraphrag\u002Fquery\u002Flocal_search\u002F",{"title":772,"url":773},"Microsoft GraphRAG documentation: DRIFT search","https:\u002F\u002Fmicrosoft.github.io\u002Fgraphrag\u002Fquery\u002Fdrift_search\u002F",{"title":775,"url":776},"Microsoft GraphRAG documentation: getting started","https:\u002F\u002Fmicrosoft.github.io\u002Fgraphrag\u002Fget_started\u002F",{"title":778,"url":779},"Microsoft Research: LazyGraphRAG, setting a new standard for quality and cost (25 November 2024)","https:\u002F\u002Fwww.microsoft.com\u002Fen-us\u002Fresearch\u002Fblog\u002Flazygraphrag-setting-a-new-standard-for-quality-and-cost\u002F",{"title":781,"url":782},"Guo et al.: LightRAG: Simple and Fast Retrieval-Augmented Generation (arXiv:2410.05779)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2410.05779",{"title":784,"url":785},"HKUDS\u002FLightRAG on GitHub","https:\u002F\u002Fgithub.com\u002FHKUDS\u002FLightRAG",{"title":787,"url":788},"Gutiérrez et al.: HippoRAG: Neurobiologically Inspired Long-Term Memory for Large Language Models (arXiv:2405.14831)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2405.14831",{"title":790,"url":791},"Xiang et al.: When to use Graphs in RAG: A Comprehensive Analysis for Graph Retrieval-Augmented Generation (arXiv:2506.05690)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2506.05690","\u002Fimages\u002Fblog\u002Fgraphrag-knowledge-graph-rag\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fgraphrag-knowledge-graph-rag\u002Fog.jpg",[73,74,75],"GraphRAG and knowledge-graph RAG: when a graph beats vector search","What Microsoft GraphRAG and LightRAG really do, what indexing costs, and when a knowledge graph beats vector RAG: multi-hop, global questions, product catalogues.","Diagram: a knowledge graph hub linked to entities, communities, local search, global search and product parts.",{"slug":799,"published":5,"minutes":577,"category":7,"tags":800,"keywords":805,"about":816,"sources":824,"cover":863,"og":864,"expertise":865,"locales":866,"lang":73,"title":867,"description":868,"coverAlt":869},"semantic-product-search-b2b",[801,652,802,803,804],"B2B search","Semantic search","Spryker","OpenSearch",[806,807,808,809,810,811,812,813,814,815],"semantic product search B2B","B2B ecommerce search","hybrid search BM25 vector","part number search ecommerce","Spryker search Elasticsearch","OpenSearch hybrid search RRF","multilingual product search German English Hungarian","zero results rate site search","LLM query understanding ecommerce","AI product search for B2B shops",[817,819,822],{"name":802,"url":818},"https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FSemantic_search",{"name":820,"url":821},"Elasticsearch","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FElasticsearch",{"name":804,"url":823},"https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FOpenSearch",[825,828,831,833,836,839,842,845,848,851,854,857,860],{"title":826,"url":827},"Baymard Institute: E-commerce search query types","https:\u002F\u002Fbaymard.com\u002Fblog\u002Fecommerce-search-query-types",{"title":829,"url":830},"Elastic Search Labs: Hybrid search in Elasticsearch","https:\u002F\u002Fwww.elastic.co\u002Fsearch-labs\u002Fblog\u002Fhybrid-search-elasticsearch",{"title":832,"url":721},"Elasticsearch documentation: kNN query (pre-filters and post-filters)",{"title":834,"url":835},"Elasticsearch documentation: Semantic reranking","https:\u002F\u002Fwww.elastic.co\u002Fdocs\u002Fsolutions\u002Fsearch\u002Franking\u002Fsemantic-reranking",{"title":837,"url":838},"Elasticsearch documentation: Word delimiter graph token filter","https:\u002F\u002Fwww.elastic.co\u002Fdocs\u002Freference\u002Ftext-analysis\u002Fanalysis-word-delimiter-graph-tokenfilter",{"title":840,"url":841},"Elasticsearch documentation: Synonym graph token filter","https:\u002F\u002Fwww.elastic.co\u002Fdocs\u002Freference\u002Ftext-analysis\u002Fanalysis-synonym-graph-tokenfilter",{"title":843,"url":844},"OpenSearch documentation: Score ranker processor (RRF)","https:\u002F\u002Fdocs.opensearch.org\u002Flatest\u002Fsearch-plugins\u002Fsearch-pipelines\u002Fscore-ranker-processor\u002F",{"title":846,"url":847},"OpenSearch documentation: Normalization processor","https:\u002F\u002Fdocs.opensearch.org\u002Flatest\u002Fsearch-plugins\u002Fsearch-pipelines\u002Fnormalization-processor\u002F",{"title":849,"url":850},"Spryker documentation: Search feature overview","https:\u002F\u002Fdocs.spryker.com\u002Fdocs\u002Fpbc\u002Fall\u002Fsearch\u002Flatest\u002Fbase-shop\u002Fsearch-feature-overview\u002Fsearch-feature-overview",{"title":852,"url":853},"Spryker documentation: Migrate from OpenSearch 1.3 to 3.5","https:\u002F\u002Fdocs.spryker.com\u002Fdocs\u002Fpbc\u002Fall\u002Fsearch\u002Flatest\u002Fbase-shop\u002Finstall-and-upgrade\u002Fmigrate-from-opensearch-1.3-to-3.5.html",{"title":855,"url":856},"Instacart via ZenML: Rebuilding query understanding for e-commerce search with LLMs","https:\u002F\u002Fwww.zenml.io\u002Fllmops-database\u002Frebuilding-query-understanding-for-e-commerce-search-with-llms",{"title":858,"url":859},"arXiv: M3-Embedding, multilingual, multi-functionality, multi-granularity text embeddings","https:\u002F\u002Farxiv.org\u002Fabs\u002F2402.03216",{"title":861,"url":862},"Algolia documentation: Search analytics metrics","https:\u002F\u002Fwww.algolia.com\u002Fdoc\u002Fguides\u002Fsearch-analytics\u002Fconcepts\u002Fmetrics\u002F","\u002Fimages\u002Fblog\u002Fsemantic-product-search-b2b\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fsemantic-product-search-b2b\u002Fog.jpg","b2b-ecommerce-developer",[73,74,75],"Semantic product search for B2B shops: part numbers, hybrid retrieval and what to measure","How to add semantic search to a B2B shop without breaking part-number search: hybrid BM25 and vectors, filters, DE\u002FEN\u002FHU, LLM query parsing, reranking and metrics.","Diagram: a search query is split into an identifier lane, lexical BM25 and vector kNN, fused, reranked and returned as results.",1791009037135]