[{"data":1,"prerenderedAt":623},["ShallowReactive",2],{"tool-ragas-en":3},{"slug":4,"published":5,"minutes":6,"category":7,"tags":8,"keywords":14,"about":22,"sources":25,"cover":43,"og":44,"expertise":45,"locales":46,"lang":47,"title":50,"description":51,"coverAlt":52,"url":53,"pricing":54,"kind":55,"metaTitle":56,"takeaways":57,"faq":63,"toc":76,"blocks":101,"others":387},"ragas","2026-05-20",10,"llmops",[9,10,11,12,13],"RAG evaluation","LLM-as-judge","Test sets","CI quality","Ragas",[4,15,16,17,18,19,20,21],"ragas tutorial","ragas vs deepeval","llm evaluation metrics","faithfulness context precision","rag evaluation framework","llm as a judge","ragas pricing",[23],{"name":13,"url":24},"https:\u002F\u002Fwww.ragas.io\u002F",[26,29,32,35,38,41],{"title":27,"url":28},"Ragas documentation: introduction","https:\u002F\u002Fdocs.ragas.io\u002Fen\u002Fstable\u002F",{"title":30,"url":31},"Ragas documentation: quick start","https:\u002F\u002Fdocs.ragas.io\u002Fen\u002Fstable\u002Fgetstarted\u002Fquickstart\u002F",{"title":33,"url":34},"Ragas documentation: list of available metrics","https:\u002F\u002Fdocs.ragas.io\u002Fen\u002Fstable\u002Fconcepts\u002Fmetrics\u002Favailable_metrics\u002F",{"title":36,"url":37},"Ragas on GitHub","https:\u002F\u002Fgithub.com\u002Fvibrantlabsai\u002Fragas",{"title":39,"url":40},"Ragas on PyPI","https:\u002F\u002Fpypi.org\u002Fproject\u002Fragas\u002F",{"title":42,"url":24},"Ragas website","\u002Fimages\u002Fblog\u002Fragas\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fragas\u002Fog.jpg","ai-engineer",[47,48,49],"en","de","hu","Ragas review: the default RAG evaluation framework","Ragas scores RAG and agent pipelines with faithfulness, context precision and context recall, generates test data and records experiments. Apache-2.0, free library, judge model billed separately.","Diagram of a Ragas experiment loop: a frozen test set feeds a run of the application, metrics call a judge model per row, scores arrive with written reasons, and the result is compared with the baseline.","https:\u002F\u002Fdocs.ragas.io","Apache-2.0 · paid platform","Evaluation framework","Ragas review: metrics, experiments, judge cost · Balázs Csorba",[58,59,60,61,62],"Ragas is an Apache-2.0 Python library at version 0.4.3 whose metrics, faithfulness, context precision and context recall, have become the standard vocabulary of RAG evaluation.","Most metrics call a judge model, so a run costs money and inherits the judge's blind spots; pin the model version before comparing two runs.","The library stores no traces and draws no dashboards, so teams that need production observability pair it with a tool such as Langfuse or build the store themselves.","The hosted platform has no published price list, which makes the judge model, not the licence, the number to budget for.","Reference-free metrics measure whether an answer is supported, not whether it is correct; keep a small labelled set for correctness.",[64,67,70,73],{"q":65,"a":66},"Is Ragas free to use?","The Python library is Apache-2.0 and free, including metrics, synthetic test-set generation and the experiments workflow. The hosted platform and enterprise support are sold directly by VibrantLabs, and no prices are published for them.",{"q":68,"a":69},"Which Ragas metrics need a golden set?","Context recall and factual correctness compare against a reference answer, while faithfulness, response relevancy and the aspect critic work from the question, the context and the response alone. Most teams start reference-free and add labels only for the rows they keep getting wrong.",{"q":71,"a":72},"Does Ragas replace an observability tool?","No. It scores a dataset you provide and writes results to CSV; it does not collect traces, sample production traffic or render dashboards. Use it as the metric layer under something that stores runs.",{"q":74,"a":75},"How much does a Ragas evaluation cost to run?","Whatever the judge model charges. Faithfulness issues several LLM calls per row because it checks claim by claim, so a 500-row set across four metrics runs to thousands of calls; caching and a smaller judge model are the levers.",[77,80,83,86,89,92,95,98],{"id":78,"title":79},"what-it-is","What it is",{"id":81,"title":82},"how-it-works","How it works",{"id":84,"title":85},"getting-started","Getting started",{"id":87,"title":88},"performance","Performance and cost",{"id":90,"title":91},"pricing","Pricing",{"id":93,"title":94},"where-it-shingles","Where it shingles",{"id":96,"title":97},"verdict","Verdict",{"id":99,"title":100},"sources","Sources",[102,111,114,117,125,147,148,151,160,163,164,167,169,192,199,200,203,250,253,261,264,265,268,278,279,286,331,334,344,345,348,361,365,366],{"type":103,"content":104},"paragraph",[105,106,110],"Ragas is an Apache-2.0 Python library for scoring LLM applications, and it has quietly become the vocabulary RAG evaluation is argued in: faithfulness, context precision and context recall all come out of its metric set. Its worth as a ",{"tag":107,"children":108},"strong",[109],"first evaluation tool for a retrieval pipeline is high; its worth as a production monitoring system is close to zero",", because it computes numbers over a dataset you hand it and leaves storage, dashboards, sampling and alerting to somebody else.",{"type":103,"content":112},[113],"It sits between the retrieval stack and the build: after LlamaIndex, LangChain or a hand-written retriever produces answers, Ragas scores them and the numbers end up next to the commit. It competes with DeepEval, TruLens and Phoenix, and with the evaluation features inside Langfuse, LangSmith and Braintrust; the difference is that Ragas stays a library, with no service to host and no vendor holding your test data.",{"type":115,"level":116,"id":78,"text":79},"heading",2,{"type":103,"content":118},[119,120,124],"The package installs from PyPI as ",{"tag":121,"children":122},"code",[123],"pip install ragas"," and runs inside the process; version 0.4.3 was published on 13 January 2026 and requires Python 3.9 or newer. Metrics come in two kinds: LLM-based, where a judge model reads a row and returns a value with a written reason, and traditional, where strings are compared directly, such as BLEU, ROUGE, exact match and semantic similarity. Around that sits an experiments workflow, dataset, run, record, compare, which is what turns a pile of metric functions into a repeatable evaluation.",{"type":126,"ordered":127,"items":128},"list",false,[129,131,133,135,137,139,145],[130],"Licence and ownership: Apache-2.0, maintained by VibrantLabs; the repository shows about 16,000 stars and 1,700 forks.",[132],"RAG metrics: context precision, context recall, context entities recall, noise sensitivity, response relevancy and faithfulness, plus multimodal variants of faithfulness and relevance.",[134],"Agent metrics: tool call accuracy, tool call F1, agent goal accuracy and topic adherence.",[136],"Grounding and comparison: factual correctness, semantic similarity, BLEU, CHRF, ROUGE, string presence and exact match.",[138],"SQL and summarisation: execution-based DataCompy scoring, SQL query equivalence and a summarisation score.",[140,141,144],"Test data and scaffolding: synthetic test-set generation from your own documents, and a ",{"tag":121,"children":142},[143],"ragas quickstart rag_eval"," template that writes a runnable project.",[146],"Output: experiment results as CSV files in your repository, diffable like any other build artefact.",{"type":115,"level":116,"id":81,"text":82},{"type":103,"content":149},[150],"A run walks the rows one at a time. Each row carries the user input, the response the application produced and whatever context the retriever supplied. A metric takes that row, makes one or more LLM calls with a prompt that spells out the criterion, and returns a value together with the judge's reason; traditional metrics skip the call and compare strings. Rows are scored independently, so the set parallelises cleanly and any single row can be repeated when the average looks wrong.",{"type":152,"attrs":153,"inner":157,"caption":158},"diagram",{"viewBox":154,"role":155,"aria-labelledby":156},"0 0 760 215","img","d-ragas-t d-ragas-d","\u003Ctitle id=\"d-ragas-t\">How a Ragas experiment runs\u003C\u002Ftitle>\u003Cdesc id=\"d-ragas-d\">A five-stage loop. A frozen test set feeds a run of the application, metrics call a judge model over each row, the scores arrive with written reasons, and the experiment is compared with the baseline before the next build runs the same rows again.\u003C\u002Fdesc>\u003Cdefs>\u003Cmarker id=\"ah-r\" viewBox=\"0 0 10 10\" refX=\"9\" refY=\"5\" markerWidth=\"7\" markerHeight=\"7\" orient=\"auto-start-reverse\">\u003Cpath d=\"M0 0L10 5L0 10z\" class=\"d-head\" \u002F>\u003C\u002Fmarker>\u003Cmarker id=\"ah-ra\" viewBox=\"0 0 10 10\" refX=\"9\" refY=\"5\" markerWidth=\"7\" markerHeight=\"7\" orient=\"auto-start-reverse\">\u003Cpath d=\"M0 0L10 5L0 10z\" class=\"d-head-accent\" \u002F>\u003C\u002Fmarker>\u003C\u002Fdefs>\u003Ctext x=\"10\" y=\"26\" class=\"d-title\">one row at a time\u003C\u002Ftext>\u003Crect x=\"10\" y=\"56\" width=\"125\" height=\"66\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"72\" y=\"86\" text-anchor=\"middle\" class=\"d-text\">test set\u003C\u002Ftext>\u003Ctext x=\"72\" y=\"108\" text-anchor=\"middle\" class=\"d-small\">queries, labels\u003C\u002Ftext>\u003Crect x=\"165\" y=\"56\" width=\"125\" height=\"66\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"227\" y=\"86\" text-anchor=\"middle\" class=\"d-text\">run app\u003C\u002Ftext>\u003Ctext x=\"227\" y=\"108\" text-anchor=\"middle\" class=\"d-small\">RAG pipeline\u003C\u002Ftext>\u003Crect x=\"320\" y=\"56\" width=\"125\" height=\"66\" rx=\"10\" class=\"d-accent\" \u002F>\u003Ctext x=\"382\" y=\"86\" text-anchor=\"middle\" class=\"d-text\">metrics\u003C\u002Ftext>\u003Ctext x=\"382\" y=\"108\" text-anchor=\"middle\" class=\"d-small\">judge calls\u003C\u002Ftext>\u003Crect x=\"475\" y=\"56\" width=\"125\" height=\"66\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"537\" y=\"86\" text-anchor=\"middle\" class=\"d-text\">scores\u003C\u002Ftext>\u003Ctext x=\"537\" y=\"108\" text-anchor=\"middle\" class=\"d-small\">value + reason\u003C\u002Ftext>\u003Crect x=\"630\" y=\"56\" width=\"125\" height=\"66\" rx=\"10\" class=\"d-gold\" \u002F>\u003Ctext x=\"692\" y=\"86\" text-anchor=\"middle\" class=\"d-text\">compare\u003C\u002Ftext>\u003Ctext x=\"692\" y=\"108\" text-anchor=\"middle\" class=\"d-small\">vs baseline\u003C\u002Ftext>\u003Cpath d=\"M137 89 H161\" class=\"d-line\" marker-end=\"url(#ah-r)\" \u002F>\u003Cpath d=\"M292 89 H316\" class=\"d-line\" marker-end=\"url(#ah-r)\" \u002F>\u003Cpath d=\"M447 89 H471\" class=\"d-line\" marker-end=\"url(#ah-r)\" \u002F>\u003Cpath d=\"M602 89 H626\" class=\"d-line\" marker-end=\"url(#ah-r)\" \u002F>\u003Cpath d=\"M692 124 V170 H72 V126\" class=\"d-line-accent\" marker-end=\"url(#ah-ra)\" \u002F>\u003Ctext x=\"382\" y=\"196\" text-anchor=\"middle\" class=\"d-small\">next experiment: same rows, new build\u003C\u002Ftext>",[159],"The loop only means anything if the rows and the judge stay fixed: change either, and two experiments are not comparable.",{"type":103,"content":161},[162],"The reason this style of evaluation spread is that the output is readable. A low faithfulness score arrives with the sentence the judge wrote about which claim was unsupported, which is a diff a human can act on rather than a number to argue about.",{"type":115,"level":116,"id":84,"text":85},{"type":103,"content":165},[166],"The smallest useful test is one metric on one row, which is also how the README introduces the library.",{"type":121,"code":168},"import asyncio\nfrom ragas.metrics.collections import AspectCritic\nfrom ragas.llms import llm_factory\n\n# The judge: llm_factory() takes OpenAI by default, Anthropic and others as options\nllm = llm_factory(\"gpt-4o\")\n\nmetric = AspectCritic(\n    name=\"summary_accuracy\",\n    definition=\"Verify if the summary is accurate and captures key information.\",\n    llm=llm,\n)\n\nrow = {\n    \"user_input\": \"summarise given text: the company reported an 8% rise in Q3 2024.\",\n    \"response\": \"The company saw an 8% increase in Q3 2024.\",\n}\n\nasync def main():\n    score = await metric.ascore(user_input=row[\"user_input\"], response=row[\"response\"])\n    print(score.value, score.reason)\n\nasyncio.run(main())",{"type":103,"content":170},[171,172,175,176,179,180,183,184,187,188,191],"For a whole project the scaffolding is faster: ",{"tag":121,"children":173},[174],"uvx ragas quickstart rag_eval"," writes a package with ",{"tag":121,"children":177},[178],"rag.py",", ",{"tag":121,"children":181},[182],"evals.py",", a datasets folder and an experiments folder, ",{"tag":121,"children":185},[186],"uv sync"," installs it, and ",{"tag":121,"children":189},[190],"uv run python evals.py"," loads the dataset, queries the application, scores every row and appends the result to CSV. The docs default the template to OpenAI and show the same three lines switched to Anthropic, Google Gemini, Ollama or any OpenAI-compatible endpoint.",{"type":193,"variant":194,"title":195,"body":196},"callout","tip","Pin the judge before you compare",[197],[198],"Two runs are only comparable when the metric implementation, the judge model and the rows are all identical. Pin the model version, keep the dataset in the repository, and treat a metric upgrade as a re-baseline rather than a regression.",{"type":115,"level":116,"id":87,"text":88},{"type":103,"content":201},[202],"Ragas adds no latency to the product; it adds cost and wall time to the build. The bill is rows times metrics times judge calls per row, and the expensive metrics are the ones that decompose an answer into claims and check each one separately.",{"type":204,"head":205,"rows":214},"table",[206,208,210,212],[207],"Metric",[209],"Judge work per row",[211],"Needs a reference answer",[213],"What a low score means",[215,224,233,242],[216,218,220,222],[217],"Faithfulness",[219],"Splits the answer into claims and checks each against the context",[221],"No",[223],"The answer says more than the retrieved text supports",[225,227,229,231],[226],"Context precision",[228],"Ranks the retrieved chunks against the question",[230],"Optional",[232],"Irrelevant chunks are being pushed into the prompt",[234,236,238,240],[235],"Context recall",[237],"Compares retrieved context with the reference answer",[239],"Yes",[241],"The retriever never saw part of what the answer needs",[243,245,247,248],[244],"Tool call accuracy",[246],"Compares the agent's tool call and arguments with the expected one",[239],[249],"The agent picks the wrong tool or passes broken arguments",{"type":103,"content":251},[252],"A 500-row set scored with four LLM metrics means thousands of judge calls per run. Teams keep it affordable by caching where the library allows it, by using a small judge model for the cheap metrics, and by running the full set only when a change is about to merge.",{"type":126,"ordered":127,"items":254},[255,257,259],[256],"Score the cheap metrics on every push; run faithfulness and the agent metrics on merge or on a schedule.",[258],"Keep the dataset small and adversarial rather than large and random: 200 rows that broke before beat 5,000 rows that never fail.",[260],"Generate synthetic rows for coverage, then freeze them; regenerating the test set every cycle makes every score look like noise.",{"type":103,"content":262},[263],"None of this is Ragas-specific: any LLM-as-judge setup bills by call and inherits the judge's blind spots. What Ragas decides is how much of that machinery you get for free, and how much you are left to build around it.",{"type":115,"level":116,"id":90,"text":91},{"type":103,"content":266},[267],"The library is free and Apache-2.0, with no seat count and no usage cap. The commercial side, the hosted platform and direct support, is sold by the maintainers through contact and office hours, and no price list is published for it. The cost that therefore matters is the judge model: every LLM metric is metered by whichever provider you point it at, and synthetic test-set generation runs on the same meter.",{"type":126,"ordered":127,"items":269},[270,272,274,276],[271],"Library: free, Apache-2.0, installed from PyPI, no account required.",[273],"Judge model: billed per call by your provider, typically the dominant line item of an evaluation budget.",[275],"Synthetic test data: also LLM work, billed the same way as scoring.",[277],"Hosted platform and support: sold direct, no public price; ask what is included before committing.",{"type":115,"level":116,"id":93,"text":94},{"type":103,"content":280},[281,282,285],"Ragas scores rows, not systems. It will not sample production traffic, keep traces or tell you that p95 latency moved, and it cannot tell you whether the test questions are representative, which is the failure that makes a green pipeline useless. The release cadence is fast, 0.3.0 in July 2025, 0.4.0 in December 2025, 0.4.3 in January 2026, so versions have to be pinned. The repository carried 381 open issues at the time of writing, which is the usual price of being the default, and the docs for ",{"tag":121,"children":283},[284],"stable"," have trailed the released API.",{"type":204,"head":287,"rows":295},[288,290,291,293],[289],"Tool",[79],[292],"Where it wins",[294],"What you give up",[296,304,313,322],[297,298,300,302],[13],[299],"Metric library with synthetic test data",[301],"The metric vocabulary and a runnable project from one command",[303],"No trace store, no dashboard, no production sampling",[305,307,309,311],[306],"DeepEval",[308],"Pytest-style metrics with a hosted platform behind it",[310],"Unit-test ergonomics and a long metric list",[312],"The platform part is a separate commercial product",[314,316,318,320],[315],"Langfuse",[317],"Open-source tracing with datasets and evaluations",[319],"Traces, prompts and scores in one place",[321],"Ragas-style metric maths is one feature among many",[323,325,327,329],[324],"Arize Phoenix",[326],"Open-source tracing and evaluation for LLM applications",[328],"Span-level debugging next to the scores",[330],"A narrower reference-free RAG metric set",{"type":103,"content":332},[333],"The real choice is where the numbers live. Ragas is the right engine if something else stores the runs, a CSV in the repository, a table in your warehouse, an observability tool that accepts imported scores. It is the wrong purchase for a team that wants to open a URL and see whether quality regressed.",{"type":193,"variant":335,"title":336,"body":337},"warn","Every row leaves your network",[338],[339,340,343],"An LLM metric sends the question, the retrieved chunks and the answer to whichever endpoint the judge uses. Where the corpus is confidential, that means a judge inside your own perimeter, the docs show Ollama and any OpenAI-compatible endpoint, or an accepted transfer to a third party. The library also reports anonymous usage data by default; ",{"tag":121,"children":341},[342],"RAGAS_DO_NOT_TRACK=true"," disables it and the collection code sits in the repository.",{"type":115,"level":116,"id":96,"text":97},{"type":103,"content":346},[347],"Ragas earns its default status for retrieval work: the metrics are well defined, the output carries reasons, and a project is one command away. It should be taken as a scoring engine inside a workflow you already own, not as an evaluation platform.",{"type":126,"ordered":349,"items":350},true,[351,353,355,357,359],[352],"Use it if you have a retrieval pipeline and no golden set yet: synthetic test-set generation plus reference-free metrics gives a first measurement in a day.",[354],"Use it if evaluation has to run in CI, where a CSV diff per pull request is exactly the right amount of signal.",[356],"Do not pick it as the only evaluation tool for an agent product: agent metrics exist, but trajectories, tool side effects and cost per task need a harness that stores runs.",[358],"Do not pick it when an answer has to be justified to a regulator or a customer from stored evidence; nothing is persisted unless you persist it.",[360],"Budget the judge model before the licence: a pinned small model for the cheap metrics and a capable one for faithfulness is the usual split.",{"type":362,"content":363},"quote",[364],"Ragas is the cheapest way to stop arguing about whether a change helped: pin the judge, freeze the rows, read the diff. Everything it does not do is work you still have to pay for.",{"type":115,"level":116,"id":99,"text":100},{"type":126,"ordered":349,"items":367},[368,372,375,378,381,384],[369],{"tag":370,"href":28,"children":371},"a",[27],[373],{"tag":370,"href":31,"children":374},[30],[376],{"tag":370,"href":34,"children":377},[33],[379],{"tag":370,"href":37,"children":380},[36],[382],{"tag":370,"href":40,"children":383},[39],[385],{"tag":370,"href":24,"children":386},[42],[388,437,520,580],{"slug":389,"published":390,"minutes":391,"category":7,"tags":392,"keywords":398,"about":405,"sources":409,"cover":428,"og":429,"expertise":45,"locales":430,"lang":47,"title":431,"description":432,"coverAlt":433,"url":434,"pricing":435,"kind":436},"ollama","2026-09-29",11,[393,394,395,396,397],"Local inference","Open models","llama.cpp","GGUF","Model serving",[389,399,400,401,402,403,404],"ollama vs lm studio","ollama vs vllm","local llm runtime","gguf model server","ollama self hosting","ollama api",[406],{"name":407,"url":408},"Ollama (software)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FOllama",[410,413,416,419,422,425],{"title":411,"url":412},"Ollama API documentation","https:\u002F\u002Fdocs.ollama.com\u002Fapi",{"title":414,"url":415},"Ollama on GitHub, with the MIT LICENSE file","https:\u002F\u002Fgithub.com\u002Follama\u002Follama",{"title":417,"url":418},"Ollama terms of service, last updated May 2026","https:\u002F\u002Follama.com\u002Fterms",{"title":420,"url":421},"Ollama pricing, cloud plans and per-token model rates","https:\u002F\u002Follama.com\u002Fpricing",{"title":423,"url":424},"Hardware support: Nvidia, AMD, Metal and Vulkan","https:\u002F\u002Fdocs.ollama.com\u002Fgpu",{"title":426,"url":427},"OpenAI compatibility, including what is not supported","https:\u002F\u002Fdocs.ollama.com\u002Fapi\u002Fopenai-compatibility","\u002Fimages\u002Fblog\u002Follama\u002Fcover.webp","\u002Fimages\u002Fblog\u002Follama\u002Fog.jpg",[47,48,49],"Ollama review: the friendly way to run open models","Ollama serves open models over one HTTP API on your own hardware. What it does well, where throughput falls short, and what the MIT licence does not cover.","Abstract cover art for the Ollama review","https:\u002F\u002Follama.com","MIT · free for personal use","Local inference runtime",{"slug":438,"published":439,"minutes":391,"category":7,"tags":440,"keywords":446,"about":454,"sources":461,"cover":513,"og":514,"expertise":45,"locales":515,"lang":47,"title":516,"description":517,"coverAlt":518,"url":457,"pricing":519,"kind":441},"portkey","2026-09-28",[441,442,443,444,445],"LLM gateway","Guardrails","Routing","Observability","Cost control",[447,448,449,450,451,452,453],"portkey ai gateway","portkey vs litellm","llm gateway comparison","llm gateway latency overhead","llm guardrails gateway","self-hosted llm gateway","portkey pricing",[455,458],{"name":456,"url":457},"Portkey","https:\u002F\u002Fportkey.ai",{"name":459,"url":460},"API gateway","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FAPI_gateway",[462,465,468,471,474,477,480,483,486,489,492,495,498,501,504,507,510],{"title":463,"url":464},"Portkey docs: AI Gateway","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fproduct\u002Fai-gateway",{"title":466,"url":467},"Portkey docs: Getting started with the AI Gateway","https:\u002F\u002Fdocs.portkey.ai\u002Fdocs\u002Fguides\u002Fgetting-started\u002Fgetting-started-with-ai-gateway",{"title":469,"url":470},"Portkey docs: Gateway config object","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fapi-reference\u002Fconfig-object",{"title":472,"url":473},"Portkey docs: Guardrails","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fproduct\u002Fguardrails",{"title":475,"url":476},"Portkey docs: Guardrail endpoints and capabilities","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fproduct\u002Fguardrails\u002Fcapabilities",{"title":478,"url":479},"Portkey docs: Cache, simple and semantic","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fproduct\u002Fai-gateway\u002Fcache-simple-and-semantic",{"title":481,"url":482},"Portkey docs: Load balancing","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fproduct\u002Fai-gateway\u002Fload-balancing",{"title":484,"url":485},"Portkey docs: Enterprise hybrid deployment architecture","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fself-hosting\u002Fhybrid-deployments\u002Farchitecture",{"title":487,"url":488},"Portkey pricing","https:\u002F\u002Fportkey.ai\u002Fpricing",{"title":490,"url":491},"Portkey gateway on GitHub, MIT licensed","https:\u002F\u002Fgithub.com\u002FPortkey-AI\u002Fgateway",{"title":493,"url":494},"Portkey's own benchmark: gateway versus direct Bedrock","https:\u002F\u002Fgithub.com\u002FPortkey-AI\u002Fbenchmark-test",{"title":496,"url":497},"Portkey status page","https:\u002F\u002Fstatus.portkey.ai\u002F",{"title":499,"url":500},"Palo Alto Networks completes acquisition of Portkey, May 2026","https:\u002F\u002Fwww.paloaltonetworks.com\u002Fcompany\u002Fpress\u002F2026\u002Fpalo-alto-networks-completes-acquisition-of-portkey-to-secure-ai-agents",{"title":502,"url":503},"Palo Alto Networks: Prisma AIRS AI Gateway","https:\u002F\u002Fwww.paloaltonetworks.com\u002Fai-security\u002Fai-gateway",{"title":505,"url":506},"Cloudflare AI Gateway pricing","https:\u002F\u002Fdevelopers.cloudflare.com\u002Fai-gateway\u002Freference\u002Fpricing\u002F",{"title":508,"url":509},"LiteLLM pricing","https:\u002F\u002Fwww.litellm.ai\u002Fpricing",{"title":511,"url":512},"OpenRouter pricing","https:\u002F\u002Fopenrouter.ai\u002Fpricing","\u002Fimages\u002Fblog\u002Fportkey\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fportkey\u002Fog.jpg",[47,48,49],"Portkey: a production LLM gateway, reviewed for routing, guardrails and cost","Portkey puts retries, fallbacks, caching, guardrails and cost tracking behind one OpenAI-compatible endpoint. What the config object does well, what the gateway costs in latency, and when to self-host.","A request path from an application through the Portkey gateway to three model providers, with the guardrail verdict and the log written below the proxy.","Free · from $49 per month",{"slug":521,"published":522,"minutes":6,"category":7,"tags":523,"keywords":529,"about":537,"sources":548,"cover":573,"og":574,"expertise":45,"locales":575,"lang":47,"title":576,"description":577,"coverAlt":578,"url":539,"pricing":579,"kind":524},"langfuse","2026-08-13",[524,525,526,527,528],"LLM observability","Tracing","OpenTelemetry","Self-hosting","Evaluation",[521,530,531,532,533,534,535,536],"langfuse vs langsmith","llm tracing tool","self-hosted llm observability","langfuse pricing","opentelemetry llm traces","llm cost tracking","prompt versioning",[538,540,542,545],{"name":315,"url":539},"https:\u002F\u002Flangfuse.com",{"name":526,"url":541},"https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FOpenTelemetry",{"name":543,"url":544},"ClickHouse","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FClickHouse",{"name":546,"url":547},"Observability (software)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FObservability_(software)",[549,552,555,558,561,564,567,570],{"title":550,"url":551},"Langfuse documentation: observability and application tracing","https:\u002F\u002Flangfuse.com\u002Fdocs\u002Fobservability\u002Foverview",{"title":553,"url":554},"Langfuse documentation: get started with tracing","https:\u002F\u002Flangfuse.com\u002Fdocs\u002Fobservability\u002Fget-started",{"title":556,"url":557},"Langfuse pricing: cloud plans, billable units and worked examples","https:\u002F\u002Flangfuse.com\u002Fpricing",{"title":559,"url":560},"Langfuse pricing: self-hosted plans and the feature comparison","https:\u002F\u002Flangfuse.com\u002Fpricing-self-host",{"title":562,"url":563},"Self-host Langfuse: deployment options, containers and storage services","https:\u002F\u002Flangfuse.com\u002Fself-hosting",{"title":565,"url":566},"Langfuse changelog: v4 is live (17 August 2026)","https:\u002F\u002Flangfuse.com\u002Fchangelog\u002F2026-08-17-langfuse-v4",{"title":568,"url":569},"Langfuse blog: Langfuse joins ClickHouse (16 January 2026)","https:\u002F\u002Flangfuse.com\u002Fblog\u002Fjoining-clickhouse",{"title":571,"url":572},"GitHub: langfuse\u002Flangfuse, the platform repository","https:\u002F\u002Fgithub.com\u002Flangfuse\u002Flangfuse","\u002Fimages\u002Fblog\u002Flangfuse\u002Fcover.webp","\u002Fimages\u002Fblog\u002Flangfuse\u002Fog.jpg",[47,48,49],"Langfuse review: tracing, prompts and evals you can host yourself","Langfuse puts LLM traces, prompt versions and experiments on one MIT-licensed platform. What self-hosting really costs, how the unit pricing adds up, and where it loses.","A pipeline from a batched application event through the Langfuse web container and object storage into ClickHouse, with Redis and PostgreSQL alongside.","MIT · paid from $59 per month",{"slug":581,"published":582,"minutes":6,"category":7,"tags":583,"keywords":588,"about":595,"sources":599,"cover":615,"og":616,"expertise":45,"locales":617,"lang":47,"title":618,"description":619,"coverAlt":620,"url":621,"pricing":622,"kind":441},"openrouter","2026-07-23",[441,584,585,586,587],"Model routing","Fallbacks","OpenAI-compatible","Pay per token",[581,589,449,590,591,592,593,594],"openrouter vs litellm","openrouter pricing","openai compatible api gateway","llm fallback routing","multi model api gateway","byok llm routing",[596],{"name":597,"url":598},"OpenRouter","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FOpenRouter",[600,603,604,607,610,613],{"title":601,"url":602},"OpenRouter documentation: quickstart","https:\u002F\u002Fopenrouter.ai\u002Fdocs\u002Fquickstart",{"title":511,"url":512},{"title":605,"url":606},"OpenRouter documentation: model fallbacks","https:\u002F\u002Fopenrouter.ai\u002Fdocs\u002Fguides\u002Frouting\u002Fmodel-fallbacks",{"title":608,"url":609},"OpenRouter documentation: provider routing","https:\u002F\u002Fopenrouter.ai\u002Fdocs\u002Fguides\u002Frouting\u002Fprovider-selection",{"title":611,"url":612},"OpenRouter documentation index","https:\u002F\u002Fopenrouter.ai\u002Fdocs\u002Fllms.txt",{"title":614,"url":598},"Wikipedia: OpenRouter","\u002Fimages\u002Fblog\u002Fopenrouter\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fopenrouter\u002Fog.jpg",[47,48,49],"OpenRouter: one API key in front of every model you might call","OpenRouter puts 500+ models from 80+ providers behind one OpenAI-compatible endpoint, with fallbacks and pass-through pricing. What it costs, where it breaks.","Request path through OpenRouter: client, router, candidate providers, fallback list and the model that finally answers.","https:\u002F\u002Fopenrouter.ai","Pay per token, no subscription",1791383549083]