[{"data":1,"prerenderedAt":627},["ShallowReactive",2],{"tool-langsmith-en":3},{"slug":4,"published":5,"minutes":6,"category":7,"tags":8,"keywords":14,"about":22,"sources":31,"cover":53,"og":54,"expertise":55,"locales":56,"lang":57,"title":60,"description":61,"coverAlt":62,"url":25,"pricing":63,"kind":64,"metaTitle":65,"takeaways":66,"faq":72,"toc":88,"blocks":113,"others":397},"langsmith","2026-05-25",10,"llmops",[9,10,11,12,13],"Tracing","LLM evals","OpenTelemetry","Datasets","Retries",[4,15,16,17,18,19,20,21],"langsmith vs langfuse","llm observability tools","llm tracing cost","langsmith pricing","opentelemetry llm tracing","llm evaluation datasets","phoenix vs langsmith",[23,26,28],{"name":24,"url":25},"LangSmith","https:\u002F\u002Fwww.langchain.com\u002Flangsmith",{"name":11,"url":27},"https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FOpenTelemetry",{"name":29,"url":30},"ClickHouse","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FClickHouse",[32,35,38,41,44,47,50],{"title":33,"url":34},"LangSmith pricing","https:\u002F\u002Fwww.langchain.com\u002Fpricing",{"title":36,"url":37},"LangSmith documentation: administration overview","https:\u002F\u002Fdocs.langchain.com\u002Flangsmith\u002Fadministration-overview",{"title":39,"url":40},"LangSmith documentation: custom instrumentation","https:\u002F\u002Fdocs.langchain.com\u002Flangsmith\u002Fannotate-code",{"title":42,"url":43},"LangSmith documentation: trace with OpenTelemetry","https:\u002F\u002Fdocs.langchain.com\u002Flangsmith\u002Ftrace-with-opentelemetry",{"title":45,"url":46},"LangSmith documentation: evaluation concepts","https:\u002F\u002Fdocs.langchain.com\u002Flangsmith\u002Fevaluation-concepts",{"title":48,"url":49},"LangSmith documentation: self-hosted deployment","https:\u002F\u002Fdocs.langchain.com\u002Flangsmith\u002Fself-hosted",{"title":51,"url":52},"LangChain: introducing end-to-end OpenTelemetry support in LangSmith","https:\u002F\u002Fwww.langchain.com\u002Fblog\u002Fend-to-end-opentelemetry-langsmith","\u002Fimages\u002Fblog\u002Flangsmith\u002Fcover.webp","\u002Fimages\u002Fblog\u002Flangsmith\u002Fog.jpg","ai-engineer",[57,58,59],"en","de","hu","LangSmith: an observability product that grew into an agent platform","LangSmith review: per-trace billing, 14 and 180 day retention, OpenTelemetry ingestion, offline and online evals, and where the platform's pull towards an agent runtime shows up.","A LangSmith trace path: traced application code posts runs through an SDK background thread into an ingest queue, which writes traces to ClickHouse for the dashboard and the API.","Developer free · from $39 per month","LLM observability","LangSmith review: tracing, evals and the bill · Balázs Csorba",[67,68,69,70,71],"The billable unit is the trace, not the span or the token. Developer includes 5,000 base traces a month for one seat, Plus costs $39 a seat with 10,000 included, and the Developer plan without a payment method on file is capped at 5,000 traces a month.","Base traces are retained for 14 days and extended traces for 180, and online evaluators, automation rules and API feedback with extend_trace_retention upgrade traces to the more expensive tier unless you opt out.","Ingest limits are hourly and per plan: 50,000 events and 500 MB an hour on Developer without payment details, 250,000 events and 2.5 GB with them, and 500,000 events and 5 GB on Plus.","The rest of the platform is billed in LangChain Standard Units at $1.00 each, including Engine, Fleet, Sandboxes and the LLM Gateway, which makes a seat price a poor guide to the invoice.","OpenTelemetry ingestion works through the SDK or an OTLP endpoint, but a span whose parent never arrives is buffered and then silently dropped, which is the failure mode to watch in a partial fan-out.",[73,76,79,82,85],{"q":74,"a":75},"How much does LangSmith cost for a small team?","The Developer plan is free for one seat with 5,000 base traces a month included, and without a payment method on file it is also capped at 5,000 traces a month. Plus is $39 per seat a month with 10,000 base traces included and unlimited extra seats. Enterprise is quoted, and it is the only tier that offers workspace role-based access control and self-hosted deployment.",{"q":77,"a":78},"Do I need LangChain to use LangSmith?","No. The langsmith SDK has a traceable decorator for Python, TypeScript, Kotlin and Java, plus a low-level RunTree API and a REST ingest path, so any code can be traced. LangChain and LangGraph just get the instrumentation for free. LangSmith also ingests OpenTelemetry spans, which the docs recommend for applications that already emit OTLP.",{"q":80,"a":81},"What happens to my traces after two weeks?","Base retention is 14 days, after which traces are no longer reachable in the UI or the API and the associated inputs and outputs are deleted within a day, while some trace metadata is kept for analytics and billing. Extended retention is 180 days and costs more; since 14 September 2026 that 180 days is the maximum for SaaS customers.",{"q":83,"a":84},"Can LangSmith run inside my own infrastructure?","Yes, but only as an Enterprise add-on with a licence key. A self-hosted instance runs the frontend, backend, platform backend, playground, queue and a code-execution service on top of ClickHouse for traces, PostgreSQL for operational data and Redis or Valkey for queues, with optional blob storage. LangChain recommends external database services in production rather than the bundled ones.",{"q":86,"a":87},"Will LangSmith train on my data?","No. The pricing FAQ states that LangSmith does not use your data to train models and that traces, prompts and outputs stay private to your organisation. That is a contractual answer rather than a technical control, so a self-hosted deployment is the option for teams who need the data inside their own perimeter.",[89,92,95,98,101,104,107,110],{"id":90,"title":91},"what-it-is","What it is",{"id":93,"title":94},"how-it-works","How it works",{"id":96,"title":97},"getting-started","Getting started",{"id":99,"title":100},"what-it-costs","What it costs",{"id":102,"title":103},"instrumentation-choices","Native tracing or OpenTelemetry",{"id":105,"title":106},"where-it-shingles","Where it shingle",{"id":108,"title":109},"verdict","Verdict",{"id":111,"title":112},"sources","Sources",[114,118,121,124,127,143,144,147,156,159,160,167,170,181,187,198,199,202,256,259,262,271,272,281,283,286,295,296,299,345,348,349,352,365,371,372],{"type":115,"content":116},"paragraph",[117],"LangSmith is a hosted tracing and evaluation platform for LLM applications: every call your agent makes becomes a tree of runs with inputs, outputs, latency and token counts that you can search, score and compare. The position after reading the documentation: it is the most capable tool of its kind, and the most expensive one to understand, because it is no longer only an observability product.",{"type":115,"content":119},[120],"It competes with Langfuse, Arize Phoenix and Helicone for the observability budget, and increasingly with its own runtime, because the same vendor now sells Deployment, Studio, the Engine, Fleet, Sandboxes and an LLM Gateway on the same pricing page. That expansion is the single most important thing to weigh, since each extra service is metered separately.",{"type":122,"level":123,"id":90,"text":91},"heading",2,{"type":115,"content":125},[126],"The platform has four parts that share one bill: tracing, evaluation, prompt management and an agent runtime. The resource model matters more than the features. An organisation holds workspaces, workspaces hold tracing projects, datasets, annotation queues and prompts, and every trace lives in one project. The core facts:",{"type":128,"ordered":129,"items":130},"list",false,[131,133,135,137,139,141],[132],"A trace is one execution, made of nested runs; a run is created and then updated as the work progresses, which is why an event limit and a trace limit are different numbers.",[134],"Tracing works through the langsmith SDK decorators, a REST ingest API or OpenTelemetry spans from any instrumented application.",[136],"Offline evals run against datasets of examples with reference outputs; online evals score live traces without references.",[138],"Evaluators can be code, LLM-as-judge, a typed decision model, pairwise or human, and one evaluator can be attached to several projects.",[140],"Annotation queues, dataset versions and splits, a prompt registry with commit tags and a playground are all included.",[142],"Self-hosted deployment exists, but only as an Enterprise add-on behind a licence key.",{"type":122,"level":123,"id":93,"text":94},{"type":115,"content":145},[146],"Instrumentation runs in your process and posts runs to LangSmith over HTTPS. The SDK sends from a background thread and batches up to 100 runs from one session into a single API call, so tracing does not sit in the request path. A server-side queue then handles ingestion, retries and integrity checks before writing into the trace store.",{"type":148,"attrs":149,"inner":153,"caption":154},"diagram",{"viewBox":150,"role":151,"aria-labelledby":152},"0 0 720 210","img","ls-diagram-t ls-diagram-d","\u003Ctitle id=\"ls-diagram-t\">The path of a LangSmith trace\u003C\u002Ftitle>\u003Cdesc id=\"ls-diagram-d\">Traced application code creates runs. The SDK posts them in batches from a background thread, because a rate limit stops the first 5000 posts to the runs endpoint in a minute. An ingest queue retries and stores the runs in ClickHouse. Dashboards, monitors and the query API read from that store.\u003C\u002Fdesc>\u003Cdefs>\u003Cmarker id=\"ah-ls\" viewBox=\"0 0 10 10\" refX=\"9\" refY=\"5\" markerWidth=\"7\" markerHeight=\"7\" orient=\"auto-start-reverse\">\u003Cpath d=\"M0 0L10 5L0 10z\" class=\"d-head\" \u002F>\u003C\u002Fmarker>\u003C\u002Fdefs>\u003Ctext x=\"16\" y=\"26\" class=\"d-title\">one request, one trace tree\u003C\u002Ftext>\u003Crect x=\"16\" y=\"70\" width=\"156\" height=\"60\" rx=\"10\" class=\"d-accent\" \u002F>\u003Ctext x=\"94\" y=\"97\" text-anchor=\"middle\" class=\"d-text\">your app\u003C\u002Ftext>\u003Ctext x=\"94\" y=\"117\" text-anchor=\"middle\" class=\"d-small\">traceable\u003C\u002Ftext>\u003Crect x=\"208\" y=\"70\" width=\"170\" height=\"60\" rx=\"10\" class=\"d-mint\" \u002F>\u003Ctext x=\"293\" y=\"97\" text-anchor=\"middle\" class=\"d-text\">SDK thread\u003C\u002Ftext>\u003Ctext x=\"293\" y=\"117\" text-anchor=\"middle\" class=\"d-small\">batched posts\u003C\u002Ftext>\u003Crect x=\"414\" y=\"70\" width=\"110\" height=\"60\" rx=\"10\" class=\"d-gold\" \u002F>\u003Ctext x=\"469\" y=\"97\" text-anchor=\"middle\" class=\"d-text\">queue\u003C\u002Ftext>\u003Ctext x=\"469\" y=\"117\" text-anchor=\"middle\" class=\"d-small\">retry\u003C\u002Ftext>\u003Crect x=\"560\" y=\"70\" width=\"144\" height=\"60\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"632\" y=\"97\" text-anchor=\"middle\" class=\"d-text\">trace store\u003C\u002Ftext>\u003Ctext x=\"632\" y=\"117\" text-anchor=\"middle\" class=\"d-small\">ClickHouse\u003C\u002Ftext>\u003Cpath d=\"M172 100L204 100\" class=\"d-line\" marker-end=\"url(#ah-ls)\" \u002F>\u003Cpath d=\"M378 100L410 100\" class=\"d-line\" marker-end=\"url(#ah-ls)\" \u002F>\u003Cpath d=\"M524 100L556 100\" class=\"d-line\" marker-end=\"url(#ah-ls)\" \u002F>\u003Ctext x=\"178\" y=\"152\" class=\"d-label\">429 after 5,000 posts per minute\u003C\u002Ftext>\u003Ctext x=\"16\" y=\"188\" class=\"d-small\">dashboards, monitors and the query API all read from the same store\u003C\u002Ftext>",[155],"Because the queue is asynchronous, a trace that is accepted with a 200 can still fail to arrive, and a short-lived process can exit before its runs are posted.",{"type":115,"content":157},[158],"That queue is why the ingestion limits are expressed as a fixed window rather than a smooth average, and why a 429 is a normal event to handle rather than an outage. The load balancer enforces fixed per-minute caps on every plan: 5000 POST or PATCH requests to the runs endpoints, 5000 to feedbacks, 2000 for any other endpoint, and 30 deletes. The SDK batches, which is what keeps a busy application under those numbers.",{"type":122,"level":123,"id":96,"text":97},{"type":115,"content":161},[162,163,164,165,166],"Two environment variables switch tracing on without touching code, which matters for local development: ","LANGSMITH_TRACING"," gates the decorator and the context manager, and ","LANGSMITH_PROJECT"," names the destination project, defaulting to default. The minimal Python setup traces a pipeline as nested runs:",{"type":168,"code":169},"code","import asyncio\n\nfrom langsmith import Client, traceable\nfrom openai import AsyncOpenAI\n\nclient = Client()\nllm = AsyncOpenAI()\n\n\n@traceable(run_type=\"retriever\", name=\"retrieve_docs\")\nasync def retrieve_docs(question: str) -> list[str]:\n    return [\"Annual report: revenue up 12 percent.\"]\n\n\n@traceable(run_type=\"llm\", name=\"answer\")\nasync def answer(question: str, context: list[str]) -> str:\n    reply = await llm.chat.completions.create(\n        model=\"gpt-5.4-mini\",\n        messages=[{\"role\": \"user\", \"content\": f\"{question}\\n{chr(10).join(context)}\"}],\n    )\n    return reply.choices[0].message.content\n\n\n@traceable(name=\"support_agent\")\nasync def support_agent(question: str) -> str:\n    return await answer(question, await retrieve_docs(question))\n\n\nasync def main() -> None:\n    try:\n        print(await support_agent(\"How did revenue move?\"))\n    finally:\n        await client.flush()  # background thread must finish before exit\n\n\nasyncio.run(main())\n",{"type":115,"content":171},[172,173,174,175,176,177,178,179,180],"The decorator propagates context, so the three functions appear as a tree without any manual parent wiring, and ","run_type"," decides how the dashboard renders a node: ","llm"," gives token counts and latency, ","retriever"," and ","tool"," mark the other kinds of step. The same client runs offline evaluations against a dataset, which is where the value starts to compound.",{"type":115,"content":182},[183,184,185,186],"The same API runs the eval loop. ","evaluate"," takes a target function, a dataset and a list of evaluators, produces an experiment with a run per example, and can be driven from CI:"," Evaluate each committed prompt change against the tagged dataset version and fail the build if the groundedness score drops by more than a point. LangSmith versions datasets automatically when examples change, so a tag can pin a CI run to one state of the data. The documentation is blunt about the starting point: write five to ten curated examples of good output before writing any evaluator.",{"type":188,"variant":189,"title":190,"body":191},"callout","warn","The trace tree is not in your process",[192],[193,194,197],"Tracing runs in a background thread so a slow or unavailable LangSmith does not block your request path, which is the right trade. The consequence is that a process can exit before its runs are posted. The SDK has a ",{"tag":168,"children":195},[196],"flush"," method, and the documentation calls it out explicitly for short-lived jobs and serverless handlers. Skipping it is the most common reason a test or a batch job produces an empty trace list.",{"type":122,"level":123,"id":99,"text":100},{"type":115,"content":200},[201],"Seats are the visible price and traces are the metered one. The Developer plan is free for a single seat with 5,000 base traces a month included; Plus is $39 per seat a month with 10,000 included and unlimited extra seats; Enterprise is quoted and adds self-hosted and hybrid deployment, custom SSO, and attribute-based and role-based access control.",{"type":203,"head":204,"rows":215},"table",[205,207,209,211,213],[206],"Plan",[208],"Seats",[210],"Traces included",[212],"Ingest ceiling per hour",[214],"Retention",[216,227,236,247],[217,219,221,223,225],[218],"Developer, no payment details",[220],"1",[222],"5,000 a month, and a monthly cap of 5,000",[224],"50,000 events and 500 MB",[226],"14 days, extended by upgrade",[228,230,231,233,235],[229],"Developer with payment details",[220],[232],"5,000 a month",[234],"250,000 events and 2.5 GB",[226],[237,239,241,243,245],[238],"Plus",[240],"Unlimited, $39 each",[242],"10,000 a month",[244],"500,000 events and 5 GB",[246],"14 or 180 days",[248,250,252,253,254],[249],"Enterprise",[251],"Custom",[251],[251],[255],"Up to 180 days, configurable",{"type":115,"content":257},[258],"An event is the creation or the update of a run, so a run created and then patched inside the same clock hour counts twice against the hourly limit, and a 2 MB run later updated to 3 MB counts 5 MB against the ingest volume limit. That is the mechanic to model in a capacity plan: trace shape, not request count, drives the ceiling.",{"type":115,"content":260},[261],"The rest of the platform is billed in LangChain Standard Units, with one LSU priced at $1.00. The Engine is scheduled every six hours and a single run is estimated at 7 to 45 LSU depending on trace volume and issue count. Fleet includes 7 LSU on Developer and 37 LSU on Plus, Sandboxes 8 LSU, and a perceived-error evaluator run 0.015 LSU. A $39 seat therefore says very little about the invoice once the runtime is in use.",{"type":188,"variant":189,"title":263,"body":264},"Retention upgrades quietly, and there is a catch",[265],[266,267,270],"Base traces live 14 days, extended 180. Online evaluators, automation rules matching any run in a trace, and feedback posted through the API with an explicit ",{"tag":168,"children":268},[269],"extend_trace_retention"," flag all upgrade a trace to the more expensive tier unless you opt out, and since 14 September 2026 the 180-day maximum for SaaS is the ceiling. The nastier edge is the extended-trace usage limit: once it is reached, rule matching, API feedback and annotation queues are switched off entirely, because each of them could create another extended trace.",{"type":122,"level":123,"id":102,"text":103},{"type":115,"content":273},[274,275,276,277,278,279,280],"LangSmith ingests OpenTelemetry spans two ways. With the SDK integration, ","LANGSMITH_OTEL_ENABLED=true"," makes LangChain and LangGraph emit spans through the LangSmith exporter, and ","LANGSMITH_OTEL_ONLY=true"," stops it sending to LangSmith's own format as well. With any other application, point a standard OTLP exporter at the base endpoint ","https:\u002F\u002Fapi.smith.langchain.com\u002Fotel","; regional endpoints exist for EU, APAC and AWS US. The exporter appends the signal path itself, so putting \u002Fv1\u002Ftraces in the base URL gives a 404.",{"type":168,"code":282},"pip install \"langsmith[otel]\"          # needs langsmith >= 0.3.18, 0.4.25 recommended\n\nexport LANGSMITH_TRACING=true\nexport LANGSMITH_OTEL_ENABLED=true\nexport LANGSMITH_ENDPOINT=https:\u002F\u002Fapi.smith.langchain.com\nexport LANGSMITH_API_KEY=...\n\n# fan out one OTLP stream to LangSmith and to the rest of the stack\nexport OTEL_EXPORTER_OTLP_ENDPOINT=https:\u002F\u002Fapi.smith.langchain.com\u002Fotel\nexport OTEL_EXPORTER_OTLP_HEADERS=\"x-api-key=...,Langsmith-Project=support\"\n\n# OTel-only, for teams that do not want a second transport\nexport LANGSMITH_OTEL_ONLY=true\n",{"type":115,"content":284},[285],"The trade is between convenience and overhead. The OTel path costs an attribute-mapping exercise, because span attributes have to be labelled with the langsmith namespace to become run types, run IDs and dotted order; LangChain's own announcement describes the OpenTelemetry route as having slightly higher overhead and recommends the native format when LangSmith is the only destination. The native format also gives pending runs that appear in the UI while the work is still going.",{"type":188,"variant":189,"title":287,"body":288},"A partial fan-out loses spans without an error",[289],[290,291,294],"The OTLP endpoint is asynchronous: it accepts a batch, answers 200 and materialises the runs in the background. A child span whose parent never arrives is buffered and then expires, and because the 200 was already sent, nothing is reported. If part of a distributed trace goes to another backend, the part that reaches LangSmith can quietly lose its parent; on self-hosted installations the buffering window is the Redis setting ",{"tag":168,"children":292},[293],"REDIS_RUNS_EXPIRY_SECONDS",", 12 hours by default.",{"type":122,"level":123,"id":105,"text":106},{"type":115,"content":297},[298],"The weaknesses first, because they decide whether you need this product. The billing unit is the trace, which quietly rewards sampling and punishes an application that traces every request; the load-balancer caps are per service key or personal access token, not per organisation, so a horizontally scaled fleet needs either more keys or the SDK's batching. Workspace role-based access control is Enterprise only, so a growing team on Plus shares one role model. And the resource hierarchy is being rewritten underneath you: workspaces were tenants, agents are in beta as a grouping above projects, and an agent-based workspace addresses traces by agent and environment rather than by project.",{"type":203,"head":300,"rows":309},[301,303,305,307],[302],"Tool",[304],"Primary strength",[306],"Hosting",[308],"What it gives up",[310,318,327,336],[311,312,314,316],[24],[313],"Tracing plus evals plus a managed agent runtime",[315],"Cloud, or self-hosted on Enterprise",[317],"Open core is not an option: the SDK is free, the platform is not",[319,321,323,325],[320],"Langfuse",[322],"MIT core you can self-host and inspect",[324],"Cloud or self-host",[326],"Less of the managed runtime around the traces",[328,330,332,334],[329],"Arize Phoenix",[331],"Apache-2.0, built on OpenTelemetry",[333],"Self-host or cloud",[335],"A narrower product surface around datasets and evals",[337,339,341,343],[338],"Helicone",[340],"Cheap request-level logging with fast setup",[342],"Cloud or proxy deployment",[344],"Less depth on run trees and evaluation workflows",{"type":115,"content":346},[347],"What LangSmith does not give up is measurement: it is the reference implementation for run trees, thread-level conversation views and dataset-driven regression testing, and the annotation queue with reservations is a genuine answer to the question of who labels what. If your application already runs on LangGraph, choosing it is the obvious path and the switching cost is the runtime, not the traces.",{"type":122,"level":123,"id":108,"text":109},{"type":115,"content":350},[351],"LangSmith is worth paying for when an agent's behaviour, not its uptime, is the thing that breaks, and when you need humans in the loop labelling runs and datasets that drive a release gate. It is worth less than its price when you only want a request log, when one trace per user request is too many traces for the budget, or when the platform sprawl makes it unclear what you are paying for.",{"type":128,"ordered":353,"items":354},true,[355,357,359,361,363],[356],"Adopt it if you need dataset-driven regression tests that run in CI, not just a trace viewer.",[358],"Adopt it if several people have to label runs, because annotation queues with reservations and dataset export are hard to rebuild.",[360],"Adopt it if you already build on LangGraph, and factor the runtime into the decision rather than pretending the traces are separate.",[362],"Be careful if you plan to trace every request: model the event and byte ceilings per hour before you turn instrumentation on in production.",[364],"Look elsewhere if you need self-hosting without a sales conversation, or workspace-level access control, or if a flat price per trace fits your workload better than an LSU-metered platform.",{"type":188,"variant":366,"title":367,"body":368},"note","The short version",[369],[370],"The best trace and evaluation tooling available, sold as a platform that now includes the runtime it observes. Take the tracing and evaluation parts seriously, keep the native transport unless you already run OpenTelemetry, and set usage limits before the first production week rather than after the invoice.",{"type":122,"level":123,"id":111,"text":112},{"type":128,"ordered":353,"items":373},[374,378,381,384,387,390,393],[375],{"tag":376,"href":34,"children":377},"a",[33],[379],{"tag":376,"href":37,"children":380},[36],[382],{"tag":376,"href":40,"children":383},[39],[385],{"tag":376,"href":43,"children":386},[42],[388],{"tag":376,"href":46,"children":389},[45],[391],{"tag":376,"href":49,"children":392},[48],[394],{"tag":376,"href":52,"children":395},[396],"LangChain: end-to-end OpenTelemetry support in LangSmith",[398,447,530,584],{"slug":399,"published":400,"minutes":401,"category":7,"tags":402,"keywords":408,"about":415,"sources":419,"cover":438,"og":439,"expertise":55,"locales":440,"lang":57,"title":441,"description":442,"coverAlt":443,"url":444,"pricing":445,"kind":446},"ollama","2026-09-29",11,[403,404,405,406,407],"Local inference","Open models","llama.cpp","GGUF","Model serving",[399,409,410,411,412,413,414],"ollama vs lm studio","ollama vs vllm","local llm runtime","gguf model server","ollama self hosting","ollama api",[416],{"name":417,"url":418},"Ollama (software)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FOllama",[420,423,426,429,432,435],{"title":421,"url":422},"Ollama API documentation","https:\u002F\u002Fdocs.ollama.com\u002Fapi",{"title":424,"url":425},"Ollama on GitHub, with the MIT LICENSE file","https:\u002F\u002Fgithub.com\u002Follama\u002Follama",{"title":427,"url":428},"Ollama terms of service, last updated May 2026","https:\u002F\u002Follama.com\u002Fterms",{"title":430,"url":431},"Ollama pricing, cloud plans and per-token model rates","https:\u002F\u002Follama.com\u002Fpricing",{"title":433,"url":434},"Hardware support: Nvidia, AMD, Metal and Vulkan","https:\u002F\u002Fdocs.ollama.com\u002Fgpu",{"title":436,"url":437},"OpenAI compatibility, including what is not supported","https:\u002F\u002Fdocs.ollama.com\u002Fapi\u002Fopenai-compatibility","\u002Fimages\u002Fblog\u002Follama\u002Fcover.webp","\u002Fimages\u002Fblog\u002Follama\u002Fog.jpg",[57,58,59],"Ollama review: the friendly way to run open models","Ollama serves open models over one HTTP API on your own hardware. What it does well, where throughput falls short, and what the MIT licence does not cover.","Abstract cover art for the Ollama review","https:\u002F\u002Follama.com","MIT · free for personal use","Local inference runtime",{"slug":448,"published":449,"minutes":401,"category":7,"tags":450,"keywords":456,"about":464,"sources":471,"cover":523,"og":524,"expertise":55,"locales":525,"lang":57,"title":526,"description":527,"coverAlt":528,"url":467,"pricing":529,"kind":451},"portkey","2026-09-28",[451,452,453,454,455],"LLM gateway","Guardrails","Routing","Observability","Cost control",[457,458,459,460,461,462,463],"portkey ai gateway","portkey vs litellm","llm gateway comparison","llm gateway latency overhead","llm guardrails gateway","self-hosted llm gateway","portkey pricing",[465,468],{"name":466,"url":467},"Portkey","https:\u002F\u002Fportkey.ai",{"name":469,"url":470},"API gateway","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FAPI_gateway",[472,475,478,481,484,487,490,493,496,499,502,505,508,511,514,517,520],{"title":473,"url":474},"Portkey docs: AI Gateway","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fproduct\u002Fai-gateway",{"title":476,"url":477},"Portkey docs: Getting started with the AI Gateway","https:\u002F\u002Fdocs.portkey.ai\u002Fdocs\u002Fguides\u002Fgetting-started\u002Fgetting-started-with-ai-gateway",{"title":479,"url":480},"Portkey docs: Gateway config object","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fapi-reference\u002Fconfig-object",{"title":482,"url":483},"Portkey docs: Guardrails","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fproduct\u002Fguardrails",{"title":485,"url":486},"Portkey docs: Guardrail endpoints and capabilities","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fproduct\u002Fguardrails\u002Fcapabilities",{"title":488,"url":489},"Portkey docs: Cache, simple and semantic","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fproduct\u002Fai-gateway\u002Fcache-simple-and-semantic",{"title":491,"url":492},"Portkey docs: Load balancing","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fproduct\u002Fai-gateway\u002Fload-balancing",{"title":494,"url":495},"Portkey docs: Enterprise hybrid deployment architecture","https:\u002F\u002Fportkey.ai\u002Fdocs\u002Fself-hosting\u002Fhybrid-deployments\u002Farchitecture",{"title":497,"url":498},"Portkey pricing","https:\u002F\u002Fportkey.ai\u002Fpricing",{"title":500,"url":501},"Portkey gateway on GitHub, MIT licensed","https:\u002F\u002Fgithub.com\u002FPortkey-AI\u002Fgateway",{"title":503,"url":504},"Portkey's own benchmark: gateway versus direct Bedrock","https:\u002F\u002Fgithub.com\u002FPortkey-AI\u002Fbenchmark-test",{"title":506,"url":507},"Portkey status page","https:\u002F\u002Fstatus.portkey.ai\u002F",{"title":509,"url":510},"Palo Alto Networks completes acquisition of Portkey, May 2026","https:\u002F\u002Fwww.paloaltonetworks.com\u002Fcompany\u002Fpress\u002F2026\u002Fpalo-alto-networks-completes-acquisition-of-portkey-to-secure-ai-agents",{"title":512,"url":513},"Palo Alto Networks: Prisma AIRS AI Gateway","https:\u002F\u002Fwww.paloaltonetworks.com\u002Fai-security\u002Fai-gateway",{"title":515,"url":516},"Cloudflare AI Gateway pricing","https:\u002F\u002Fdevelopers.cloudflare.com\u002Fai-gateway\u002Freference\u002Fpricing\u002F",{"title":518,"url":519},"LiteLLM pricing","https:\u002F\u002Fwww.litellm.ai\u002Fpricing",{"title":521,"url":522},"OpenRouter pricing","https:\u002F\u002Fopenrouter.ai\u002Fpricing","\u002Fimages\u002Fblog\u002Fportkey\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fportkey\u002Fog.jpg",[57,58,59],"Portkey: a production LLM gateway, reviewed for routing, guardrails and cost","Portkey puts retries, fallbacks, caching, guardrails and cost tracking behind one OpenAI-compatible endpoint. What the config object does well, what the gateway costs in latency, and when to self-host.","A request path from an application through the Portkey gateway to three model providers, with the guardrail verdict and the log written below the proxy.","Free · from $49 per month",{"slug":531,"published":532,"minutes":6,"category":7,"tags":533,"keywords":536,"about":544,"sources":552,"cover":577,"og":578,"expertise":55,"locales":579,"lang":57,"title":580,"description":581,"coverAlt":582,"url":546,"pricing":583,"kind":64},"langfuse","2026-08-13",[64,9,11,534,535],"Self-hosting","Evaluation",[531,537,538,539,540,541,542,543],"langfuse vs langsmith","llm tracing tool","self-hosted llm observability","langfuse pricing","opentelemetry llm traces","llm cost tracking","prompt versioning",[545,547,548,549],{"name":320,"url":546},"https:\u002F\u002Flangfuse.com",{"name":11,"url":27},{"name":29,"url":30},{"name":550,"url":551},"Observability (software)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FObservability_(software)",[553,556,559,562,565,568,571,574],{"title":554,"url":555},"Langfuse documentation: observability and application tracing","https:\u002F\u002Flangfuse.com\u002Fdocs\u002Fobservability\u002Foverview",{"title":557,"url":558},"Langfuse documentation: get started with tracing","https:\u002F\u002Flangfuse.com\u002Fdocs\u002Fobservability\u002Fget-started",{"title":560,"url":561},"Langfuse pricing: cloud plans, billable units and worked examples","https:\u002F\u002Flangfuse.com\u002Fpricing",{"title":563,"url":564},"Langfuse pricing: self-hosted plans and the feature comparison","https:\u002F\u002Flangfuse.com\u002Fpricing-self-host",{"title":566,"url":567},"Self-host Langfuse: deployment options, containers and storage services","https:\u002F\u002Flangfuse.com\u002Fself-hosting",{"title":569,"url":570},"Langfuse changelog: v4 is live (17 August 2026)","https:\u002F\u002Flangfuse.com\u002Fchangelog\u002F2026-08-17-langfuse-v4",{"title":572,"url":573},"Langfuse blog: Langfuse joins ClickHouse (16 January 2026)","https:\u002F\u002Flangfuse.com\u002Fblog\u002Fjoining-clickhouse",{"title":575,"url":576},"GitHub: langfuse\u002Flangfuse, the platform repository","https:\u002F\u002Fgithub.com\u002Flangfuse\u002Flangfuse","\u002Fimages\u002Fblog\u002Flangfuse\u002Fcover.webp","\u002Fimages\u002Fblog\u002Flangfuse\u002Fog.jpg",[57,58,59],"Langfuse review: tracing, prompts and evals you can host yourself","Langfuse puts LLM traces, prompt versions and experiments on one MIT-licensed platform. What self-hosting really costs, how the unit pricing adds up, and where it loses.","A pipeline from a batched application event through the Langfuse web container and object storage into ClickHouse, with Redis and PostgreSQL alongside.","MIT · paid from $59 per month",{"slug":585,"published":586,"minutes":6,"category":7,"tags":587,"keywords":592,"about":599,"sources":603,"cover":619,"og":620,"expertise":55,"locales":621,"lang":57,"title":622,"description":623,"coverAlt":624,"url":625,"pricing":626,"kind":451},"openrouter","2026-07-23",[451,588,589,590,591],"Model routing","Fallbacks","OpenAI-compatible","Pay per token",[585,593,459,594,595,596,597,598],"openrouter vs litellm","openrouter pricing","openai compatible api gateway","llm fallback routing","multi model api gateway","byok llm routing",[600],{"name":601,"url":602},"OpenRouter","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FOpenRouter",[604,607,608,611,614,617],{"title":605,"url":606},"OpenRouter documentation: quickstart","https:\u002F\u002Fopenrouter.ai\u002Fdocs\u002Fquickstart",{"title":521,"url":522},{"title":609,"url":610},"OpenRouter documentation: model fallbacks","https:\u002F\u002Fopenrouter.ai\u002Fdocs\u002Fguides\u002Frouting\u002Fmodel-fallbacks",{"title":612,"url":613},"OpenRouter documentation: provider routing","https:\u002F\u002Fopenrouter.ai\u002Fdocs\u002Fguides\u002Frouting\u002Fprovider-selection",{"title":615,"url":616},"OpenRouter documentation index","https:\u002F\u002Fopenrouter.ai\u002Fdocs\u002Fllms.txt",{"title":618,"url":602},"Wikipedia: OpenRouter","\u002Fimages\u002Fblog\u002Fopenrouter\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fopenrouter\u002Fog.jpg",[57,58,59],"OpenRouter: one API key in front of every model you might call","OpenRouter puts 500+ models from 80+ providers behind one OpenAI-compatible endpoint, with fallbacks and pass-through pricing. What it costs, where it breaks.","Request path through OpenRouter: client, router, candidate providers, fallback list and the model that finally answers.","https:\u002F\u002Fopenrouter.ai","Pay per token, no subscription",1791383549041]