[{"data":1,"prerenderedAt":585},["ShallowReactive",2],{"blog-local-text-to-speech-pipeline-en":3},{"slug":4,"published":5,"minutes":6,"category":7,"tags":8,"keywords":14,"about":25,"sources":35,"cover":57,"og":58,"expertise":59,"locales":60,"lang":61,"title":64,"description":65,"coverAlt":66,"metaTitle":67,"takeaways":68,"faq":74,"toc":90,"blocks":118,"others":359},"local-text-to-speech-pipeline","2026-10-01",11,"llmops",[9,10,11,12,13],"Text-to-speech","Audio","Kokoro","FFmpeg","Vue",[15,16,17,18,19,20,21,22,23,24],"local text-to-speech","text-to-speech pipeline","kokoro tts","piper tts","bark tts comparison","narrate blog posts","wav to m4a ffmpeg","aac 192 kbit\u002Fs faststart","vue audio player","onnx tts mac",[26,29,32],{"name":27,"url":28},"Speech synthesis","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FSpeech_synthesis",{"name":30,"url":31},"Advanced Audio Coding","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FAdvanced_Audio_Coding",{"name":33,"url":34},"ESpeak","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FESpeak",[36,39,42,45,48,51,54],{"title":37,"url":38},"Kokoro onnx: runtime bindings","https:\u002F\u002Fgithub.com\u002Fthewh1teagle\u002Fkokoro-onnx",{"title":40,"url":41},"Kokoro-82M model card","https:\u002F\u002Fhuggingface.co\u002Fhexgrad\u002FKokoro-82M",{"title":43,"url":44},"Piper text-to-speech","https:\u002F\u002Fgithub.com\u002Frhasspy\u002Fpiper",{"title":46,"url":47},"Bark by Suno","https:\u002F\u002Fgithub.com\u002Fsuno-ai\u002Fbark",{"title":49,"url":50},"FFmpeg AAC encoder documentation","https:\u002F\u002Fffmpeg.org\u002Fffmpeg-codecs.html#aac",{"title":52,"url":53},"torch.load weights_only documentation","https:\u002F\u002Fpytorch.org\u002Fdocs\u002Fstable\u002Fgenerated\u002Ftorch.load.html",{"title":55,"url":56},"ESpeak NG phonemizer","https:\u002F\u002Fgithub.com\u002Fespeak-ng\u002Fespeak-ng","\u002Fimages\u002Fblog\u002Flocal-text-to-speech-pipeline\u002Fcover.webp","\u002Fimages\u002Fblog\u002Flocal-text-to-speech-pipeline\u002Fog.jpg","ai-engineer",[61,62,63],"en","de","hu","Local text-to-speech at scale: narrating 95 articles with open models","Three open TTS models on one laptop: how 95 articles became 153 minutes of narration, from SQLite rows through chunked synthesis and AAC at 192 kbit\u002Fs to a player that follows you down the page.","Wave diagram of the narration pipeline: database rows are split into sentences, synthesized by a local model, encoded to AAC and played from a sticky tab at the screen edge.","Local TTS at scale: 95 articles, 3 models · Balázs Csorba",[69,70,71,72,73],"Local TTS inverts the economics of narration: 95 articles and 153 minutes of audio cost about 35 minutes of generation and zero API fees — the marginal cost of one more article is electricity.","Model choice at this scale is a speed-quality Pareto, not a benchmark fight: Piper was fastest, Bark never finished, and Kokoro v1.0 (82M, ONNX) won on prosody at about four times real time.","Chunk at sentence boundaries up to 400 characters with 80 ms of silence between pieces; smarter break points buy almost nothing a listener would notice.","WAV is a working format, not a shipping format: AAC at 192 kbit\u002Fs with faststart cut 406 MB to 181 MB and makes duration and seeking work before the file is fully downloaded.","In a prerendered player, media events beat assumptions: loadedmetadata can fire before hydration, so duration must be re-read on durationchange, canplay and mount — subscribe to state, not notifications.",[75,78,81,84,87],{"q":76,"a":77},"Why not use a hosted text-to-speech API?","Because at 95 articles the metered cost starts to matter with every edit and re-render, and because reproducibility matters more: the same text and model version produce the same file next year. Hosted voices still win on peak expressiveness; they do not win on the economics of a whole archive.",{"q":79,"a":80},"Which model sounds the best?","Kokoro v1.0, and it is what ships: 82 million parameters, natural emphasis, clean handling of numbers and product names. Piper with the LibriTTS voice is a close second and noticeably faster. Bark is the most expressive of the three but was too slow to be usable at this volume.",{"q":82,"a":83},"How long does the full run take?","About 35 minutes for all 95 articles on one Apple Silicon laptop — roughly 4.5 times real time — with peak memory under 2 GB. A single article regenerates in about 20 seconds.",{"q":85,"a":86},"Why AAC (M4A) instead of MP3 or Opus?","At 192 kbit\u002Fs the codec differences are inaudible for speech, so the decision is about packaging: AAC plays natively everywhere including Safari, faststart exposes the duration before the download completes, and Opus would be smaller but has uneven browser support outside WebM.",{"q":88,"a":89},"Does the narration exist in German and Hungarian?","Not yet. The narration reads the English original of each article; the player interface itself is fully localized in English, German and Hungarian.",[91,94,97,100,103,106,109,112,115],{"id":92,"title":93},"why-local","Why local text-to-speech",{"id":95,"title":96},"model-shootout","Three models, one laptop",{"id":98,"title":99},"the-pipeline","The generation pipeline",{"id":101,"title":102},"wav-to-aac","WAV is a working format, not a shipping format",{"id":104,"title":105},"the-player","The player",{"id":107,"title":108},"what-broke","What broke on the way",{"id":110,"title":111},"the-numbers","The numbers",{"id":113,"title":114},"verdict","What I would do differently",{"id":116,"title":117},"sources","Sources",[119,123,126,129,132,135,142,143,146,149,152,155,200,201,204,207,211,214,221,227,228,231,234,236,239,240,243,246,249,255,256,280,281,323,326,327,330,333,334],{"type":120,"content":121},"paragraph",[122],"Every article on this site now has a narration: 95 posts, 153 minutes of audio, generated on one laptop without a single API call. This post is the full account — how three open text-to-speech models were compared, how the pipeline turns SQLite rows into sentence-sized synthesis chunks, why the shipped format is AAC at 192 kbit\u002Fs, and how the player is built so the audio follows you down the page.",{"type":120,"content":124},[125],"The numbers first, because they frame every decision below: 95 articles (45 blog posts and 50 tool reviews), 153 minutes of finished narration, about 35 minutes of generation time, 406 MB of intermediate WAV reduced to 181 MB of shipped AAC, and zero euros in API fees. Everything ran on a single Apple Silicon laptop with 16 GB of memory.",{"type":127,"level":128,"id":92,"text":93},"heading",2,{"type":120,"content":130},[131],"Hosted text-to-speech is excellent and improving every quarter — but it is metered. At this volume the bill scales with every word, every re-render after an edit, every experiment with voice or speed. A local pipeline inverts the economics: the marginal cost of one more article is a few cents of electricity and two minutes of waiting. It is also reproducible — the same text, the same model version and the same voice produce the same file next year, which matters when an edited article must be re-narrated without the voice drifting.",{"type":120,"content":133},[134],"There is a second reason: the text never leaves the machine. The draft of an unpublished article is read by the same process that will publish it, on the same machine, and when a paragraph changes, rerunning one slug regenerates one file — not ninety-five.",{"type":136,"variant":137,"title":138,"body":139},"callout","note","The constraint",[140],[141],"No GPU cluster and no cloud batch job: one 16 GB Apple Silicon laptop, shared with everything else. That bounded the choice to models that fit comfortably in memory and run near real time on CPU — a model of a few hundred million parameters, not a few billion.",{"type":127,"level":128,"id":95,"text":96},{"type":120,"content":144},[145],"Piper was the starting point because it is boring in the best way: a small ONNX model (the LibriTTS high voice is about 137 MB), espeak-ng phonemization, 22.05 kHz output, generation at roughly five times real time. The result is clean and consistent — a competent newsreader — but it is also even. Long articles sound slightly flat.",{"type":120,"content":147},[148],"The first obstacle was packaging rather than quality: the prebuilt macOS binary ships x86_64 only, and on Apple Silicon it refuses to link against the arm64 espeak-ng library with an architecture mismatch. The Python package sidesteps the binary entirely and was running in minutes. A useful reminder that \"it does not run\" is often a toolchain problem, not a model problem.",{"type":120,"content":150},[151],"Bark was the most tempting of the three: expressive, capable of laughs and sighs, with published samples that make any other model sound like a metronome. It was also the one that never produced a usable sentence. The current PyTorch changed the default of torch.load to weights_only, so checkpoint unpickling failed until patched; once loading worked, inference on CPU was slow enough that 95 articles would have taken most of a day. Expressiveness was not worth that bill — Bark stays a research toy for this workload.",{"type":120,"content":153},[154],"Kokoro v1.0 through the kokoro-onnx bindings was the find: 82 million parameters, about 325 MB of ONNX, 24 kHz output, 54 voices, and generation at about four times real time with noticeably better prosody than Piper — sentences land their emphasis, and numbers and abbreviations are handled without the halting quality typical of smaller models. It is what every article on this site sounds like now.",{"type":156,"head":157,"rows":168},"table",[158,160,162,164,166],[159],"Model",[161],"Weights",[163],"Output",[165],"Speed on this machine",[167],"Verdict",[169,180,190],[170,172,174,176,178],[171],"Piper, LibriTTS high",[173],"about 137 MB, ONNX",[175],"22.05 kHz WAV",[177],"about 5x real time",[179],"Fast and consistent, slightly flat",[181,182,184,186,188],[46],[183],"about 2 GB, PyTorch",[185],"24 kHz WAV",[187],"never finished in budget",[189],"Expressive, impractical here",[191,193,195,196,198],[192],"Kokoro v1.0",[194],"82M, about 325 MB, ONNX",[185],[197],"about 4x real time",[199],"Natural prosody; the shipped voice",{"type":127,"level":128,"id":98,"text":99},{"type":120,"content":202},[203],"The content database is the single source of truth, so the pipeline reads from it directly: title, description, key takeaways and FAQ on one side, the article block list on the other. Nothing is scraped from the rendered page. If a paragraph exists in the post, it is narrated — and if it is a table, a callout or a code block, it is flattened into something a listener can follow: table rows become comma-separated lines, callout titles are spoken like headings, code is read as written.",{"type":120,"content":205},[206],"Two constraints shape the middle of the pipeline. Models truncate long input, so articles must be split; and prosody must not be cut in half, so splits must land on sentence boundaries. The chunker is deliberately plain: accumulate sentences up to 400 characters, flush at the boundary, join the pieces with 80 milliseconds of silence. Smarter break points — paragraph ends, headings as intonation resets — bought almost nothing; the sentence is the unit listeners actually notice.",{"type":127,"level":208,"id":209,"text":210},3,"the-chunker","The chunker",{"type":212,"code":213},"code","def chunks(text, limit=400):\n    \"\"\"Accumulate sentences up to limit characters; never split mid-sentence.\"\"\"\n    out, buf = [], \"\"\n    for sentence in text.split(\". \"):\n        cand = (buf + \" \" + sentence).strip()\n        if len(cand) > limit and buf:\n            out.append(buf + \".\")\n            buf = sentence\n        else:\n            buf = cand\n    if buf:\n        out.append(buf)\n    return out\n\npieces = []\nfor part in chunks(article_text):\n    audio, _ = model.create(part, voice=\"af_bella\")\n    pieces.append(audio)\n    pieces.append(np.zeros(int(24000 * 0.08)))   # 80 ms between chunks\nsf.write(tmp_wav, np.concatenate(pieces), 24000)",{"type":120,"content":215},[216,220],{"tag":217,"children":218},"strong",[219],"One voice reads all 95 articles."," It is af_bella, chosen after rendering the same paragraph in several of the 54 voices. Consistency is the point: the archive should sound like one publication, not a lottery. The narration reads the English text of each article — including from the German and Hungarian pages — and is labelled as the English narration of the piece, not as a translation of it.",{"type":136,"variant":222,"title":223,"body":224},"tip","Rerun only what changed",[225],[226],"The generator skips every slug whose output file already exists, so an edit to one article costs one regeneration, not a full run. A complete fresh pass over all 95 articles is about 35 minutes — cheap enough that a full rebuild is also a reasonable default.",{"type":127,"level":128,"id":101,"text":102},{"type":120,"content":229},[230],"The synthesis step writes 24 kHz, 16-bit PCM: correct, seekable, uncompressed — and about 406 MB for 153 minutes of audio. On disk that is harmless; over the wire it is a mistake, and on a site that prerenders everything it would be the single largest class of asset by an order of magnitude.",{"type":120,"content":232},[233],"The shipped format is AAC in an M4A container at 192 kbit\u002Fs, mono. That bitrate is generous for speech — 96 to 128 kbit\u002Fs already sounds transparent at 24 kHz — but 192 leaves headroom and costs about a megabyte per article. The container matters as much as the codec: faststart moves the moov atom to the front of the file, so a browser knows the duration and can seek before the download finishes. Without it, players sit at 0:00 and refuse to scrub until the last byte arrives.",{"type":212,"code":235},"ffmpeg -i article.wav -c:a aac -b:a 192k -ac 1 -movflags +faststart article.m4a",{"type":120,"content":237},[238],"406 MB became 181 MB — a 55 percent reduction, every file under 1.5 MB, delivered with ordinary range requests and cache headers. The WAV files never reach the build output: they are deleted the moment the encoder succeeds.",{"type":127,"level":128,"id":104,"text":105},{"type":120,"content":241},[242],"Audio that arrives late is audio nobody hears, so the player loads metadata only: the header is fetched, the duration is learned, and no actual samples are downloaded until the reader presses play. Nothing autoplays — a page that talks at you is a page people close.",{"type":120,"content":244},[245],"Each article gets an inline bar under the table of contents: play and pause, a seekable progress line with current and total time, and a speed toggle cycling 1x, 1.25x, 1.5x and 2x. The seek control is a native range input styled to the site, so keyboard navigation and screen-reader semantics are inherited rather than reimplemented.",{"type":120,"content":247},[248],"The inline bar is useless once you scroll past it, so it hands over to a sticky tab pinned to the right edge of the viewport. An IntersectionObserver watches the bar: the moment it leaves the screen, the tab appears; scroll back up, and it steps aside. The tab fills from the bottom as playback advances — progress is legible at a glance — and it honours prefers-reduced-motion like the rest of the interface.",{"type":136,"variant":250,"title":251,"body":252},"warn","The 0:00 bug",[253],[254],"The first deployed version showed 0:00 for the total duration on every page. The audio element fires loadedmetadata during page load, and on a fast connection with a faststart file that happens before the JavaScript has hydrated and attached its handlers — so the one event carrying the duration was missed. The fix is to stop trusting a single event: duration is now re-read on durationchange and canplay, and checked once on mount. In a player, media events are a stream, not a message — subscribe to the state, not to the notification.",{"type":127,"level":128,"id":107,"text":108},{"type":257,"ordered":258,"items":259},"list",false,[260,265,270,275],[261,264],{"tag":217,"children":262},[263],"Packaging beat models twice."," Piper shipped an x86_64-only binary for an arm64 machine, and Bark’s checkpoint loading failed after PyTorch flipped the weights_only default. Neither failure had anything to do with speech.",[266,269],{"tag":217,"children":267},[268],"Version drift is real."," The kokoro-onnx package passed the speed parameter as int32 while the shipped model expects float — a one-line patch, found by reading a two-line stack trace instead of guessing.",[271,274],{"tag":217,"children":272},[273],"Verify the audio, not just the file."," ffprobe on every output caught duration problems before they reached a browser; a file that exists is not a file that plays.",[276,279],{"tag":217,"children":277},[278],"Old formats die hard."," The first pass left WAVs in the build output, caught only because a request for one returned the wrong byte count.",{"type":127,"level":128,"id":110,"text":111},{"type":156,"head":282,"rows":287},[283,285],[284],"Metric",[286],"Value",[288,293,298,303,308,313,318],[289,291],[290],"Articles narrated",[292],"95 (45 blog posts, 50 tool reviews)",[294,296],[295],"Finished audio",[297],"153 minutes",[299,301],[300],"Generation wall time",[302],"about 35 minutes, 4.5x real time",[304,306],[305],"Peak memory",[307],"under 2 GB",[309,311],[310],"Intermediate WAV",[312],"406 MB",[314,316],[315],"Shipped AAC",[317],"181 MB, about 1 MB per article",[319,321],[320],"API cost",[322],"0 EUR",{"type":120,"content":324},[325],"Read together, the numbers say that the cost of this pipeline is patience, not money. A full re-run — after a voice change or an edit that touches every article — is well under an hour on a machine that stays usable throughout. That is the argument for local text-to-speech at this scale: not that it beats a frontier hosted voice on expressiveness, but that it makes narration a default rather than a budget line.",{"type":127,"level":128,"id":113,"text":114},{"type":120,"content":328},[329],"Start with Kokoro. The shootout was not wasted — comparing models is what makes the choice defensible — but the shipped pipeline would have been identical with Kokoro as the only candidate. Second: encode straight from synthesis to AAC and never persist WAVs; the intermediate format added a cleanup step and 406 MB of files that never needed to exist. Third: treat the player’s media events as state to subscribe to, not notifications to catch — and the duration bug never happens.",{"type":120,"content":331},[332],"What remains is scope. The narration is English-only for now — one voice, one language, 95 articles — and per-locale voices are the obvious next step if the German and Hungarian audiences ask for them. Until then the English narration doubles as the pronunciation guide for every product name on the site, which is its own quiet utility.",{"type":127,"level":128,"id":116,"text":117},{"type":257,"ordered":335,"items":336},true,[337,341,344,347,350,353,356],[338],{"tag":339,"href":38,"children":340},"a",[37],[342],{"tag":339,"href":41,"children":343},[40],[345],{"tag":339,"href":44,"children":346},[43],[348],{"tag":339,"href":47,"children":349},[46],[351],{"tag":339,"href":50,"children":352},[49],[354],{"tag":339,"href":53,"children":355},[52],[357],{"tag":339,"href":56,"children":358},[55],[360,416,504,540],{"slug":361,"published":362,"updated":363,"minutes":364,"category":7,"tags":365,"keywords":370,"about":381,"sources":391,"cover":410,"og":411,"expertise":59,"locales":412,"lang":61,"title":413,"description":414,"coverAlt":415},"artificial-analysis-leaderboard-claude-opus-5-5","2026-09-07","2026-09-11",8,[366,367,368,369],"Claude Opus 5.5","Artificial Analysis","LLM benchmarks","LLM cost",[371,372,366,373,374,375,376,377,378,379,380],"Claude Opus 5.5 benchmark","Claude Opus 5.5 pricing","Artificial Analysis Intelligence Index","AI model leaderboard","Opus 5.5 benchmark","Opus 5.5 vs GPT-6 Astra","LLM cost per task","reasoning effort setting","Opus 5.5 pricing","best LLM September 2026",[382,385,388],{"name":383,"url":384},"Claude (language model)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FClaude_(language_model)",{"name":386,"url":387},"Large language model","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FLarge_language_model",{"name":389,"url":390},"Benchmark (computing)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FBenchmark_(computing)",[392,395,398,401,404,407],{"title":393,"url":394},"Artificial Analysis: Claude Opus 5.5 takes the top spot on the Artificial Analysis Intelligence Index (22 September 2026)","https:\u002F\u002Fartificialanalysis.ai\u002Farticles\u002Fclaude-opus-5-5",{"title":396,"url":397},"Artificial Analysis: Claude Opus 5.5 (max) model page","https:\u002F\u002Fartificialanalysis.ai\u002Fmodels\u002Fclaude-opus-5-5",{"title":399,"url":400},"Artificial Analysis: Claude Opus 5.5 (medium) model page","https:\u002F\u002Fartificialanalysis.ai\u002Fmodels\u002Fclaude-opus-5-5-medium",{"title":402,"url":403},"Artificial Analysis: Claude Opus 5 (max) model page","https:\u002F\u002Fartificialanalysis.ai\u002Fmodels\u002Fclaude-opus-5",{"title":405,"url":406},"OfficeChai: Claude Opus 5.5 creates a 5-point lead over GPT-6 Astra","https:\u002F\u002Fofficechai.com\u002Fai\u002Fclaude-opus-5-5-creates-5-point-lead-over-gpt-6-astra-jumps-to-top-spot-on-artificial-analysis-intelligence-index\u002F",{"title":408,"url":409},"Claude API docs: Models overview, context windows and prices (as of September 2026)","https:\u002F\u002Fplatform.claude.com\u002Fdocs\u002Fen\u002Fabout-claude\u002Fmodels\u002Foverview","\u002Fimages\u002Fblog\u002Fartificial-analysis-leaderboard-claude-opus-5-5\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fartificial-analysis-leaderboard-claude-opus-5-5\u002Fog.jpg",[61,62,63],"Claude Opus 5.5 takes #1 on Artificial Analysis, and medium effort is the real story","Claude Opus 5.5 is #1 of 211 models on Artificial Analysis with 58 points. At medium effort it matches Opus 5 for $1.34 per task instead of $5.86.","Horizontal bars of Intelligence Index scores: Claude Opus 5.5 at max effort 58, GPT-6 Astra and Claude Fable 5.1 53, Opus 5 51, Opus 5.5 at medium effort 51.",{"slug":417,"published":418,"minutes":419,"category":7,"tags":420,"keywords":426,"about":437,"sources":446,"cover":498,"og":499,"expertise":59,"locales":500,"lang":61,"title":501,"description":502,"coverAlt":503},"agent-observability-opentelemetry","2026-08-06",12,[421,422,423,424,425],"OpenTelemetry","LLM observability","AI agents","Tracing","Evals",[427,428,429,430,431,432,433,434,435,436],"LLM agent observability OpenTelemetry","OpenTelemetry GenAI semantic conventions","how to trace LLM agents","gen_ai semantic conventions attributes","LLM token usage and cost metrics","LLM tracing sampling","PII in LLM traces","Langfuse vs Arize Phoenix vs Datadog","AI agent tracing evals","OpenTelemetry LLM tracing",[438,440,443],{"name":421,"url":439},"https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FOpenTelemetry",{"name":441,"url":442},"Observability","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FObservability_(software)",{"name":444,"url":445},"Intelligent agent","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FIntelligent_agent",[447,450,453,456,459,462,465,468,471,474,477,480,483,486,489,492,495],{"title":448,"url":449},"OpenTelemetry: GenAI semantic conventions repository (open-telemetry\u002Fsemantic-conventions-genai)","https:\u002F\u002Fgithub.com\u002Fopen-telemetry\u002Fsemantic-conventions-genai",{"title":451,"url":452},"GenAI conventions: overview (status Development)","https:\u002F\u002Fgithub.com\u002Fopen-telemetry\u002Fsemantic-conventions-genai\u002Fblob\u002Fmain\u002Fdocs\u002Fgen-ai\u002FREADME.md",{"title":454,"url":455},"GenAI conventions: model spans, execute_tool and content capture","https:\u002F\u002Fgithub.com\u002Fopen-telemetry\u002Fsemantic-conventions-genai\u002Fblob\u002Fmain\u002Fdocs\u002Fgen-ai\u002Fgen-ai-spans.md",{"title":457,"url":458},"GenAI conventions: agent spans","https:\u002F\u002Fgithub.com\u002Fopen-telemetry\u002Fsemantic-conventions-genai\u002Fblob\u002Fmain\u002Fdocs\u002Fgen-ai\u002Fgen-ai-agent-spans.md",{"title":460,"url":461},"GenAI conventions: metrics","https:\u002F\u002Fgithub.com\u002Fopen-telemetry\u002Fsemantic-conventions-genai\u002Fblob\u002Fmain\u002Fdocs\u002Fgen-ai\u002Fgen-ai-metrics.md",{"title":463,"url":464},"GenAI conventions: inference token metrics","https:\u002F\u002Fgithub.com\u002Fopen-telemetry\u002Fsemantic-conventions-genai\u002Fblob\u002Fmain\u002Fdocs\u002Fgen-ai\u002Fgen-ai-token-metrics.md",{"title":466,"url":467},"GenAI conventions: events (gen_ai.evaluation.result)","https:\u002F\u002Fgithub.com\u002Fopen-telemetry\u002Fsemantic-conventions-genai\u002Fblob\u002Fmain\u002Fdocs\u002Fgen-ai\u002Fgen-ai-events.md",{"title":469,"url":470},"GenAI conventions: Model Context Protocol","https:\u002F\u002Fgithub.com\u002Fopen-telemetry\u002Fsemantic-conventions-genai\u002Fblob\u002Fmain\u002Fdocs\u002Fgen-ai\u002Fmcp.md",{"title":472,"url":473},"OpenTelemetry docs: GenAI conventions moved notice","https:\u002F\u002Fopentelemetry.io\u002Fdocs\u002Fspecs\u002Fsemconv\u002Fgen-ai\u002F",{"title":475,"url":476},"OpenTelemetry docs: Sampling","https:\u002F\u002Fopentelemetry.io\u002Fdocs\u002Fconcepts\u002Fsampling\u002F",{"title":478,"url":479},"John Hodge: The state of the OpenTelemetry GenAI semantic conventions (July 2026)","https:\u002F\u002Fjohn-hodge.com\u002Fblog\u002Fopentelemetry-genai-semantic-conventions\u002F",{"title":481,"url":482},"Langfuse docs: OpenTelemetry integration","https:\u002F\u002Flangfuse.com\u002Fdocs\u002Fopentelemetry\u002Fget-started",{"title":484,"url":485},"Langfuse repository and licence","https:\u002F\u002Fgithub.com\u002Flangfuse\u002Flangfuse",{"title":487,"url":488},"Arize Phoenix repository","https:\u002F\u002Fgithub.com\u002FArize-ai\u002Fphoenix",{"title":490,"url":491},"Arize OpenInference repository","https:\u002F\u002Fgithub.com\u002FArize-ai\u002Fopeninference",{"title":493,"url":494},"Datadog docs: OpenTelemetry instrumentation for LLM Observability","https:\u002F\u002Fdocs.datadoghq.com\u002Fllm_observability\u002Finstrumentation\u002Fotel_instrumentation\u002F",{"title":496,"url":497},"Honeycomb docs: Send data with OpenTelemetry","https:\u002F\u002Fdocs.honeycomb.io\u002Fsend-data\u002Fopentelemetry\u002F","\u002Fimages\u002Fblog\u002Fagent-observability-opentelemetry\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fagent-observability-opentelemetry\u002Fog.jpg",[61,62,63],"Observability for LLM agents with OpenTelemetry: traces, tokens, PII and evals","How to trace LLM agents with OpenTelemetry: GenAI semantic conventions and their status, span tree, token metrics, sampling, PII, evals and tool options.","Diagram: an agent run fans out into OpenTelemetry spans for model calls, tool calls, token metrics and evaluation results, exported to a trace backend.",{"slug":505,"published":506,"minutes":507,"category":7,"tags":508,"keywords":512,"about":521,"sources":526,"cover":534,"og":535,"expertise":59,"locales":536,"lang":61,"title":537,"description":538,"coverAlt":539},"llm-cost-latency-prompt-caching-routing","2026-06-01",9,[509,510,369,511],"Prompt caching","Model routing","Latency",[513,514,515,516,517,518,519,520],"prompt caching","LLM cost optimization","LLM latency","model routing","batch API LLM","LLM cost per request","cheaper LLM model","cache hit rate LLM",[522,523],{"name":386,"url":387},{"name":524,"url":525},"Latency (engineering)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FLatency_(engineering)",[527,530,533],{"title":528,"url":529},"Claude API docs: Prompt caching","https:\u002F\u002Fplatform.claude.com\u002Fdocs\u002Fen\u002Fdocs\u002Fbuild-with-claude\u002Fprompt-caching",{"title":531,"url":532},"OpenAI API docs: Prompt caching","https:\u002F\u002Fdevelopers.openai.com\u002Fapi\u002Fdocs\u002Fguides\u002Fprompt-caching",{"title":408,"url":409},"\u002Fimages\u002Fblog\u002Fllm-cost-latency-prompt-caching-routing\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fllm-cost-latency-prompt-caching-routing\u002Fog.jpg",[61,62,63],"Prompt caching and model routing: cutting LLM cost and latency","Prompt caching, cheap-model-first routing and batch APIs are the levers that cut LLM cost and latency in production. Here is how to use each one.","Four relative cost bars for one request: expensive model without a cache, cheaper model, cached prefix, and cached prefix in a batch job.",{"slug":541,"published":542,"minutes":364,"category":7,"tags":543,"keywords":548,"about":555,"sources":560,"cover":579,"og":580,"expertise":59,"locales":581,"lang":61,"title":582,"description":583,"coverAlt":584},"llm-evals-for-product-features","2026-05-15",[544,545,546,547],"LLM evals","LLM-as-judge","Error analysis","CI",[544,549,545,550,551,552,553,554],"AI evals","eval-driven development","pass^k vs pass@k","agent evaluation harness","regression eval suite","error analysis LLM",[556,557],{"name":386,"url":387},{"name":558,"url":559},"Software testing","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FSoftware_testing",[561,564,567,570,573,576],{"title":562,"url":563},"Anthropic: Demystifying evals for AI agents (2026)","https:\u002F\u002Fwww.anthropic.com\u002Fengineering\u002Fdemystifying-evals-for-ai-agents",{"title":565,"url":566},"Hamel Husain: LLM evals FAQ (updated September 2026)","https:\u002F\u002Fhamel.dev\u002Fblog\u002Fposts\u002Fevals-faq\u002F",{"title":568,"url":569},"Shankar et al., Who Validates the Validators? (2024)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2404.12272",{"title":571,"url":572},"Zheng et al., Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (2023)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2306.05685",{"title":574,"url":575},"Yao et al., tau-bench: A Benchmark for Tool-Agent-User Interaction (2024)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2406.12045",{"title":577,"url":578},"Anthropic: Quantifying infrastructure noise in agentic coding evals (2026)","https:\u002F\u002Fwww.anthropic.com\u002Fengineering\u002Finfrastructure-noise","\u002Fimages\u002Fblog\u002Fllm-evals-for-product-features\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fllm-evals-for-product-features\u002Fog.jpg",[61,62,63],"LLM evals for product features: from hand-read traces to a CI gate","LLM evals turn a vibe check into a test suite: error analysis on real traces, grader choice, a validated LLM judge, pass^k and a CI gate.","A pipeline from real traces through open coding and a counted failure taxonomy to graders, a validated judge and a CI gate.",1791386958061]