[{"data":1,"prerenderedAt":624},["ShallowReactive",2],{"blog-artificial-analysis-leaderboard-claude-opus-5-5-en":3},{"slug":4,"published":5,"minutes":6,"category":7,"tags":8,"keywords":13,"about":22,"sources":32,"cover":51,"og":52,"expertise":53,"locales":54,"lang":55,"title":58,"description":59,"coverAlt":60,"metaTitle":61,"takeaways":62,"faq":68,"toc":84,"blocks":109,"others":397},"artificial-analysis-leaderboard-claude-opus-5-5","2026-09-28",8,"llmops",[9,10,11,12],"Claude Opus 5.5","Artificial Analysis","LLM benchmarks","LLM cost",[9,14,15,16,17,18,19,20,21],"Artificial Analysis Intelligence Index","AI model leaderboard","Opus 5.5 benchmark","Opus 5.5 vs GPT-6 Astra","LLM cost per task","reasoning effort setting","Opus 5.5 pricing","best LLM September 2026",[23,26,29],{"name":24,"url":25},"Claude (language model)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FClaude_(language_model)",{"name":27,"url":28},"Large language model","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FLarge_language_model",{"name":30,"url":31},"Benchmark (computing)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FBenchmark_(computing)",[33,36,39,42,45,48],{"title":34,"url":35},"Artificial Analysis: Claude Opus 5.5 takes the top spot on the Artificial Analysis Intelligence Index (22 September 2026)","https:\u002F\u002Fartificialanalysis.ai\u002Farticles\u002Fclaude-opus-5-5",{"title":37,"url":38},"Artificial Analysis: Claude Opus 5.5 (max) model page","https:\u002F\u002Fartificialanalysis.ai\u002Fmodels\u002Fclaude-opus-5-5",{"title":40,"url":41},"Artificial Analysis: Claude Opus 5.5 (medium) model page","https:\u002F\u002Fartificialanalysis.ai\u002Fmodels\u002Fclaude-opus-5-5-medium",{"title":43,"url":44},"Artificial Analysis: Claude Opus 5 (max) model page","https:\u002F\u002Fartificialanalysis.ai\u002Fmodels\u002Fclaude-opus-5",{"title":46,"url":47},"OfficeChai: Claude Opus 5.5 creates a 5-point lead over GPT-6 Astra","https:\u002F\u002Fofficechai.com\u002Fai\u002Fclaude-opus-5-5-creates-5-point-lead-over-gpt-6-astra-jumps-to-top-spot-on-artificial-analysis-intelligence-index\u002F",{"title":49,"url":50},"Claude API docs: Models overview, context windows and prices (as of September 2026)","https:\u002F\u002Fplatform.claude.com\u002Fdocs\u002Fen\u002Fabout-claude\u002Fmodels\u002Foverview","\u002Fimages\u002Fblog\u002Fartificial-analysis-leaderboard-claude-opus-5-5\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fartificial-analysis-leaderboard-claude-opus-5-5\u002Fog.jpg","ai-engineer",[55,56,57],"en","de","hu","Claude Opus 5.5 takes #1 on Artificial Analysis, and medium effort is the real story","Claude Opus 5.5 scores 58 on the Artificial Analysis Intelligence Index, five points clear. The bigger story is what medium effort buys you per task.","Horizontal bars of Intelligence Index scores: Claude Opus 5.5 at max effort 58, GPT-6 Astra and Claude Fable 5.1 53, Opus 5 51, Opus 5.5 at medium effort 51.","Claude Opus 5.5 tops Artificial Analysis · Balázs Csorba",[63,64,65,66,67],"Claude Opus 5.5 at max effort scores 58 on the Artificial Analysis Intelligence Index v4.3.2, first of 211 models and five points ahead of GPT-6 Astra and Claude Fable 5.1 on 53.","At medium effort Opus 5.5 scores 51, the same as Opus 5 at max effort, for $1.34 per task instead of $5.86: about 77% cheaper for the same score.","Four of its five effort settings sit on Artificial Analysis' intelligence versus cost per task frontier.","The catch is verbosity and latency: about 119,000 output tokens per task at max effort and a reported time to first token of 682.71 seconds, against 13.20 seconds at medium.","Prices fell to $4 and $20 per million tokens and cache reads by 60% to $0.20, so medium effort plus a cached prefix is the sensible production default.",[69,72,75,78,81],{"q":70,"a":71},"What is the Artificial Analysis Intelligence Index?","It is a composite score published by Artificial Analysis, an independent benchmarking company that runs the same evaluations on every major model. Version 4.3.2 combines ten evaluations, including Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDPval-AA and AA-Omniscience, and is published next to cost per task, output tokens and speed.",{"q":73,"a":74},"Which model is number one on the Artificial Analysis leaderboard?","As of 28 September 2026, Claude Opus 5.5 at max effort, with an Intelligence Index score of 58. GPT-6 Astra and Claude Fable 5.1 follow on 53 and Claude Opus 5 at max effort scores 51. Artificial Analysis published the result on 22 September 2026.",{"q":76,"a":77},"Is Claude Opus 5.5 more expensive to run than Opus 5?","At max effort it costs about the same per task, $5.98 against $5.86, because the lower token price offsets the roughly 1.6 times more output tokens it uses. At medium effort it matches Opus 5's max-effort score for $1.34 per task, so for the same quality it is much cheaper.",{"q":79,"a":80},"Which Opus 5.5 effort setting should I use in production?","Start with medium for interactive features: it scores 51 on the index and Artificial Analysis reports a time to first token of 13.20 seconds. Reserve high, xhigh and max for background work where more than ten minutes of latency is acceptable, escalate by a rule you define in advance, and confirm the choice on your own eval set.",{"q":82,"a":83},"How much did Opus 5.5 prices change?","Input and output fell 20%, from $5 and $25 to $4 and $20 per million tokens. Cache reads fell 60%, from $0.50 to $0.20 per million tokens, and writes to the five-minute cache from $6.25 to $5, which makes a stable, cached prompt prefix worth more than on any previous Claude model.",[85,88,91,94,97,100,103,106],{"id":86,"title":87},"what-the-index-measures","What the Artificial Analysis Intelligence Index measures",{"id":89,"title":90},"the-leaderboard","The leaderboard: 58, and a five-point lead",{"id":92,"title":93},"where-the-points-come-from","Where the points come from",{"id":95,"title":96},"medium-effort-is-the-headline","Medium effort is the real headline",{"id":98,"title":99},"the-catch","The catch: tokens and waiting time",{"id":101,"title":102},"the-price-cut","The price cut, and why caching matters more now",{"id":104,"title":105},"what-i-would-do","What I would actually do with this",{"id":107,"title":108},"sources","Sources",[110,122,134,137,140,143,144,155,164,171,172,175,204,213,214,217,228,235,238,292,293,296,299,306,307,314,322,323,326,364,372,373],{"type":111,"content":112},"paragraph",[113,116,117,121],{"tag":114,"children":115},"strong",[9]," is the new number one on the ",{"tag":118,"href":119,"children":120},"a","https:\u002F\u002Fartificialanalysis.ai\u002F",[10]," leaderboard, and not by a rounding error. At its maximum effort setting it scores 58 on the Artificial Analysis Intelligence Index, five points clear of the next models, the highest score the index has recorded. That is the headline you have probably already seen.",{"type":111,"content":123},[124,125,129,130,133],"It is not the interesting part. The interesting part is further down the same page: at ",{"tag":126,"children":127},"em",[128],"medium"," effort, Opus 5.5 scores exactly what last generation's flagship scored at ",{"tag":126,"children":131},[132],"maximum"," effort, for less than a quarter of the cost per task. This article walks through the leaderboard, the per-benchmark numbers, the efficiency data and the catch, all as published by Artificial Analysis as of 28 September 2026, and ends with what I would actually change in a production setup because of it.",{"type":135,"level":136,"id":86,"text":87},"heading",2,{"type":111,"content":138},[139],"Artificial Analysis is an independent benchmarking company that runs the same evaluations against every major model and publishes intelligence, speed and price side by side. Its Intelligence Index is a single composite number. Version 4.3.2 combines ten evaluations: AA-Briefcase, GDPval-AA, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR. The mix leans on agentic and knowledge-work tasks rather than trivia: office work, terminal work, scientific coding, long-context reasoning and a hallucination-aware knowledge test.",{"type":111,"content":141},[142],"Two properties make it more useful than a vendor's own launch chart. The same harness runs every model, so the numbers are comparable across providers. And the index is published next to cost per task, output tokens used and speed, so you can see what a score costs, which vendor charts rarely show. Keep both in mind for the rest of this article, because the efficiency story only exists because of the second one.",{"type":135,"level":136,"id":89,"text":90},{"type":111,"content":145},[146,147,150,151,154],"Artificial Analysis published its evaluation on 22 September 2026 under the title ",{"tag":118,"href":35,"children":148},[149],"Claude Opus 5.5 takes the top spot",". Opus 5.5 at max effort scores 58 and ranks first of 211 models on the index, where the median model scores 26. Behind it, GPT-6 Astra and Claude Fable 5.1 are tied on 53, and Claude Opus 5, the model Opus 5.5 replaces, sits on 51 and ",{"tag":118,"href":44,"children":152},[153],"ranks eleventh",".",{"type":156,"attrs":157,"inner":161,"caption":162},"diagram",{"viewBox":158,"role":159,"aria-labelledby":160},"0 0 720 330","img","d1-aa-t d1-aa-d","\u003Ctitle id=\"d1-aa-t\">Artificial Analysis Intelligence Index, top of the leaderboard\u003C\u002Ftitle>\u003Cdesc id=\"d1-aa-d\">Horizontal bars show Intelligence Index scores as of September 2026. Claude Opus 5.5 at max effort scores 58, GPT-6 Astra 53, Claude Fable 5.1 53, Claude Opus 5 at max effort 51, Claude Opus 5.5 at medium effort also 51, and the median of 211 models 26.\u003C\u002Fdesc>\u003Ctext x=\"20\" y=\"28\" class=\"d-title\">INTELLIGENCE INDEX V4.3.2\u003C\u002Ftext>\u003Ctext x=\"700\" y=\"28\" text-anchor=\"end\" class=\"d-label\">source: Artificial Analysis\u003C\u002Ftext>\u003Ctext x=\"20\" y=\"72\" class=\"d-text\">Opus 5.5 (max)\u003C\u002Ftext>\u003Crect x=\"200\" y=\"50\" width=\"435\" height=\"34\" rx=\"8\" class=\"d-accent\" \u002F>\u003Ctext x=\"647\" y=\"73\" class=\"d-label\">58\u003C\u002Ftext>\u003Ctext x=\"20\" y=\"116\" class=\"d-small\">GPT-6 Astra\u003C\u002Ftext>\u003Crect x=\"200\" y=\"94\" width=\"398\" height=\"34\" rx=\"8\" class=\"d-box\" \u002F>\u003Ctext x=\"610\" y=\"117\" class=\"d-label\">53\u003C\u002Ftext>\u003Ctext x=\"20\" y=\"160\" class=\"d-small\">Claude Fable 5.1\u003C\u002Ftext>\u003Crect x=\"200\" y=\"138\" width=\"398\" height=\"34\" rx=\"8\" class=\"d-box\" \u002F>\u003Ctext x=\"610\" y=\"161\" class=\"d-label\">53\u003C\u002Ftext>\u003Ctext x=\"20\" y=\"204\" class=\"d-small\">Opus 5 (max)\u003C\u002Ftext>\u003Crect x=\"200\" y=\"182\" width=\"383\" height=\"34\" rx=\"8\" class=\"d-box\" \u002F>\u003Ctext x=\"595\" y=\"205\" class=\"d-label\">51\u003C\u002Ftext>\u003Ctext x=\"20\" y=\"248\" class=\"d-small\">Opus 5.5 (medium)\u003C\u002Ftext>\u003Crect x=\"200\" y=\"226\" width=\"383\" height=\"34\" rx=\"8\" class=\"d-mint\" \u002F>\u003Ctext x=\"595\" y=\"249\" class=\"d-label\">51\u003C\u002Ftext>\u003Ctext x=\"20\" y=\"292\" class=\"d-small\">median, 211 models\u003C\u002Ftext>\u003Crect x=\"200\" y=\"270\" width=\"195\" height=\"34\" rx=\"8\" class=\"d-gold\" \u002F>\u003Ctext x=\"407\" y=\"293\" class=\"d-label\">26\u003C\u002Ftext>",[163],"Opus 5.5 at max effort leads by five points; at medium effort it matches the previous flagship at max.",{"type":111,"content":165},[166,167,170],"Five points on a composite index is a large gap. For comparison, the whole spread between Opus 5 and the joint second place is two points. The ",{"tag":118,"href":47,"children":168},[169],"OfficeChai coverage"," framed it as a five-point lead over GPT-6 Astra, and that is the fair reading: it is a new top score, not a tie broken by noise.",{"type":135,"level":136,"id":92,"text":93},{"type":111,"content":173},[174],"A composite can hide a single outlier benchmark, so the per-evaluation numbers matter more than the total. Artificial Analysis reports the following for Opus 5.5 at max effort:",{"type":176,"ordered":177,"items":178},"list",false,[179,184,189,194,199],[180,183],{"tag":114,"children":181},[182],"Humanity's Last Exam:"," 61.4%, against a previous best of 59.1%.",[185,188],{"tag":114,"children":186},[187],"SciCode:"," 66.9%, against a previous best of 63.1%.",[190,193],{"tag":114,"children":191},[192],"Terminal-Bench 4.0:"," 59.6%, level with GPT-6 Astra at xhigh effort and 11 points above Opus 5.",[195,198],{"tag":114,"children":196},[197],"AA-Briefcase:"," an Elo of 1822, 143 points above Claude Fable 5.1.",[200,203],{"tag":114,"children":201},[202],"GDPval-AA v2.1:"," 1,846, which is 111 above Fable 5.1 and 138 above Opus 5.",{"type":111,"content":205},[206,207,212],"The pattern is broad rather than spiky. The largest gains are in knowledge work (AA-Briefcase and GDPval-AA, both built around realistic office tasks) and in terminal work, which is where coding agents spend their time. Terminal-Bench is the one place where it does not lead outright: GPT-6 Astra matches it. If your workload is an ",{"tag":208,"to":209,"children":210},"link","\u002Fblog\u002Fagent-loop-explained",[211],"agent loop"," driving a shell, the honest summary is \"joint best\", not \"best\".",{"type":135,"level":136,"id":95,"text":96},{"type":111,"content":215},[216],"Opus 5.5 exposes five effort settings: low, medium, high, xhigh and max. Effort controls how much the model is allowed to think before answering, and therefore how many output tokens you pay for. Artificial Analysis measured all five. The index scores are 42 at low, 51 at medium, 54 at high, 56 at xhigh and 58 at max.",{"type":111,"content":218},[219,220,223,224,227],"Now put two of those numbers next to each other. Opus 5 at max effort scores 51 at ",{"tag":118,"href":44,"children":221},[222],"$5.86 per task",". Opus 5.5 at medium effort also scores 51, at ",{"tag":118,"href":41,"children":225},[226],"$1.34 per task",", ranking eighth on the whole leaderboard. Same score, about 77% cheaper per task, which is roughly a 4.4x difference. That is the sentence I would put on the launch slide, and it is not on the launch slide.",{"type":156,"attrs":229,"inner":232,"caption":233},{"viewBox":230,"role":159,"aria-labelledby":231},"0 0 720 320","d2-aa-t d2-aa-d","\u003Ctitle id=\"d2-aa-t\">Intelligence Index score against cost per task\u003C\u002Ftitle>\u003Cdesc id=\"d2-aa-d\">A scatter chart with cost per task on the horizontal axis and Intelligence Index score on the vertical axis. Claude Opus 5.5 at medium effort sits at 1.34 dollars and a score of 51. Claude Opus 5 at max effort sits at 5.86 dollars and the same score of 51. Claude Opus 5.5 at max effort sits at 5.98 dollars and a score of 58. The medium setting reaches the previous flagship's score at about a quarter of the cost.\u003C\u002Fdesc>\u003Ctext x=\"20\" y=\"28\" class=\"d-title\">SCORE VS COST PER TASK\u003C\u002Ftext>\u003Ctext x=\"700\" y=\"28\" text-anchor=\"end\" class=\"d-label\">source: Artificial Analysis\u003C\u002Ftext>\u003Cpath d=\"M90 270 H680\" class=\"d-line\" \u002F>\u003Cpath d=\"M90 270 V50\" class=\"d-line\" \u002F>\u003Ctext x=\"90\" y=\"292\" text-anchor=\"middle\" class=\"d-label\">$0\u003C\u002Ftext>\u003Ctext x=\"290\" y=\"292\" text-anchor=\"middle\" class=\"d-label\">$2\u003C\u002Ftext>\u003Ctext x=\"490\" y=\"292\" text-anchor=\"middle\" class=\"d-label\">$4\u003C\u002Ftext>\u003Ctext x=\"690\" y=\"292\" text-anchor=\"end\" class=\"d-label\">$6 per task\u003C\u002Ftext>\u003Ctext x=\"80\" y=\"274\" text-anchor=\"end\" class=\"d-label\">45\u003C\u002Ftext>\u003Ctext x=\"80\" y=\"174\" text-anchor=\"end\" class=\"d-label\">52\u003C\u002Ftext>\u003Ctext x=\"80\" y=\"74\" text-anchor=\"end\" class=\"d-label\">59\u003C\u002Ftext>\u003Cpath d=\"M224 184 H676\" class=\"d-line d-dash\" \u002F>\u003Ccircle cx=\"224\" cy=\"184\" r=\"11\" class=\"d-mint\" \u002F>\u003Ctext x=\"224\" y=\"220\" text-anchor=\"middle\" class=\"d-text\">Opus 5.5 medium\u003C\u002Ftext>\u003Ctext x=\"224\" y=\"238\" text-anchor=\"middle\" class=\"d-small\">51 at $1.34\u003C\u002Ftext>\u003Ccircle cx=\"676\" cy=\"184\" r=\"11\" class=\"d-box\" \u002F>\u003Ctext x=\"660\" y=\"220\" text-anchor=\"end\" class=\"d-text\">Opus 5 max\u003C\u002Ftext>\u003Ctext x=\"660\" y=\"238\" text-anchor=\"end\" class=\"d-small\">51 at $5.86\u003C\u002Ftext>\u003Ccircle cx=\"688\" cy=\"84\" r=\"11\" class=\"d-accent\" \u002F>\u003Ctext x=\"668\" y=\"74\" text-anchor=\"end\" class=\"d-text\">Opus 5.5 max\u003C\u002Ftext>\u003Ctext x=\"668\" y=\"92\" text-anchor=\"end\" class=\"d-small\">58 at $5.98\u003C\u002Ftext>\u003Ctext x=\"450\" y=\"178\" text-anchor=\"middle\" class=\"d-label\">same score, about 4.4x cheaper\u003C\u002Ftext>",[234],"Opus 5.5 at medium effort matches Opus 5 at max effort for about a quarter of the cost per task.",{"type":111,"content":236},[237],"Artificial Analysis also notes that four of the five effort settings land on its intelligence versus cost per task frontier. In plain terms: for four of the five settings, no other measured model gives you a higher score for the same money. The setting ladder is not marketing; each step up buys real points, and you can choose where on the curve to sit.",{"type":239,"head":240,"rows":253},"table",[241,243,245,247,249,251],[242],"Model and setting",[244],"Index score",[246],"Rank",[248],"Cost per task",[250],"Output tokens, full index",[252],"Output speed",[254,267,280],[255,257,259,261,263,265],[256],"Claude Opus 5.5, max",[258],"58",[260],"1 of 211",[262],"$5.98",[264],"260M",[266],"95.5 tokens\u002Fs",[268,270,272,274,276,278],[269],"Claude Opus 5.5, medium",[271],"51",[273],"8",[275],"$1.34",[277],"38M",[279],"81.7 tokens\u002Fs",[281,283,284,286,288,290],[282],"Claude Opus 5, max",[271],[285],"11",[287],"$5.86",[289],"140M",[291],"60.5 tokens\u002Fs",{"type":135,"level":136,"id":98,"text":99},{"type":111,"content":294},[295],"Now the part the headline leaves out. Opus 5.5 at max effort is the most verbose model at the top of the table. Artificial Analysis counts about 119,000 output tokens per index task, against about 73,000 for Opus 5, about 78,000 for Fable 5.1 and about 27,000 for GPT-6 Astra. Over the full index it used 260 million output tokens, where the median model used 88 million. It stays level with Opus 5 on cost per task only because the price per token dropped, not because it thinks less.",{"type":111,"content":297},[298],"The second cost is time. The max-effort model page reports a time to first token of 682.71 seconds, which is more than eleven minutes before the first token arrives, even though the output speed of 95.5 tokens per second is faster than Opus 5. At medium effort the same page reports 13.20 seconds. For a background job, eleven minutes is fine. For anything a user is watching, it is not a setting, it is an outage.",{"type":300,"variant":301,"title":302,"body":303},"callout","warn","A benchmark is not your workload",[304],[305],"The index measures ten specific evaluations under one harness. Your prompts, tools and failure modes are different, and a five-point lead on a composite can shrink or vanish on a narrow task. Treat the leaderboard as a shortlist, and decide with your own eval set.",{"type":135,"level":136,"id":101,"text":102},{"type":111,"content":308},[309,310,313],"Opus 5.5 is priced at $4 per million input tokens and $20 per million output tokens, down 20% from Opus 5's $5 and $25, according to the ",{"tag":118,"href":50,"children":311},[312],"Claude models overview"," and the Artificial Analysis model page. Cache reads dropped harder, by 60%, from $0.50 to $0.20 per million tokens, and writes to the five-minute cache went from $6.25 to $5. The context window is one million tokens.",{"type":111,"content":315},[316,317,321],"The cache numbers are the ones to act on. A long-running agent resends its history on every turn, and a more verbose model makes that history grow faster. At $0.20 per million cached tokens, a stable prefix is close to free, and an uncached one is twenty times more expensive. If you have not already stabilized your prompt prefix, the ",{"tag":208,"to":318,"children":319},"\u002Fblog\u002Fllm-cost-latency-prompt-caching-routing",[320],"prompt caching and routing guide"," explains how, and on Opus 5.5 the saving is larger than on any previous Claude model.",{"type":135,"level":136,"id":104,"text":105},{"type":111,"content":324},[325],"The leaderboard tells you which model is strongest. It does not tell you which setting to run, and that is where the money is. This is the order I would work in:",{"type":176,"ordered":327,"items":328},true,[329,334,339,349,354],[330,333],{"tag":114,"children":331},[332],"Make medium the default."," It matches the previous flagship at a quarter of the cost and answers in seconds, not minutes. Most interactive features should start here.",[335,338],{"tag":114,"children":336},[337],"Reserve max for work nobody is waiting for:"," overnight analysis, large refactors run by a coding agent, eval generation. Budget the tokens explicitly, because 119,000 output tokens per task adds up.",[340,343,344,348],{"tag":114,"children":341},[342],"Escalate by rule, not by feel."," Run medium first and move to high or max only when a check fails, with the escalation condition written down in advance. A small typed decision model, as in the ",{"tag":208,"to":345,"children":346},"\u002Fblog\u002Fjev-typed-decisions-llm-routing",[347],"Jev routing article",", is a cheap way to make that call.",[350,353],{"tag":114,"children":351},[352],"Cache aggressively."," At a 60% lower cache read price, prefix stability is now the biggest single saving on long agent sessions.",[355,358,359,363],{"tag":114,"children":356},[357],"Re-run your own evals before switching."," Compare Opus 5.5 medium against whatever you run today on the same graded set; the ",{"tag":208,"to":360,"children":361},"\u002Fblog\u002Fllm-evals-for-product-features",[362],"evals guide"," shows how to build one in an afternoon.",{"type":111,"content":365},[366,367,371],"The pattern behind all five is the same: the effort setting is now a product decision, not a model detail. The team that picks it per feature will get last generation's flagship quality at a fraction of last generation's bill. The team that leaves everything on max will pay for eleven-minute answers. If you want help making that decision for a real system, the ",{"tag":208,"to":368,"children":369},"\u002Fexpertise\u002Fai-engineer",[370],"AI engineering"," page describes how I approach it.",{"type":135,"level":136,"id":107,"text":108},{"type":176,"ordered":327,"items":374},[375,378,381,384,387,391,394],[376],{"tag":118,"href":35,"children":377},[34],[379],{"tag":118,"href":38,"children":380},[37],[382],{"tag":118,"href":41,"children":383},[40],[385],{"tag":118,"href":44,"children":386},[43],[388],{"tag":118,"href":119,"children":389},[390],"Artificial Analysis: model leaderboard and Intelligence Index",[392],{"tag":118,"href":47,"children":393},[46],[395],{"tag":118,"href":50,"children":396},[49],[398,452,496,531],{"slug":399,"published":400,"minutes":401,"category":7,"tags":402,"keywords":408,"about":418,"sources":427,"cover":446,"og":447,"expertise":53,"locales":448,"lang":55,"title":449,"description":450,"coverAlt":451},"jev-typed-decisions-llm-routing","2026-09-27",10,[403,404,405,406,407],"LLM routing","Classification","Calibration","OpenRouter","Human in the loop",[409,410,411,412,413,414,415,416,417],"llm routing","llm classification","calibrated confidence llm","typed llm output","jev model","openrouter decisions api","ai triage of code review findings","how to route low-confidence llm answers to a human","llm classifier vs structured output",[419,422,425],{"name":420,"url":421},"Calibration (statistics)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FCalibration_(statistics)",{"name":423,"url":424},"Statistical classification","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FStatistical_classification",{"name":406,"url":426},"https:\u002F\u002Fopenrouter.ai\u002F",[428,431,434,437,440,443],{"title":429,"url":430},"OpenRouter: Jev documentation","https:\u002F\u002Fopenrouter.ai\u002Fdocs\u002Fguides\u002Fcommunity\u002Fjev",{"title":432,"url":433},"OpenRouter: What is Jev?","https:\u002F\u002Fopenrouter.ai\u002Fblog\u002Finsights\u002Fwhat-is-jev\u002F",{"title":435,"url":436},"OpenRouter: How to use Jev","https:\u002F\u002Fopenrouter.ai\u002Fblog\u002Ftutorials\u002Fhow-to-use-jev\u002F",{"title":438,"url":439},"Guo et al.: On Calibration of Modern Neural Networks","https:\u002F\u002Farxiv.org\u002Fabs\u002F1706.04599",{"title":441,"url":442},"Claude API: Structured outputs","https:\u002F\u002Fplatform.claude.com\u002Fdocs\u002Fen\u002Fbuild-with-claude\u002Fstructured-outputs",{"title":444,"url":445},"AI SDK: Generating structured data","https:\u002F\u002Fai-sdk.dev\u002Fdocs\u002Fai-sdk-core\u002Fgenerating-structured-data","\u002Fimages\u002Fblog\u002Fjev-typed-decisions-llm-routing\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fjev-typed-decisions-llm-routing\u002Fog.jpg",[55,56,57],"Typed decisions for LLM routing and triage: calibrated confidence with Jev","LLM routing with typed decisions: Choice, Score and yes-probability answers with calibrated confidence, thresholds and human hand-off, using Jev.","Fan-out diagram: one diff hunk as state feeds a Choice, a Score and a Noul question in a single decisions call.",{"slug":453,"published":400,"minutes":6,"category":7,"tags":454,"keywords":459,"about":466,"sources":471,"cover":490,"og":491,"expertise":53,"locales":492,"lang":55,"title":493,"description":494,"coverAlt":495},"llm-evals-for-product-features",[455,456,457,458],"LLM evals","LLM-as-judge","Error analysis","CI",[455,460,456,461,462,463,464,465],"AI evals","eval-driven development","pass^k vs pass@k","agent evaluation harness","regression eval suite","error analysis LLM",[467,468],{"name":27,"url":28},{"name":469,"url":470},"Software testing","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FSoftware_testing",[472,475,478,481,484,487],{"title":473,"url":474},"Anthropic: Demystifying evals for AI agents (2026)","https:\u002F\u002Fwww.anthropic.com\u002Fengineering\u002Fdemystifying-evals-for-ai-agents",{"title":476,"url":477},"Hamel Husain: LLM evals FAQ (updated September 2026)","https:\u002F\u002Fhamel.dev\u002Fblog\u002Fposts\u002Fevals-faq\u002F",{"title":479,"url":480},"Shankar et al., Who Validates the Validators? (2024)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2404.12272",{"title":482,"url":483},"Zheng et al., Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (2023)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2306.05685",{"title":485,"url":486},"Yao et al., tau-bench: A Benchmark for Tool-Agent-User Interaction (2024)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2406.12045",{"title":488,"url":489},"Anthropic: Quantifying infrastructure noise in agentic coding evals (2026)","https:\u002F\u002Fwww.anthropic.com\u002Fengineering\u002Finfrastructure-noise","\u002Fimages\u002Fblog\u002Fllm-evals-for-product-features\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fllm-evals-for-product-features\u002Fog.jpg",[55,56,57],"LLM evals for product features: from hand-read traces to a CI gate","LLM evals turn a vibe check into a test suite: error analysis on real traces, grader choice, a validated LLM judge, pass^k and a CI gate.","A pipeline from real traces through open coding and a counted failure taxonomy to graders, a validated judge and a CI gate.",{"slug":497,"published":400,"minutes":498,"category":7,"tags":499,"keywords":503,"about":512,"sources":517,"cover":525,"og":526,"expertise":53,"locales":527,"lang":55,"title":528,"description":529,"coverAlt":530},"llm-cost-latency-prompt-caching-routing",9,[500,501,12,502],"Prompt caching","Model routing","Latency",[504,505,506,507,508,509,510,511],"prompt caching","LLM cost optimization","LLM latency","model routing","batch API LLM","LLM cost per request","cheaper LLM model","cache hit rate LLM",[513,514],{"name":27,"url":28},{"name":515,"url":516},"Latency (engineering)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FLatency_(engineering)",[518,521,524],{"title":519,"url":520},"Claude API docs: Prompt caching","https:\u002F\u002Fplatform.claude.com\u002Fdocs\u002Fen\u002Fdocs\u002Fbuild-with-claude\u002Fprompt-caching",{"title":522,"url":523},"OpenAI API docs: Prompt caching","https:\u002F\u002Fdevelopers.openai.com\u002Fapi\u002Fdocs\u002Fguides\u002Fprompt-caching",{"title":49,"url":50},"\u002Fimages\u002Fblog\u002Fllm-cost-latency-prompt-caching-routing\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fllm-cost-latency-prompt-caching-routing\u002Fog.jpg",[55,56,57],"Prompt caching and model routing: cutting LLM cost and latency","Prompt caching, cheap-model-first routing and batch APIs are the levers that cut LLM cost and latency in production. Here is how to use each one.","Four relative cost bars for one request: expensive model without a cache, cheaper model, cached prefix, and cached prefix in a batch job.",{"slug":532,"published":5,"minutes":533,"category":534,"tags":535,"keywords":540,"about":551,"sources":557,"cover":618,"og":619,"expertise":53,"locales":620,"lang":55,"title":621,"description":622,"coverAlt":623},"token-saving-tools-coding-agents-top-20",17,"agents",[536,537,538,539],"Claude Code","token usage","context engineering","developer tools",[541,542,543,544,545,546,547,548,504,549,550,538],"reduce Claude Code token usage","token saving tools for coding agents","rtk token killer","lean-ctx","context-mode MCP","Serena MCP","Repomix compress","MCP tool search","ccusage","coding agent cost",[552,553,554],{"name":27,"url":28},{"name":24,"url":25},{"name":555,"url":556},"Prompt engineering","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FPrompt_engineering",[558,561,564,567,570,573,576,579,582,585,588,591,594,597,600,603,606,609,612,615],{"title":559,"url":560},"rtk: a CLI proxy that filters shell output for coding agents (GitHub)","https:\u002F\u002Fgithub.com\u002Frtk-ai\u002Frtk",{"title":562,"url":563},"lean-ctx: context read modes and shell compression for coding agents (GitHub)","https:\u002F\u002Fgithub.com\u002Fyvgude\u002Flean-ctx",{"title":565,"url":566},"context-mode: sandboxed tool output with SQLite FTS5 search (GitHub)","https:\u002F\u002Fgithub.com\u002Fmksglu\u002Fcontext-mode",{"title":568,"url":569},"Serena: semantic code retrieval and editing over MCP (GitHub)","https:\u002F\u002Fgithub.com\u002Foraios\u002Fserena",{"title":571,"url":572},"token-savior: symbol index, memory and bash compaction over MCP (GitHub)","https:\u002F\u002Fgithub.com\u002FMibayy\u002Ftoken-savior",{"title":574,"url":575},"code-review-graph: a code graph for blast-radius reviews (GitHub)","https:\u002F\u002Fgithub.com\u002Ftirth8205\u002Fcode-review-graph",{"title":577,"url":578},"Aider documentation: repository map","https:\u002F\u002Faider.chat\u002Fdocs\u002Frepomap.html",{"title":580,"url":581},"Repomix: pack a repository into one AI-friendly file (GitHub)","https:\u002F\u002Fgithub.com\u002Fyamadashy\u002Frepomix",{"title":583,"url":584},"Context7: up-to-date library documentation for LLMs (GitHub)","https:\u002F\u002Fgithub.com\u002Fupstash\u002Fcontext7",{"title":586,"url":587},"claude-context: hybrid code search MCP (GitHub)","https:\u002F\u002Fgithub.com\u002Fzilliztech\u002Fclaude-context",{"title":589,"url":590},"caveman: terse output modes for coding agents (GitHub)","https:\u002F\u002Fgithub.com\u002FJuliusBrussee\u002Fcaveman",{"title":592,"url":593},"claude-token-efficient: an eight-rule CLAUDE.md (GitHub)","https:\u002F\u002Fgithub.com\u002Fdrona23\u002Fclaude-token-efficient",{"title":595,"url":596},"claude-code-router: a local model gateway for coding agents (GitHub)","https:\u002F\u002Fgithub.com\u002Fmusistudio\u002Fclaude-code-router",{"title":598,"url":599},"ccusage: token and cost reports from local agent logs (GitHub)","https:\u002F\u002Fgithub.com\u002Fryoppippi\u002Fccusage",{"title":601,"url":602},"LLMLingua: prompt compression (Microsoft, GitHub)","https:\u002F\u002Fgithub.com\u002Fmicrosoft\u002FLLMLingua",{"title":604,"url":605},"ComputingForGeeks: tools that reduce Claude Code token usage, tested (April 2026)","https:\u002F\u002Fcomputingforgeeks.com\u002Freduce-claude-code-token-usage-tools\u002F",{"title":607,"url":608},"Claude Code docs: manage costs effectively","https:\u002F\u002Fcode.claude.com\u002Fdocs\u002Fen\u002Fcosts",{"title":610,"url":611},"Claude Code docs: MCP, output limits and tool search","https:\u002F\u002Fcode.claude.com\u002Fdocs\u002Fen\u002Fmcp",{"title":613,"url":614},"Claude API docs: tool search tool","https:\u002F\u002Fplatform.claude.com\u002Fdocs\u002Fen\u002Fagents-and-tools\u002Ftool-use\u002Ftool-search-tool",{"title":616,"url":617},"Claude API docs: prompt caching","https:\u002F\u002Fplatform.claude.com\u002Fdocs\u002Fen\u002Fbuild-with-claude\u002Fprompt-caching","\u002Fimages\u002Fblog\u002Ftoken-saving-tools-coding-agents-top-20\u002Fcover.webp","\u002Fimages\u002Fblog\u002Ftoken-saving-tools-coding-agents-top-20\u002Fog.jpg",[55,56,57],"Top 20 ways to cut coding-agent tokens: rtk, lean-ctx, Serena and more, ranked by evidence","rtk, lean-ctx, context-mode, Serena and 16 more token savers for coding agents, ranked by evidence, with my own measurements on a real Nuxt codebase.","Bar chart falling from a 15,100-token build log to about 370 tokens after filtering, under the heading Top 20 token savers.",1790582553881]