Blog/LLMOps & evals
Claude Opus 5.5 takes #1 on Artificial Analysis, and medium effort is the real story
Claude Opus 5.5 scores 58 on the Artificial Analysis Intelligence Index, five points clear. The bigger story is what medium effort buys you per task.
Balázs Csorba··8 min read
- Claude Opus 5.5
- Artificial Analysis
- LLM benchmarks
- LLM cost

Key takeaways
- Claude Opus 5.5 at max effort scores 58 on the Artificial Analysis Intelligence Index v4.3.2, first of 211 models and five points ahead of GPT-6 Astra and Claude Fable 5.1 on 53.
- At medium effort Opus 5.5 scores 51, the same as Opus 5 at max effort, for $1.34 per task instead of $5.86: about 77% cheaper for the same score.
- Four of its five effort settings sit on Artificial Analysis' intelligence versus cost per task frontier.
- The catch is verbosity and latency: about 119,000 output tokens per task at max effort and a reported time to first token of 682.71 seconds, against 13.20 seconds at medium.
- Prices fell to $4 and $20 per million tokens and cache reads by 60% to $0.20, so medium effort plus a cached prefix is the sensible production default.
Claude Opus 5.5 is the new number one on the Artificial Analysis leaderboard, and not by a rounding error. At its maximum effort setting it scores 58 on the Artificial Analysis Intelligence Index, five points clear of the next models, the highest score the index has recorded. That is the headline you have probably already seen.
It is not the interesting part. The interesting part is further down the same page: at medium effort, Opus 5.5 scores exactly what last generation's flagship scored at maximum effort, for less than a quarter of the cost per task. This article walks through the leaderboard, the per-benchmark numbers, the efficiency data and the catch, all as published by Artificial Analysis as of 28 September 2026, and ends with what I would actually change in a production setup because of it.
What the Artificial Analysis Intelligence Index measures
Artificial Analysis is an independent benchmarking company that runs the same evaluations against every major model and publishes intelligence, speed and price side by side. Its Intelligence Index is a single composite number. Version 4.3.2 combines ten evaluations: AA-Briefcase, GDPval-AA, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR. The mix leans on agentic and knowledge-work tasks rather than trivia: office work, terminal work, scientific coding, long-context reasoning and a hallucination-aware knowledge test.
Two properties make it more useful than a vendor's own launch chart. The same harness runs every model, so the numbers are comparable across providers. And the index is published next to cost per task, output tokens used and speed, so you can see what a score costs, which vendor charts rarely show. Keep both in mind for the rest of this article, because the efficiency story only exists because of the second one.
The leaderboard: 58, and a five-point lead
Artificial Analysis published its evaluation on 22 September 2026 under the title Claude Opus 5.5 takes the top spot. Opus 5.5 at max effort scores 58 and ranks first of 211 models on the index, where the median model scores 26. Behind it, GPT-6 Astra and Claude Fable 5.1 are tied on 53, and Claude Opus 5, the model Opus 5.5 replaces, sits on 51 and ranks eleventh.
Five points on a composite index is a large gap. For comparison, the whole spread between Opus 5 and the joint second place is two points. The OfficeChai coverage framed it as a five-point lead over GPT-6 Astra, and that is the fair reading: it is a new top score, not a tie broken by noise.
Where the points come from
A composite can hide a single outlier benchmark, so the per-evaluation numbers matter more than the total. Artificial Analysis reports the following for Opus 5.5 at max effort:
- Humanity's Last Exam: 61.4%, against a previous best of 59.1%.
- SciCode: 66.9%, against a previous best of 63.1%.
- Terminal-Bench 4.0: 59.6%, level with GPT-6 Astra at xhigh effort and 11 points above Opus 5.
- AA-Briefcase: an Elo of 1822, 143 points above Claude Fable 5.1.
- GDPval-AA v2.1: 1,846, which is 111 above Fable 5.1 and 138 above Opus 5.
The pattern is broad rather than spiky. The largest gains are in knowledge work (AA-Briefcase and GDPval-AA, both built around realistic office tasks) and in terminal work, which is where coding agents spend their time. Terminal-Bench is the one place where it does not lead outright: GPT-6 Astra matches it. If your workload is an agent loop driving a shell, the honest summary is "joint best", not "best".
Medium effort is the real headline
Opus 5.5 exposes five effort settings: low, medium, high, xhigh and max. Effort controls how much the model is allowed to think before answering, and therefore how many output tokens you pay for. Artificial Analysis measured all five. The index scores are 42 at low, 51 at medium, 54 at high, 56 at xhigh and 58 at max.
Now put two of those numbers next to each other. Opus 5 at max effort scores 51 at $5.86 per task. Opus 5.5 at medium effort also scores 51, at $1.34 per task, ranking eighth on the whole leaderboard. Same score, about 77% cheaper per task, which is roughly a 4.4x difference. That is the sentence I would put on the launch slide, and it is not on the launch slide.
Artificial Analysis also notes that four of the five effort settings land on its intelligence versus cost per task frontier. In plain terms: for four of the five settings, no other measured model gives you a higher score for the same money. The setting ladder is not marketing; each step up buys real points, and you can choose where on the curve to sit.
| Model and setting | Index score | Rank | Cost per task | Output tokens, full index | Output speed |
|---|---|---|---|---|---|
| Claude Opus 5.5, max | 58 | 1 of 211 | $5.98 | 260M | 95.5 tokens/s |
| Claude Opus 5.5, medium | 51 | 8 | $1.34 | 38M | 81.7 tokens/s |
| Claude Opus 5, max | 51 | 11 | $5.86 | 140M | 60.5 tokens/s |
The catch: tokens and waiting time
Now the part the headline leaves out. Opus 5.5 at max effort is the most verbose model at the top of the table. Artificial Analysis counts about 119,000 output tokens per index task, against about 73,000 for Opus 5, about 78,000 for Fable 5.1 and about 27,000 for GPT-6 Astra. Over the full index it used 260 million output tokens, where the median model used 88 million. It stays level with Opus 5 on cost per task only because the price per token dropped, not because it thinks less.
The second cost is time. The max-effort model page reports a time to first token of 682.71 seconds, which is more than eleven minutes before the first token arrives, even though the output speed of 95.5 tokens per second is faster than Opus 5. At medium effort the same page reports 13.20 seconds. For a background job, eleven minutes is fine. For anything a user is watching, it is not a setting, it is an outage.
The price cut, and why caching matters more now
Opus 5.5 is priced at $4 per million input tokens and $20 per million output tokens, down 20% from Opus 5's $5 and $25, according to the Claude models overview and the Artificial Analysis model page. Cache reads dropped harder, by 60%, from $0.50 to $0.20 per million tokens, and writes to the five-minute cache went from $6.25 to $5. The context window is one million tokens.
The cache numbers are the ones to act on. A long-running agent resends its history on every turn, and a more verbose model makes that history grow faster. At $0.20 per million cached tokens, a stable prefix is close to free, and an uncached one is twenty times more expensive. If you have not already stabilized your prompt prefix, the prompt caching and routing guide explains how, and on Opus 5.5 the saving is larger than on any previous Claude model.
What I would actually do with this
The leaderboard tells you which model is strongest. It does not tell you which setting to run, and that is where the money is. This is the order I would work in:
- Make medium the default. It matches the previous flagship at a quarter of the cost and answers in seconds, not minutes. Most interactive features should start here.
- Reserve max for work nobody is waiting for: overnight analysis, large refactors run by a coding agent, eval generation. Budget the tokens explicitly, because 119,000 output tokens per task adds up.
- Escalate by rule, not by feel. Run medium first and move to high or max only when a check fails, with the escalation condition written down in advance. A small typed decision model, as in the Jev routing article, is a cheap way to make that call.
- Cache aggressively. At a 60% lower cache read price, prefix stability is now the biggest single saving on long agent sessions.
- Re-run your own evals before switching. Compare Opus 5.5 medium against whatever you run today on the same graded set; the evals guide shows how to build one in an afternoon.
The pattern behind all five is the same: the effort setting is now a product decision, not a model detail. The team that picks it per feature will get last generation's flagship quality at a fraction of last generation's bill. The team that leaves everything on max will pay for eleven-minute answers. If you want help making that decision for a real system, the AI engineering page describes how I approach it.
Sources
- Artificial Analysis: Claude Opus 5.5 takes the top spot on the Artificial Analysis Intelligence Index (22 September 2026)
- Artificial Analysis: Claude Opus 5.5 (max) model page
- Artificial Analysis: Claude Opus 5.5 (medium) model page
- Artificial Analysis: Claude Opus 5 (max) model page
- Artificial Analysis: model leaderboard and Intelligence Index
- OfficeChai: Claude Opus 5.5 creates a 5-point lead over GPT-6 Astra
- Claude API docs: Models overview, context windows and prices (as of September 2026)
Frequently asked questions
What is the Artificial Analysis Intelligence Index?
It is a composite score published by Artificial Analysis, an independent benchmarking company that runs the same evaluations on every major model. Version 4.3.2 combines ten evaluations, including Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDPval-AA and AA-Omniscience, and is published next to cost per task, output tokens and speed.
Which model is number one on the Artificial Analysis leaderboard?
As of 28 September 2026, Claude Opus 5.5 at max effort, with an Intelligence Index score of 58. GPT-6 Astra and Claude Fable 5.1 follow on 53 and Claude Opus 5 at max effort scores 51. Artificial Analysis published the result on 22 September 2026.
Is Claude Opus 5.5 more expensive to run than Opus 5?
At max effort it costs about the same per task, $5.98 against $5.86, because the lower token price offsets the roughly 1.6 times more output tokens it uses. At medium effort it matches Opus 5's max-effort score for $1.34 per task, so for the same quality it is much cheaper.
Which Opus 5.5 effort setting should I use in production?
Start with medium for interactive features: it scores 51 on the index and Artificial Analysis reports a time to first token of 13.20 seconds. Reserve high, xhigh and max for background work where more than ten minutes of latency is acceptable, escalate by a rule you define in advance, and confirm the choice on your own eval set.
How much did Opus 5.5 prices change?
Input and output fell 20%, from $5 and $25 to $4 and $20 per million tokens. Cache reads fell 60%, from $0.50 to $0.20 per million tokens, and writes to the five-minute cache from $6.25 to $5, which makes a stable, cached prompt prefix worth more than on any previous Claude model.