Anthropic has a new model on the block, and it comes with a score that sounds impressive, a price that sounds steep, and a name that asks to be read twice. Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) is the latest in Anthropic’s line of reasoning models, released on September 22, 2026, and it now has its first proper assessment. The verdict is simple: it is smart, it is verbose, and it costs more than most of its rivals.
The model holds a score of 58 on the Artificial Analysis Intelligence Index, putting it well above average among comparable models, where the median sits at 25. It generated 260 million tokens during evaluation, which is far more verbose than the median of 88 million. It costs $4 per 1 million input tokens and $20 per 1 million output tokens, which puts it well above the median of $2 for input and $10 for output.
The Intelligence Index Behind The Score
The Artificial Analysis Intelligence Index is the benchmark here, and it comes with its own title problem. “Artificial Analysis Intelligence Index” is a tautology wrapped in a double-barrelled name, which is either a stroke of branding genius or a slip of language. Either way, the number matters more than the label. The model’s 58 puts it well above a median of 25, which is a real margin over its peers.
The index does not say how it measures intelligence. That is a real gap in the reporting, and it is one worth keeping in mind rather than closing with a guess. What the index does measure, according to artificialanalysis.ai’s report, is performance across specific capabilities and industries. The components of the test are:
- Agentic knowledge work
- Agentic real-world work tasks
- Agentic SaaS workflows
- Agentic coding and terminal use
- Professional document reasoning
- Physics reasoning
- Quantitative analysis on spreadsheets and documents
- Kubernetes incident root-cause analysis
- Medical long-context reasoning
That is a wide spread of skills.
The model also holds a score on the AA-Briefcase v1.1, which the report describes as an “AA-Briefcase Elo,” though it does not explain what that means in terms of how the model performed there. The AA-Omniscience and AA-Omniscience Index are also part of the comparison set, and the report does not say what they test beyond naming them.
Price And Token Use
The price of Claude Opus 5.5 is the headline that will give pause to anyone who wants to run it. At $4 per 1 million input tokens and $20 per 1 million output tokens, it sits well above the median of $2 and $10. For a blended rate of 7:2:1 cache hit/input/output, the effective price works out to $2.94 per 1 million tokens. The report notes that pricing may vary by provider, which means the actual cost could shift depending on where a user happens to be buying access.
The token numbers are worth keeping in mind. The model scored 58 on the Intelligence Index and generated 260 million output tokens, which is at the higher end of what other reasoning models in its price tier produce. The median was 88 million. That level of output is not wasteful in the abstract — it could be a sign of thoroughness. But it is a heavy hit against the wallet, since the output tokens cost $20 per million.
The context window is 1 million tokens, which determines how much text and conversation history the model can process in a single request. That window is standard for a model of this size, and it is not the point of comparison here — the point is the intelligence the model delivers inside that window.
What Claude Opus 5.5 Actually Does
Anthropic describes the model as a reasoning model that uses extended thinking or chain-of-thought reasoning to work through complex problems before providing an answer. It supports text and image input, and it can process both in a single request. The model can analyze, describe, and answer questions about images, though the report does not give an example of what that looks like in practice. It can also generate text output.
The model is proprietary — Anthropic has not disclosed the model size or parameter count. It is not open weights, which means its architecture and training data are not available for inspection by outside researchers. That is a significant constraint for anyone who wants to test its limits or compare its behavior against rivals.
How It Compares To Other Models
The report compares Claude Opus 5.5 against a few groups of models, and the patterns are consistent across all of them. Claude Opus 5.5 is expensive by the benchmark’s measure, and it talks a lot more than its rivals.
The report groups models into several categories for comparison:
- Non-reasoning models, compared only with other non-reasoning models
- Reasoning models, compared across both reasoning and non-reasoning
- Open weights models, compared only with other open weights models of the same size class, broken down by parameter count (tiny: ≤4B, small: 4B–40B, medium: 40B–150B, large: >150B)
- Proprietary models, compared across proprietary and open weights models of the same price range, using a blended 3:1 input/output price ratio
Claude Opus 5.5 sits in the reasoning-model comparison, and its price is measured against the $0.15–$1 per 1 million tokens tier. The report does not say where the model lands on the larger proprietary model comparison, which uses a different blended price ratio.
The Business Of Being Expensive
The report’s own numbers tell a clear story about where Claude Opus 5.5 sits. It is smarter than its rivals by the Intelligence Index. It talks more than its rivals. And it costs more than most of them. That is a useful combination if the problem at hand is worth the money, but it is a warning sign if the question could be answered by a cheaper model.
The blended price of $2.94 per 1 million tokens, based on Anthropic’s API, is the number to hold onto. It is higher than the median across the board, and it applies to any workload that mixes input and output costs. The report notes that pricing may vary by provider, so the number could shift depending on where a user buys access. That is a real consideration for anyone building a workflow around the model.
What This Means For Users
The bottom line is simple. Claude Opus 5.5 is a model for people who need answers that come from reasoning through a problem, not just fetching a known fact. It is a model for users who want the most verbose, thorough response they can get, and who are willing to pay for it. The model’s score of 58 on the Intelligence Index suggests it delivers on that promise, but the cost of 260 million output tokens is a real burden for any workflow that runs it at scale.
The model is available through 5 API providers, which gives users some room to shop around for the best price. Anthropic’s own API is the source of the price figures, and the report compares providers for users who want to test the numbers against each other. That is a useful service, but it does not change the underlying math.
Claude Opus 5.5 is a model that scores well, talks a lot, and costs more than most of its rivals. If the problem is worth solving and the budget is not a concern, it is a model to consider. If the problem can be solved with a cheaper model, it is a model to avoid. The report has done the work of measuring the intelligence; the decision about which trade-off to make is up to the user.
Source material: “Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max),” artificialanalysis.ai.
Get the Notebook.
The day's best stories and every fresh verdict, in plain English, in your inbox by seven. One email a day, no more.

