All articles
LLM API Comparison & Cost/

Claude Opus 5.5: 40% Lower Compute Costs, Comparable Performance

Anthropic's latest model cuts inference costs to $4/$20 per million tokens while matching Fable 5.1 performance—with 1M context and 30% faster output.

5 minutes read

Claude Opus 5.5: 40% Lower Compute Costs, Comparable Performance. A single high-density GPU server module positioned next to stacked cooling pipes and power distribution units, with LED indicators showing reduced thermal output and power — GPUniq
Data as of 3 sources

Drafted with AI tools and edited by a human. The figures and conclusions were checked by a GPUniq editor.

TL;DR

Claude Opus 5.5 costs $4 per million input tokens and $20 per million output tokens, with Anthropic citing roughly 40% lower compute and running costs compared to Opus 5. It matches Claude Fable 5.1 performance on most tasks, generates output about 30% faster, and keeps the 1M-token context window intact.

Pricing breakdown

Five line items matter here.

ItemPrice, $/M tokens
Input tokens$4.00
Output tokens$20.00
Cache reads$0.20
5-minute cache writes$5.00
1-hour cache writes$8.00

Cache reads at $0.20/M are 20x cheaper than standard input. That gap is where real savings accumulate in multi-turn or RAG workloads. The 5-minute window costs $5.00/M to write; if your session runs longer, the 1-hour retention tier at $8.00/M is the alternative. Neither is free, so do the arithmetic: write cost only pays off if the same cached prefix gets read enough times to offset it.

Performance vs. alternatives

According to Anthropic's official announcement, Opus 5.5 delivers comparable performance to Claude Fable 5.1 on most tasks. That's a meaningful claim: a model with roughly 40% lower compute and running costs than its predecessor matching a separate model line entirely.

Output generation is about 30% faster than Opus 5, per the same source. Fewer tokens consumed per task is also cited as a capability improvement.

On Anthropic's Automated Behavioral Audit, Opus 5.5 scores highest among all tested Claude models to date. That's the alignment and safety suite, not a general reasoning benchmark, so weight it accordingly.

DimensionOpus 5.5Opus 5
Input price, $/M$4.00-
Output price, $/M$20.00-
Context window1M tokens-
Max output per request128K tokens-
Output speed~30% fasterbaseline
Behavioral Audithighest tested-

Max output is 128K tokens per request, or up to 300K via Batch API beta for outputs, according to Anthropic.

Availability

Released September 22, 2026. The model ID is claude-opus-5-5.

Available on:

  • Claude API (all customers)
  • Amazon Bedrock as anthropic.claude-opus-5-5
  • Google Cloud as claude-opus-5-5
  • Microsoft Foundry as claude-opus-5-5
  • Claude.ai on Pro, Max, Team, and Enterprise plans

If a specific platform restricts a feature, check that platform's documentation directly.

Adaptive thinking is enabled by default and cannot be turned off. You can control it via the effort setting, but you cannot disable it entirely.

Should I migrate from Opus 5?

For high-volume inference, the math is straightforward. For high-volume inference, Anthropic cites roughly 40% lower compute and running costs versus Opus 5, which can translate to significant savings at scale before any caching. Add prompt caching on a typical RAG pipeline and you can cut further.

No breaking changes to the core API contract. The model ID changes and that's mostly it for standard text workloads. Adaptive thinking always being on is the main behavioral shift: simple queries that previously returned quickly may consume more tokens for internal reasoning. Budget for that.

I'd start with a shadow-test on your existing prompt suite before cutting over production traffic. Edge cases in structured output or tool use are where surprises tend to appear.

What breaks when upgrading?

Three things to check before you flip the switch.

First, the computer_20251124 tool is gone. If you're using computer use, update to a supported tool version per Anthropic's documentation. Calls using the old tool ID will fail.

Second, adaptive thinking cannot be disabled. If your application depends on predictable low-latency responses for simple queries, test actual latency under the new model. The reasoning overhead is real.

Third, cache write costs may differ from prior beta rates. If you built cost models around earlier beta pricing, recalculate against the table above.

Thinking blocks are tied to the model and conversation context and cannot be edited. Knowledge cutoff is June 2026.

# Old tool call - will fail on Opus 5.5
tool = {"type": "computer_20251124", ...}

# Use this instead
tool = {"type": "computer_20250102", ...}

How caching saves money

Cache reads at $0.20/M versus standard input at $4.00/M is a 20x difference. For any prompt with a large static prefix, that gap compounds fast.

Concrete example using the pricing table: a 50K-token system prompt read 100 times costs $20.00 at standard input rates (50,000 x 100 x $4/M). With cache reads it costs $1.00 (50,000 x 100 x $0.20/M). The cache write is a one-time $0.25 (50,000 x $5/M). Net saving on that session: $18.75.

The 5-minute window covers most interactive sessions. For batch jobs running longer, the 1-hour retention tier at $8.00/M write cost is worth it if the same prefix appears across many requests in that window.

Multi-turn conversations benefit most. Each turn that reuses accumulated context avoids re-billing the full history at input rates.

Cloud platform support

Multiple platforms are confirmed available, per Anthropic's announcement.

PlatformModel IDNotes
Claude APIclaude-opus-5-5All customers
Amazon Bedrockanthropic.claude-opus-5-5On-demand inference
Google Cloudclaude-opus-5-5-
Microsoft Foundryclaude-opus-5-5-

For provisioned throughput pricing on Bedrock or committed-use discounts on Vertex, check those platforms directly.

Estimated cost per task

Using the published rates and straightforward token arithmetic:

TaskApprox. input tokensApprox. output tokensEstimated cost, $
Simple query1,0005000.014
RAG with cached prompt50K cached + 1K fresh5000.024
Long-context analysis100,00010,0000.60
Batch job, per request5,0002,0000.060

The RAG row assumes the 50K-token prompt is already cached (read at $0.20/M) and only the 1K fresh tokens are billed at $4.00/M. These numbers are illustrative; your token counts will vary.

Methodology

Facts are compiled from the vendor's own announcement and documentation, linked below, at the time of writing; where the vendor hasn't published a number it is marked as such. GPUniq prices, when the model is already available here, come from our live catalog.

Sources

  1. 1.Claude Opus 5.5 – Anthropic announcement — Anthropic (accessed )
  2. 2.Claude Opus 5.5 – Claude Platform Docs overview — Anthropic (accessed )
  3. 3.Claude Opus 5.5 – Amazon Bedrock model card — Amazon Web Services (accessed )

Want cheap GPUs for your next project?

Browse live GPU prices and rent the right card in seconds — H100, A100, RTX 4090, and 50+ more models.

More from the GPUniq blog