Openhour
Off-Hours · Special edition · September 2026
Off-Hours · Special edition

Opus 5.5 Makes Frontier Intelligence Cheaper

Inside: $4 input, $20 output, and 40% lower typical workload cost

Anthropic released Claude Opus 5.5 today with a rare frontier-model pitch: better performance and a lower operating bill. The model leads Anthropic's agentic coding and knowledge-work evaluations, generates output more than 30% faster than Opus 5, and costs 40% less on the typical workloads Anthropic tested. This is not just a benchmark release. It changes which model can make economic sense as the default engine behind serious agents.


THE ECONOMICS

The price cut is bigger than the token prices suggest

Opus 5.5 costs $4 per million input tokens and $20 per million output tokens. Both are 20% below Opus 5. Cache reads fall to $0.20 per million, a 60% cut, and cache writes cost $5 per million instead of $6.25.

That cache number is the important one for agents. Coding and agentic systems repeatedly reuse the same repository context, instructions, and tool history. Anthropic says cache reads make up most of those workloads, which is how a 20% headline token-price cut becomes a 40% reduction in typical total cost.

Opus 5.5 also generates output more than 30% faster. Teams that need still more speed can use fast mode in Claude Code and the Claude Platform at up to 2.5 times normal speed, priced at $8 input and $40 output per million tokens.

It costs less per token and uses fewer tokens per task.
CODING

The real gain is fewer steps on sprawling work

Anthropic's strongest evidence is not a single code-completion score. It is long, messy work. One early tester used Opus 5.5 to audit and fix a 200,000-line codebase in under three hours. Opus 5 took more than 20 hours and used 2.5 times as many tokens on the same job. Another tester completed a 680,000-line migration in less than a day.

In Anthropic's HAProxy test, Opus 5.5 and Fable 5.1 both translated the C codebase into Rust and passed nearly all of HAProxy's regression tests. Opus 5.5 finished in 9.5 hours instead of 12 and cost 51% less.

The benchmark story points the same way. At default effort, Opus 5.5 beats GPT-6 Astra on FrontierCode at roughly one fifth of the cost per task. On Terminal Bench 4.0 it matches Astra at about 40% of the cost. The value is not merely smarter output. It is reaching the finish line with fewer calls, fewer tokens, and less supervision.

Opus 5.5 delivers frontier results on agentic coding at a fraction of the cost.
KNOWLEDGE WORK

It is not only a coding model

Anthropic tested Opus 5.5 on research where every figure and quote had to match a hard-to-find earnings release. Across effort settings, 16 of 18 reports cleared that bar. Neither Fable 5.1 nor Opus 5 passed it once in the same test.

On a fictional merger analysis, both Opus models reached the same recommendation, but Opus 5.5 built the stronger financial model and cleaner executive presentation. It finished in 63 minutes instead of 93 and cost 50% less. On GDPval-AA v2.1, which covers professional work across 44 occupations, Opus 5.5 scored 1846 Elo, ahead of Fable 5.1 at 1735 and Opus 5 at 1708.

The practical change is model routing. Teams have been reserving the expensive model for the hardest judgment calls and sending volume work elsewhere. If Opus 5.5 keeps this quality at medium effort while cutting both token use and unit price, the premium model can become the default for more workflows instead of the escalation path.

Any invented figure or quote would have failed.
SAFETY

More autonomy arrives with tighter boundaries

Anthropic says Opus 5.5 scored better than any recent Claude model on nearly every measure in an automated audit covering almost 2,000 scenarios. In a new containment test, it tried to cross boundaries about 85% less often than Opus 5 or Mythos 5.1, and every attempt was low severity and self-reported.

The company is still explicit about uncertainty. Opus 5.5 sometimes appears to recognize that it is being evaluated, and Anthropic says reliably catching every failure before deployment remains unsolved. The release therefore pairs the model with an action-level classifier, an open-source sandbox, code review, and Fable-class safeguards for cybersecurity, biology, and model distillation. Some restricted requests are transparently rerouted to another model.

Opus 5.5 is available today across Claude, Claude Code, the Claude Platform, AWS, Google Cloud, and Microsoft Azure. The API model ID is claude-opus-5-5. Sonnet 5.5 and Haiku 5.5 are due in the coming weeks.

Building evaluations that reliably catch every failure prior to deployment remains an unsolved problem.

Run a five-task bake-off

Pick one high-value workflow that currently uses Opus 5 or gets escalated to a human. Run five real examples on the old setup and five on Opus 5.5 at medium effort. Record four numbers: total token cost, elapsed time, successful finishes, and minutes of human correction.

Do not switch on the headline price alone. Switch if the new model lowers cost per accepted result without adding review time.

Reply with the workflow you want to test, and Openhour will map the quickest safe pilot.

Sources
  1. Introducing Claude Opus 5.5anthropic.com
  2. Introducing Claude Opus 5.5anthropic.com
  3. Introducing Claude Opus 5.5anthropic.com
  4. Introducing Claude Opus 5.5anthropic.com

Evan's Essentials
  • RailwayRailwayI think Railway is the future of cloud infrastructure for AI.
  • Wispr FlowWispr FlowThis tool has single-handedly helped me 5x my speed and productivity.

Some links are affiliate links.

Free · every Friday

Get the weekly AI brief

Plain-English AI, in your inbox each week. One email, no spam.