studio raoulsoftware innovation, architecture & development
← Blog
AI in Practice · September 25, 2026

The Most Capable AI Just Got 40% Cheaper. Here Is What That Changes.

Anthropic's flagship Claude Opus 5.5 is now roughly 40% cheaper to run, with cache reads down 60% — and that changes the economics of running AI agents in production.

Key facts
$4
input price per million tokens
–40%
lower cost on typical workloads vs Opus 5
–60%
drop in cache-read price
66.4%
score on Terminal-Bench 4.0 coding test

Running the most capable AI on the market just got significantly cheaper. On 22 September, Anthropic released Claude Opus 5.5 — its most powerful model to date — at prices that work out to roughly 40% less on typical workloads compared to the version it replaces. Input tokens now cost $4 per million, down from $5. Output tokens fell from $25 to $20.

The steepest cut is on cache reads — down 60%, from $0.50 to $0.20 per million tokens. A cache read is what you pay when an AI re-reads the same document or the same instructions across multiple requests. Long agent tasks — scanning contracts, analysing a codebase, processing large reports — depend on this heavily. That 60% cut makes document-heavy AI workflows much cheaper to run at scale.

Anthropic says Opus 5.5 performs at the same level as their best research model on most tasks. On Terminal-Bench 4.0, a standard test for coding agents, it scores 66.4%. The lower price is not a trade-off for less capability — the model is cheaper and just as capable as before.

There is a real catch, though. Anthropic changed four things about how the API works that will break older Opus 5 code. The 40% savings headline also depends on how you set the model's effort level — a setting that changed with this release. Test on your actual workload before you migrate, and do not assume the headline number is what you will see.

If you have AI workflows that felt too expensive to run as often as you wanted, now is a good time to redo the numbers. The drop in cache reads matters most for anything that processes long documents repeatedly — legal review, code analysis, support pipelines. One clear action: estimate what last month's usage would have cost at the new prices. Then test a migration in staging before switching anything in production.

Sources
← Most companies don't know how many AI agents they're running← Nearly a Third of Every Mission Goes to Human Operators. One Startup Is Done With That.
The Most Capable AI Just Got 40% Cheaper. Here Is What That Changes.