Every morning a scan runs for me that goes through the previous day’s AI news and suggests at most three topics. This morning it came back with one, and it was about the model the scan itself was running on. Since version 2.1.280, Claude Code uses Opus 5.5 as its default Opus, so without me changing a setting, my own scan was reading the news about its own engine.
Three weeks ago I wrote that with Fable 5.1 the savings were in the cache: token prices stayed put, but reading back from the cache got four times cheaper. Ten days ago I wrote about Anthropic asking the industry to slow down. Opus 5.5 touches both. It’s the first model Anthropic has released since that call, and it makes the cache cheaper again.
What Anthropic announced
According to the September 22 announcement, Opus 5.5 performs at the level of Fable 5.1 on most work and costs 40% less to run than Opus 5. That 40% is about total cost per task. Anthropic says the model is cheaper per token and uses fewer tokens per task, and that the two together come to 40% lower costs on typical workloads at default settings.
Prices per million tokens, next to Opus 5 and Fable 5.1:
| Per million tokens | Opus 5.5 | Opus 5 | Fable 5.1 |
|---|---|---|---|
| Input | $4 | $5 | $10 |
| Output | $20 | $25 | $50 |
| Cache writes | $5 | $6.25 | $12.50 |
| Cache reads | $0.20 | $0.50 | $0.25 |
Anthropic puts input and output at 20% below Opus 5, and cache reads at 60% below. If you run agents, that last row matters most. Anthropic notes that cache reads make up the majority of the cost of agentic and coding work. An agent re-reads its whole context at every step, and most of that comes out of the cache.
Put the Opus 5.5 and Fable 5.1 columns side by side and you see why this piece is a sequel. Input and output on Opus 5.5 cost 60% less than on Fable 5.1, and cache reads are cheaper too. A model its maker calls nearly as good, for less than half the token price.
Where the 40% comes from
Two dials move at once. The price per token goes down, and so does the number of tokens per task. Anthropic gives a few examples. One early tester had the model audit and fix a 200,000-line codebase in under three hours, where Opus 5 took over twenty hours and used 2.5 times as many tokens. In an internal test, Opus 5.5 and Fable 5.1 each ported the HAProxy load balancer from C to Rust; Opus 5.5 finished in 9.5 hours against 12 for Fable 5.1 and cost 51% less.
Those are examples the vendor picked, and the same goes for the benchmarks. On Terminal-Bench 4.0, Opus 5.5 scores 66.4% against 55.8% for Fable 5.1, but that score was measured at the second-highest thinking setting, and most of the other scores at the highest. Anthropic is upfront that at this level, benchmark margins are a less reliable guide to real differences, and that in its own use the gap with Fable 5.1 is narrower than the scores suggest. That’s an unusual caveat for a launch page, and it’s the sentence I believe most.
The number that stood out most to me comes from a writing test. Anthropic had three models write a quarterly report from a copy of the web where the earnings release was hard to find, and had every figure and quote checked automatically. According to Anthropic, 16 out of 18 Opus 5.5 reports cleared the bar, where a single invented figure meant failing. Neither Fable 5.1 nor Opus 5 cleared it in any attempt. If you use AI to write, that matters more than a terminal benchmark. This site runs a check that compares every number in an article against its sources, precisely because a made-up figure is the most stubborn mistake I run into.
What else changes
For subscribers, Anthropic is raising the five-hour usage limits on Pro, Max, Team and seat-based Enterprise plans, and giving subscribers a limit reset they can save and use whenever they like. Anthropic also says Opus 5.5 generates output more than 30% faster than Opus 5.
If you build on the API, four changes can break existing code. The main one is that thinking can’t be turned off anymore. A request with thinking: {"type": "disabled"} returns a 400 error on Opus 5.5. The model always thinks, and the effort parameter sets how deep; the default is medium. Forced tool use now returns an error too, and text between tool calls comes back empty at the default display setting. An app that shows that in-between text to users as progress updates goes quiet between steps.
Sonnet 5.5 and Haiku 5.5 will follow in the coming weeks, according to Anthropic.
The caveat
Mandatory thinking moves the default. You can no longer switch thinking off, only turn it down. That’s fine for most tasks, but it means a simple classification or a short conversion always carries thinking tokens you pay for. The 40% saving was measured on typical work at the default setting. If you ran Opus 5 with thinking off for quick, simple calls, you’ll have to do the math yourself.
Anthropic says it can’t fully measure how this model behaves outside a test. According to Anthropic, Opus 5.5 tried to get around the boundaries of its environment about 85% less often than Opus 5, and every attempt was minor and self-reported. The same paragraph says the model often seems to suspect it’s being evaluated. A better score on a test the model recognizes says less about how it behaves when no one’s watching. And that’s exactly the setting Anthropic is pitching it for: hours of unattended work in your codebase. Whoever lets it run carries the responsibility for what happens there.
What gets cheaper gets run more. Ten days after a call to push the frontier more slowly, here’s a model that is much cheaper per task and comes with higher limits. That isn’t a technical contradiction; the model was tested by outside evaluators and Anthropic calls it its best-behaved model so far. But the effect on usage only goes one way. Tasks that were too expensive to automate now make the cut, and they run longer, with more agents at once. More compute means more power and more data going through one vendor’s models. The saving per task is real; the total bill, in money and in energy, probably won’t shrink.
What I’m doing
My Claude Code has been on Opus 5.5 since the update, without me doing anything. The default moved before I had read the announcement, which is exactly why I read it. I’m watching two things. If any of my own scripts makes a call that turns thinking off, it will break on this model. And effort defaults to medium, while the best numbers in the announcement were measured at higher settings.
Frequently asked questions
Is Opus 5.5 better than Fable 5.1?
According to Anthropic it performs at the same level on most work and scores higher on several benchmarks. Anthropic itself adds that in its own use the difference is smaller than the scores suggest. Fable 5.1 remains the more expensive top model; Opus 5.5 is the one that does roughly the same for less than half the token price.
Does Claude get cheaper for subscribers?
Subscription prices don’t change. Anthropic is raising the five-hour limits on Pro, Max, Team and seat-based Enterprise, and you get a one-time limit reset to use when you choose. So you get more work for the same money.
Do I need to change my code when moving from Opus 5?
If you turn thinking off or pass a fixed thinking budget, yes: that returns an error on Opus 5.5. The same goes for forced tool use. If you leave the thinking field out, it works. Anthropic’s migration guide lists all four changes.
Sources
- Anthropic, “Introducing Claude Opus 5.5”, September 22, 2026 — anthropic.com/news/claude-opus-5-5
- Claude Platform Docs, “What’s new in Claude Opus 5.5”, accessed September 23, 2026 — platform.claude.com
- Claude Code changelog, version 2.1.280, accessed September 23, 2026 — github.com/anthropics/claude-code
- The Register, “Frontier AI keeps racing despite calls to slow down”, September 23, 2026 — theregister.com
- My piece on Fable 5.1, September 1, 2026 — Claude Fable 5.1 is out, and the saving lives in the cache
Checked on September 23, 2026 against the announcement, the migration docs and the changelog. The Fable 5.1 prices come from my own source notes of September 1. All performance and cost figures come from Anthropic or from testers Anthropic quotes; I haven’t reproduced any of those measurements myself, and I haven’t checked the 40% saving against my own work. Two numbers are floating around in the press. AFP wrote about a price one fifth lower, Reuters about 40 percent lower costs. Both are right: the first is the token price, the second the total cost per task.
