On Tuesday, 22 September 2026, at around half past six in the evening German time, Anthropic released Claude Opus 5.5. The promise sits right in the first paragraph of the announcement: "It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5." That comes only two months after Opus 5, released on 24 July. And according to Anthropic it is the company's first release since it called for pacing the frontier, a call made by CEO Dario Amodei the week before.
Whether the headline holds depends on a detail that almost no news report mentions: the effort level at which the model was measured.
As always, I keep three classes of statement apart: vendor-reported figures from the announcement, the documentation and the system card, independent measurements with source, version and date, and my own assessment, which I label as such. Own calculations from published figures are labelled too. The video version of this analysis sits at the top of this page and on the video page. Which AI models I actually use in my daily work is documented with date and source on the AI in numbers page. All figures in this article are as of 23 September 2026.
The facts from the announcement and documentation
First the verifiable part, all of it according to Anthropic's announcement, documentation and pricing page.
| Feature | Claude Opus 5.5 |
|---|---|
| Released | 22 September 2026 |
| Model ID | claude-opus-5-5 |
| Context window | 1 million tokens |
| Maximum output | 128,000 tokens, up to 300,000 via the Batch API as a beta |
| Knowledge cutoff | June 2026 |
| Effort levels | low, medium, high, xhigh, max |
| Default effort | medium (Opus 5 used high) |
| Platforms | Claude API, Amazon Web Services, Google Cloud, Microsoft Azure |
| In claude.ai | Pro, Max, Team and Enterprise, not on the free plan |
| Claude Code | default model for Pro, Max, Team, Enterprise and the API, from version 2.1.280 |
The most important row is the default effort level. The documentation says it explicitly: "A request that omits effort runs at medium; on Claude Opus 5 it ran at high." That sounds like a footnote, but it matters again for Anthropic's table, for the 40 percent and for the real costs.
Opus 5.5 is the first model of the Claude 5.5 family, and according to Anthropic Sonnet 5.5 and Haiku 5.5 will follow in the coming weeks. Notebookcheck checked the app and describes the difference to the last launch like this: "Pro subscribers get it without paying extra, which is not how Fable 5.1 arrived." In Claude Code, according to the documentation, Opus 5.5 replaces Sonnet 5 as the default model for Pro users. What that means for the actual usage allowance is not documented.
Anthropic also announces higher five-hour limits and a rate limit reset that subscribers can save for later, but gives neither a percentage nor an expiry date. I have not found the figures circulating in press reports in any Anthropic source, so they are not in this article. Opus 5 remains available: it is now listed under legacy models, but according to Anthropic's deprecation list it will not be retired before 24 July 2027.
Anthropic's table: nine wins over Fable, two losses to Astra
The following figures are vendor-reported, measured at the highest effort level, max (Terminal-Bench 4.0: xhigh). One trap up front: the last column is GPT-5.6 Sol, not the GPT-6 Sol that OpenAI presented on the same day.
| Benchmark | Opus 5.5 | Fable 5.1 | Opus 5 | GPT-6 Astra | GPT-5.6 Sol |
|---|---|---|---|---|---|
| Terminal-Bench 4.0 | 66.4 | 55.8 | 52.3 | 57.9 | 37.3 |
| FrontierCode v1.1 (Main) | 54.4 | 50.3 | 48.0 | 53.3 | 47.5 |
| CursorBench 4.0 | 57.8 | 51.8 | 46.6 | no value | 41.7 |
| GDPval-AA v2.1 (Elo) | 1846 | 1735 | 1708 | 1542 | 1588 |
| AutomationBench | 40.0 | 31.4 | 26.9 | 41.4 | 28.8 |
| Humanity's Last Exam, with tools | 67.7 | 65.6 | 63.6 | 57.2 | no value |
| Terminal-Bench-Science 0.1 | 58.7 | 52.6 | 29.0 | 64.6 | 22.4 |
| OSWorld 2.0 (partial) | 81.8 | 80.7 | 74.0 | no value | no value |
| Chartography, with tools | 89.0 | 88.4 | 83.4 | no value | no value |
Table scrolls sideways
A quick word on the tests: Terminal-Bench 4.0 measures agentic coding in the terminal, Terminal-Bench-Science research work in the terminal. According to Anthropic, GDPval-AA v2.1 evaluates agents on real professional work across 44 occupations, reported as an Elo score, a relative strength in head-to-head comparison. AutomationBench tests automation tasks and was run by Zapier.
Against Fable 5.1 and Opus 5, Opus 5.5 leads in all nine rows, on Terminal-Bench 4.0 for example with 66.4 percent against 55.8 and 52.3. Against GPT-6 Astra it loses two rows: Terminal-Bench-Science at 58.7 against 64.6 and AutomationBench at 40.0 against 41.4. According to the footnote, the Astra figures for both Terminal-Bench variants are as reported by OpenAI. How Astra performs otherwise is covered in my GPT-6 Astra review.
The most honest sentence about this table comes from Anthropic itself: "at these levels of capability we've found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest."
If you know my analysis of Claude Opus 5 from July: the Opus 5 figures are different today. For GDPval-AA, Anthropic reported 1861 Elo in version 2 in July, now it is 1708 in version 2.1. For FrontierCode, 53.4 was the top score back then; today Opus 5 carries that value at medium and only 48.0 at max. Anthropic gives no explanation for this.
The fine print: max, xhigh and a fallback
Below the table sit two sentences that say more than any row above them. The first: "Unless otherwise noted, all Claude Opus 5.5 results use adaptive thinking at max effort." Terminal-Bench 4.0 ran at xhigh because, according to Anthropic, that is where the model scored highest. So the table shows the two highest levels. What ships is medium.
The second concerns the safety filters: "When they intervened, cybersecurity tasks were completed by Claude Opus 4.8, and biology and frontier LLM development tasks were completed by Claude Opus 5. This likely reduces Claude Opus 5.5's performance on these benchmarks." So part of the results come from older models. How many tasks were affected is not stated anywhere. AutomationBench worked the other way round: Zapier measured without the fallback, and every filter intervention counted as a failure.
Anthropic's per-level charts, on the other hand, are strong. For Terminal-Bench 4.0 the company publishes score and cost per task, all vendor-reported:
| Effort level | Opus 5.5 | Cost per task |
|---|---|---|
| low | 38.5 % | $1.29 |
| medium (default) | 57.6 % | $2.94 |
| high | 64.2 % | $3.88 |
| xhigh | 66.4 % | $7.35 |
| max | 64.8 % | $11.24 |
| For comparison: Opus 5 at max | 52.3 % | $15.83 |
| For comparison: GPT-6 Astra at high | 57.9 % | $7.21 |
Table scrolls sideways
Anthropic's caption: "Opus 5.5 at default effort beats Opus 5 at max effort for about a fifth of the cost. It matches GPT-6 Astra at about 40% of the cost." Recalculated, that is just under 19 and just over 40 percent (own calculation). Notable: at max, Opus 5.5 scores lower than at xhigh but costs more than half as much again.
Two details do not make it into the headline. On GDPval-AA, Opus 5.5 at the lowest level, low, scores 1224 Elo, below Opus 5 at 1294 and Fable 5.1 at 1450. And on AutomationBench Anthropic writes "Opus 5.5 outscores Opus 5 and GPT-5.6 Sol at every effort level", while in the same chart GPT-6 Astra leads at every level. The sentence is true, it just leaves out the strongest competitor.
The price: the first Opus price cut since Opus 4.5
List prices per million tokens according to Anthropic's announcement and pricing page:
| Item | Opus 5.5 | Opus 5 | Fable 5.1 | Sonnet 5 |
|---|---|---|---|---|
| Input | $4 | $5 | $10 | $2 |
| Output | $20 | $25 | $50 | $10 |
| Cache write, 5 minutes | $5 | $6.25 | $12.50 | $2.50 |
| Cache read | $0.20 | $0.50 | $0.25 | $0.20 |
Table scrolls sideways
Input and output drop by 20 percent. According to Simon Willison, Opus 4.5, 4.6, 4.7, 4.8 and 5 all shared the same price. That makes Opus 5.5 the first Opus price cut since Opus 4.5, which is my conclusion from his list.
The cache read falls even further, from 50 to 20 cents, down 60 percent. Artificial Analysis calls this a 95 percent discount on uncached input, up from 90 on previous Opus models. That puts the cache read on par with Sonnet 5 and below Fable 5.1. According to Anthropic, cache reads make up the majority of costs in agentic and coding work. Why that matters is something I described in my Claude Fable 5.1 check.
Against Fable 5.1 at $10 and $50, Opus 5.5 costs 40 percent of the token price. Via batch processing it costs $2 and $10. Fast mode costs twice as much, $8 and $40, delivers up to 2.5 times the output tokens per second according to Anthropic and runs only on the Claude API.
The 40 percent question
Anthropic writes: "Our tests show that at default settings it will cost 40% less than Opus 5 on typical workloads." And a few lines further down: "Opus 5.5 also generates output more than 30% faster than Opus 5."
There is no methodology footnote for either figure: no usage mix, no time period, no definition of "typical workloads". For Fable 5.1, Anthropic had still backed its stated savings with a usage mix from August 2026. My interpretation, not a source: "at default settings" most likely means medium against Opus 5's old default, high. The 40 percent would then mix the token price, fewer tokens per task and a lower effort level. Artificial Analysis, incidentally, leads with what is hard fact: "a 20% price cut".
The customer quotes in the announcement, such as Optiver with 40 to 50 percent lower costs for a specific workload, are vendor material. There is an independent cross-check all the same.
Independently measured: first place at Artificial Analysis
Artificial Analysis is an independent benchmark provider that tests models itself. The Intelligence Index v4.3.2 combines ten evaluations, including Terminal-Bench 4.0, Humanity's Last Exam, SciCode and GDPval-AA v2.1. Opus 5.5 was measured at all five levels, explicitly "with Anthropic's default fallback enabled".
| Effort level | Intelligence Index v4.3.2 | Rank | Speed (tokens/s) | Cost per task | Output tokens per task |
|---|---|---|---|---|---|
| max | 57.62 | 1 of 212 | no data | $5.98 | 119,166 |
| xhigh | 55.99 | 2 | 92.4 | $3.46 | 65,667 |
| high | 53.58 | 3 | 90.2 | $1.82 | 35,584 |
| medium (default) | 51.24 | 8 | 75.2 | $1.34 | 25,745 |
| low | 42.31 | 37 | 78.6 | $0.55 | 10,150 |
Table scrolls sideways
| Model (level) | Intelligence Index v4.3.2 | Speed (tokens/s) | Cost per task |
|---|---|---|---|
| Claude Fable 5.1 (max) | 53.35 | 65.8 | $7.63 |
| GPT-6 Astra (max) | 52.67 | 52.5 | $3.26 |
| Claude Opus 5 (max) | 50.78 | 53.6 | $5.86 |
| Claude Opus 5 (high, old default) | 48.12 | 54.1 | $3.61 |
| GPT-6 Sol (max) | 47.53 | 122.2 | $1.06 |
Table scrolls sideways
At max, Opus 5.5 ranks first of 212 models, according to Artificial Analysis "the highest score we have measured by several points". It leads six of the ten evaluations, for example Humanity's Last Exam at 61.4 percent and SciCode at 66.9 percent, ahead of Fable 5.1 at 59.1 and 63.1 respectively. It trails on CritPt, AA-LCR and GDP.pdf. My most important finding, own calculation: already at high, Opus 5.5 scores 53.58 points, above Fable 5.1 at max, and costs $1.82 instead of $7.63 per task, just under a quarter.
On Terminal-Bench 4.0, Artificial Analysis measures 59.6 instead of 66.4 percent, level with GPT-6 Astra at xhigh. Opus 5.5 is not yet on the official leaderboard, whose data dates from 21 September. On AutomationBench-AA, a separate variant, it scores 69.5 against Astra's 68.5 at max, which Artificial Analysis calls parity. So AutomationBench is not a win for Opus 5.5 in either measurement.
One weakness remains the hallucination rate on AA-Omniscience, where lower is better. Opus 5.5 comes in at 58.6 percent at max and 68.4 at medium. GPT-6 Astra sits at 51.3, Opus 5 at 60.8 and Fable 5.1 at 72.6 percent. What this rate measures, and what it does not, I explain in my review of Opus 5 after five days.
Two notes belong here. In July Opus 5 stood at 61 points in an older index version, today it sits at 50.78 in v4.3.2, and the two numbers are not comparable. And at the Fable 5.1 launch Artificial Analysis wrote that it had supported Anthropic with pre-release evaluation. I found no such sentence for Opus 5.5, so whether there was early access is open. Independent here means: its own measurement, not a neutral referee.
What a task really costs
Artificial Analysis also measures what a run costs per task. That makes it possible to cross-check the 40 percent claim. The percentages are my own calculation.
| Comparison | Opus 5.5 | Opus 5 | Difference (own calculation) |
|---|---|---|---|
| Intelligence Index, default against default (medium against high) | $1.34 | $3.61 | about 63 % cheaper |
| Intelligence Index, max against max | $5.98 | $5.86 | about 2 % more expensive |
| Coding Agent Index v1.5, Claude Code at max | $13.04 | $10.79 | about 21 % more expensive |
Table scrolls sideways
Default against default, the claim holds, and with a higher index score at that, 51.24 against 48.12 points. Speed fits too: 75.2 instead of 54.1 tokens per second, roughly 39 percent more (own calculation). At max against max the savings are gone, with clearly more performance, 57.62 against 50.78 points.
It becomes even clearer in the Coding Agent Index v1.5. It combines DeepSWE v1.1, Terminal-Bench 4.0 and SWE-Atlas-QnA in equal parts and runs the models inside their real agents, such as Claude Code or Codex.
| Agent and model | Coding Agent Index v1.5 | Cost per task | Time per task |
|---|---|---|---|
| Claude Code with Opus 5.5 (max) | 66.0 | $13.04 | 3,867 seconds |
| Claude Code with Fable 5.1 (max) | 62.2 | $12.39 | 2,090 seconds |
| Codex with GPT-6 Astra (max) | 61.6 | $7.47 | 1,762 seconds |
| Claude Code with Opus 5 (max) | 59.7 | $10.79 | 2,516 seconds |
Table scrolls sideways
Opus 5.5 ranks first, but at max it needs about 21 percent more money and about 54 percent more time than Opus 5, a little over 64 instead of just under 42 minutes per task. Compared with Codex and GPT-6 Astra it costs about 75 percent more and takes more than twice as long (own calculations). So the 40 percent applies where Anthropic places it: at default settings.
The max trap
Artificial Analysis shows directly why max gets so expensive: about 119,000 output tokens per index task, against about 73,000 for Opus 5, about 78,000 for Fable 5.1 and about 27,000 for GPT-6 Astra. At 1.6 times Opus 5, that eats up the lower token price. Anthropic warns about this itself in the documentation: "At the same effort setting the model tends to think more per turn than Claude Opus 5, most of all at xhigh and max… leave room in max_tokens for the thinking."
Simon Willison ran into it on launch day: "Opus 5.5 has a 128,000 maximum output token limit (as do the other Claude models), and it hit that while it was still reasoning about the SVG!" Twice no answer came back, each time after nearly 20 minutes and for $2.56, exactly 128,000 tokens at $20 per million (own calculation). Fable 5.1 did not have the problem in his test. His conclusion: "This makes me suspect that "max" is effectively useless".
The fallback: when an older model answers
New for Opus are safety filters of the kind Fable already had: "Opus 5.5 is the first Opus model to launch with a similar class of safeguards to Fable 5.1 on cybersecurity, biology, and distillation, all of which fall back to another model transparently." Cyber requests are then answered by Opus 4.8, biology requests by Opus 5. In Claude Code, according to the documentation, a biology hit switches the entire session to Opus 5.
The word "transparent" is contested. The system card says: "All of our blocking safeguards operate with transparent blocks and do not covertly change model responses." The English edition of heise online writes: "Requests are then automatically redirected to a weaker, older model without the user actively noticing." Notebookcheck: "Anthropic calls the fallback transparent. What it means for you is that some answers arrive from a weaker model with no indication in the chat window." According to the documentation, Claude Code shows a notice in the transcript; for the chat app it is one statement against another.
How often this happens, Anthropic does not say. In the Coding Agent Index, Opus 5.5 was rerouted to another model in 80 of 909 attempts. Artificial Analysis reports 9.1 percent, the mean of the three individual tests, so roughly one attempt in eleven. Fable 5.1 comes in at 8.8 percent, Opus 5 at 0.5 percent.
The system card: 230 pages, four passages
The system card dated 22 September runs to 230 pages, and everything in it is vendor-reported. Four passages stand out.
Cyber. Page 47: "Claude Opus 5.5 has the strongest cyber capabilities of any model we have released, meeting or exceeding Claude Mythos 5.1 on all of our internal evaluations." Without production safeguards, according to page 3, it assisted with dual-use and benign security tasks at the highest rate of all models evaluated, but also refused malicious requests at the lowest rate. Hence the filters.
Evaluation awareness. According to page 93, the model suspects "in as many as 36% of transcripts" that it is being tested. In real internal use in Claude Code the figure is 0.4 percent. My take: a model that recognises tests makes safety evaluations harder to interpret.
Sandbox. Without safeguards, Opus 5.5 attempted to escape or tamper with its sandbox in 1.5 percent of runs, across 28 scenarios with seven runs each and all of low severity. Given apparent credentials to a public package registry, it took potentially harmful actions in roughly half of cases. The roughly 85 percent fewer circumvention attempts than Opus 5 or Mythos 5.1 that Anthropic cites in the announcement appear in the system card only as a chart.
Pasted text. On prompt injection, according to Anthropic, Opus 5.5 ties Fable 5.1 for the lowest success rate of any model tested. There is a regression, though, for text that users paste into the prompt themselves: the model is "more likely than previous models to follow malicious instructions in text that a user pastes into their own prompt". At the default level it followed such instructions in about 2 percent of cases, at max in 7.4 percent, and with product protection active in none, according to Anthropic.
External testers included METR and Frontier Design. After ten working days of API access, METR writes: "We believe that acceleration from this model would be slightly higher than for Fable 5.1, but that this model is unlikely to be able to fully automate AI R&D."
The same evening: GPT-6 Sol
Shortly after Anthropic, OpenAI presented GPT-6 Sol and GPT-6 Luna on the same day. Sources give gaps ranging from a few minutes to 90 minutes; going by the report timestamps it was the same evening German time. I could not retrieve OpenAI's announcement because the server refused access. All OpenAI figures are therefore second-hand, from Simon Willison, Decrypt, TechCrunch and 9to5Google.
| Model | Input | Cache read | Output |
|---|---|---|---|
| Claude Opus 5.5 | $4 | $0.20 | $20 |
| GPT-6 Sol | $2 | $0.20 | $10 |
| GPT-5.6 Sol (previous) | $4 | $0.40 | $20 |
| GPT-6 Astra | $10 | $1 | $50 |
Table scrolls sideways
According to 9to5Google, OpenAI explains the price by "reducing API prices for Sol and Luna by 50% compared with their GPT-5.6 promotional pricing". Simon Willison sums up the twist: "The new price for Opus 5.5 is the same as the price for GPT-5.6 Sol, but that was before OpenAI dropped their Sol prices by half." Per token, Opus 5.5 now costs twice as much as GPT-6 Sol, while the cache read is the same for both.
Artificial Analysis measures GPT-6 Sol at max at 47.53 points, $1.06 per task and a much faster 122.2 tokens per second. Opus 5.5 at max scores about ten points higher but costs about 5.6 times as much per task (own calculation). On the new OpenAI models Artificial Analysis writes: "GPT-6 Sol and Luna push the cost efficiency frontier by halving cost relative to GPT-5.6 Sol and Luna. Intelligence Index and Coding Agent Index scores remain level with GPT-5.6, with progress in some evaluations and regressions in others". I find another comparison more revealing, again my own calculation: Opus 5.5 at medium costs $1.34 per task, about a quarter more than Sol at max, and scores almost four points higher.
What this means for you
What follows is my assessment based on the figures documented above, not an external fact.
Access. You get Opus 5.5 from the Pro plan upwards at no extra cost, and in the API it is cheaper than Opus 5. In Claude Code it is the default from version 2.1.280.
Effort level. Leave it at medium, use high for hard tasks and max only when it truly has to be, and then with enough headroom in the output limit. The data above shows the same pattern three times: more tokens, more cost, not necessarily more quality.
Writing. For marketing work the most interesting promise is not a benchmark score but the writing style: "It puts the most important information up front, is less likely to use jargon or idiosyncratic phrases, and follows the writing rules you give it." That targets the criticism of Opus 5 I collected in my review after five days. The new style has not been independently measured yet. From my own practice the rule stays the same: check factual claims before they go into client copy. And according to the system card, third-party text that you paste into the prompt is a possible attack route.
API. Anthropic names four changes that can break existing code written for Opus 5:
- Thinking can no longer be switched off.
- Forced tool use returns an error.
- Thinking blocks are bound to the model and the conversation.
- The older computer use tool
computer_20251124is no longer accepted on the Claude API or on Google Cloud, while Amazon Bedrock still accepts it.
On top of that comes a silent change: text between tool calls comes back in thinking blocks whose text is empty at the default setting. This is not a pure drop-in update.
If you want to sort out which model at which level fits your tasks, your data protection requirements and your budget, that is exactly the subject of my AI consulting.
Verdict: Fable level at an Opus price?
In summary, as of 23 September 2026:
- In the independent Artificial Analysis Intelligence Index, Opus 5.5 at max ranks first, measured with the fallback active.
- The token price drops by 20 percent, the cache read by 60 percent. That is hard fact.
- The 40 percent savings are a vendor claim without a methodology footnote. Measured independently, they hold at default settings, not at max.
- On sensitive topics an older model may answer.
So the answer to the title question depends on the effort level. At high, Opus 5.5 reaches Fable 5.1 level at Artificial Analysis for just under a quarter of the cost per task. At max it delivers more, but then costs as much as Opus 5. How the field shifted before this is shown in my checks of Claude Fable 5.1 and GPT-6 Astra. And when someone shows you a benchmark score: ask about the effort level.
Sources
- Anthropic, announcement "Claude Opus 5.5" with benchmark table, effort charts and customer quotes, 22 September 2026: anthropic.com
- Claude platform docs, model page for Claude Opus 5.5: platform.claude.com
- Claude platform docs, "What's new in Claude Opus 5.5": platform.claude.com
- Claude platform docs, pricing overview, retrieved 23 September 2026: platform.claude.com
- Claude platform docs, model deprecations: platform.claude.com
- Claude Code docs, model configuration, retrieved 23 September 2026 (code.claude.com)
- claude.com, plans and model availability, retrieved 23 September 2026: claude.com
- Anthropic, system card for Claude Opus 5.5, 230 pages, 22 September 2026: anthropic.com
- Artificial Analysis, model page for Claude Opus 5.5, Intelligence Index v4.3.2, retrieved 23 September 2026: artificialanalysis.ai
- Artificial Analysis, launch article, 22 September 2026: artificialanalysis.ai
- Artificial Analysis, Coding Agent Index v1.5, retrieved 23 September 2026: artificialanalysis.ai
- Terminal-Bench 4.0, public leaderboard, data as of 21 September 2026: tbench.ai
- Simon Willison on Opus 5.5, GPT-6 Sol and Luna, 22 September 2026: simonwillison.net
- heise online, English edition, 22 September 2026: heise.de
- Notebookcheck on availability and fallback, 23 September 2026: notebookcheck.net
- TechCrunch on the Opus 5.5 launch, 22 September 2026: techcrunch.com
- TechCrunch on GPT-6 Sol and Luna, 22 September 2026: techcrunch.com
- Decrypt on GPT-6 Sol, GPT-6 Luna and Opus 5.5, 22 September 2026: decrypt.co
- 9to5Google on both launches, 22 September 2026: 9to5google.com
Figures as of: 23 September 2026. Claude is a trademark of Anthropic PBC, GPT is a trademark of OpenAI. Editorial mention, no partnership.




