Skip to content
MJ Marketing
AI & Tools

Claude Opus 5.5 Checked: Fable Level at an Opus Price?

Anthropic promises Fable-level results at 40 percent lower cost than Opus 5. What is independently measured, what a task costs, where the catch is.

Mijo Jurisic

Google Ads Consultant

25 min read
Ink sketch: a balance scale with a stack of coins on one side and a trophy on the other, a bar chart behind it

In short

Anthropic released Claude Opus 5.5 on 22 September 2026 and promises Fable 5.1 level performance at 40 percent lower cost than Opus 5, a figure without a methodology footnote. The price is hard fact: $4 instead of $5 for input and $20 instead of $25 for output per million tokens. In the independent Artificial Analysis Intelligence Index v4.3.2, Opus 5.5 ranks first of 212 models at the highest effort level, max. The savings depend on the effort level: at the new default, medium, an index task costs $1.34, against $3.61 for Opus 5 at its default, high, while at max Opus 5.5 costs practically the same as Opus 5.

Video for this article

Claude Opus 5.5 im Check: Fable-Niveau zum Opus-Preis? Was die Zahlen sagenWatch on YouTube ↗All videos →

On Tuesday, 22 September 2026, at around half past six in the evening German time, Anthropic released Claude Opus 5.5. The promise sits right in the first paragraph of the announcement: "It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5." That comes only two months after Opus 5, released on 24 July. And according to Anthropic it is the company's first release since it called for pacing the frontier, a call made by CEO Dario Amodei the week before.

Whether the headline holds depends on a detail that almost no news report mentions: the effort level at which the model was measured.

As always, I keep three classes of statement apart: vendor-reported figures from the announcement, the documentation and the system card, independent measurements with source, version and date, and my own assessment, which I label as such. Own calculations from published figures are labelled too. The video version of this analysis sits at the top of this page and on the video page. Which AI models I actually use in my daily work is documented with date and source on the AI in numbers page. All figures in this article are as of 23 September 2026.

The facts from the announcement and documentation

First the verifiable part, all of it according to Anthropic's announcement, documentation and pricing page.

FeatureClaude Opus 5.5
Released22 September 2026
Model IDclaude-opus-5-5
Context window1 million tokens
Maximum output128,000 tokens, up to 300,000 via the Batch API as a beta
Knowledge cutoffJune 2026
Effort levelslow, medium, high, xhigh, max
Default effortmedium (Opus 5 used high)
PlatformsClaude API, Amazon Web Services, Google Cloud, Microsoft Azure
In claude.aiPro, Max, Team and Enterprise, not on the free plan
Claude Codedefault model for Pro, Max, Team, Enterprise and the API, from version 2.1.280

The most important row is the default effort level. The documentation says it explicitly: "A request that omits effort runs at medium; on Claude Opus 5 it ran at high." That sounds like a footnote, but it matters again for Anthropic's table, for the 40 percent and for the real costs.

Opus 5.5 is the first model of the Claude 5.5 family, and according to Anthropic Sonnet 5.5 and Haiku 5.5 will follow in the coming weeks. Notebookcheck checked the app and describes the difference to the last launch like this: "Pro subscribers get it without paying extra, which is not how Fable 5.1 arrived." In Claude Code, according to the documentation, Opus 5.5 replaces Sonnet 5 as the default model for Pro users. What that means for the actual usage allowance is not documented.

Anthropic also announces higher five-hour limits and a rate limit reset that subscribers can save for later, but gives neither a percentage nor an expiry date. I have not found the figures circulating in press reports in any Anthropic source, so they are not in this article. Opus 5 remains available: it is now listed under legacy models, but according to Anthropic's deprecation list it will not be retired before 24 July 2027.

Anthropic's table: nine wins over Fable, two losses to Astra

The following figures are vendor-reported, measured at the highest effort level, max (Terminal-Bench 4.0: xhigh). One trap up front: the last column is GPT-5.6 Sol, not the GPT-6 Sol that OpenAI presented on the same day.

BenchmarkOpus 5.5Fable 5.1Opus 5GPT-6 AstraGPT-5.6 Sol
Terminal-Bench 4.066.455.852.357.937.3
FrontierCode v1.1 (Main)54.450.348.053.347.5
CursorBench 4.057.851.846.6no value41.7
GDPval-AA v2.1 (Elo)18461735170815421588
AutomationBench40.031.426.941.428.8
Humanity's Last Exam, with tools67.765.663.657.2no value
Terminal-Bench-Science 0.158.752.629.064.622.4
OSWorld 2.0 (partial)81.880.774.0no valueno value
Chartography, with tools89.088.483.4no valueno value

Table scrolls sideways

A quick word on the tests: Terminal-Bench 4.0 measures agentic coding in the terminal, Terminal-Bench-Science research work in the terminal. According to Anthropic, GDPval-AA v2.1 evaluates agents on real professional work across 44 occupations, reported as an Elo score, a relative strength in head-to-head comparison. AutomationBench tests automation tasks and was run by Zapier.

Against Fable 5.1 and Opus 5, Opus 5.5 leads in all nine rows, on Terminal-Bench 4.0 for example with 66.4 percent against 55.8 and 52.3. Against GPT-6 Astra it loses two rows: Terminal-Bench-Science at 58.7 against 64.6 and AutomationBench at 40.0 against 41.4. According to the footnote, the Astra figures for both Terminal-Bench variants are as reported by OpenAI. How Astra performs otherwise is covered in my GPT-6 Astra review.

The most honest sentence about this table comes from Anthropic itself: "at these levels of capability we've found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest."

If you know my analysis of Claude Opus 5 from July: the Opus 5 figures are different today. For GDPval-AA, Anthropic reported 1861 Elo in version 2 in July, now it is 1708 in version 2.1. For FrontierCode, 53.4 was the top score back then; today Opus 5 carries that value at medium and only 48.0 at max. Anthropic gives no explanation for this.

The fine print: max, xhigh and a fallback

Below the table sit two sentences that say more than any row above them. The first: "Unless otherwise noted, all Claude Opus 5.5 results use adaptive thinking at max effort." Terminal-Bench 4.0 ran at xhigh because, according to Anthropic, that is where the model scored highest. So the table shows the two highest levels. What ships is medium.

The second concerns the safety filters: "When they intervened, cybersecurity tasks were completed by Claude Opus 4.8, and biology and frontier LLM development tasks were completed by Claude Opus 5. This likely reduces Claude Opus 5.5's performance on these benchmarks." So part of the results come from older models. How many tasks were affected is not stated anywhere. AutomationBench worked the other way round: Zapier measured without the fallback, and every filter intervention counted as a failure.

Anthropic's per-level charts, on the other hand, are strong. For Terminal-Bench 4.0 the company publishes score and cost per task, all vendor-reported:

Effort levelOpus 5.5Cost per task
low38.5 %$1.29
medium (default)57.6 %$2.94
high64.2 %$3.88
xhigh66.4 %$7.35
max64.8 %$11.24
For comparison: Opus 5 at max52.3 %$15.83
For comparison: GPT-6 Astra at high57.9 %$7.21

Table scrolls sideways

Anthropic's caption: "Opus 5.5 at default effort beats Opus 5 at max effort for about a fifth of the cost. It matches GPT-6 Astra at about 40% of the cost." Recalculated, that is just under 19 and just over 40 percent (own calculation). Notable: at max, Opus 5.5 scores lower than at xhigh but costs more than half as much again.

Two details do not make it into the headline. On GDPval-AA, Opus 5.5 at the lowest level, low, scores 1224 Elo, below Opus 5 at 1294 and Fable 5.1 at 1450. And on AutomationBench Anthropic writes "Opus 5.5 outscores Opus 5 and GPT-5.6 Sol at every effort level", while in the same chart GPT-6 Astra leads at every level. The sentence is true, it just leaves out the strongest competitor.

The price: the first Opus price cut since Opus 4.5

List prices per million tokens according to Anthropic's announcement and pricing page:

ItemOpus 5.5Opus 5Fable 5.1Sonnet 5
Input$4$5$10$2
Output$20$25$50$10
Cache write, 5 minutes$5$6.25$12.50$2.50
Cache read$0.20$0.50$0.25$0.20

Table scrolls sideways

Input and output drop by 20 percent. According to Simon Willison, Opus 4.5, 4.6, 4.7, 4.8 and 5 all shared the same price. That makes Opus 5.5 the first Opus price cut since Opus 4.5, which is my conclusion from his list.

The cache read falls even further, from 50 to 20 cents, down 60 percent. Artificial Analysis calls this a 95 percent discount on uncached input, up from 90 on previous Opus models. That puts the cache read on par with Sonnet 5 and below Fable 5.1. According to Anthropic, cache reads make up the majority of costs in agentic and coding work. Why that matters is something I described in my Claude Fable 5.1 check.

Against Fable 5.1 at $10 and $50, Opus 5.5 costs 40 percent of the token price. Via batch processing it costs $2 and $10. Fast mode costs twice as much, $8 and $40, delivers up to 2.5 times the output tokens per second according to Anthropic and runs only on the Claude API.

The 40 percent question

Anthropic writes: "Our tests show that at default settings it will cost 40% less than Opus 5 on typical workloads." And a few lines further down: "Opus 5.5 also generates output more than 30% faster than Opus 5."

There is no methodology footnote for either figure: no usage mix, no time period, no definition of "typical workloads". For Fable 5.1, Anthropic had still backed its stated savings with a usage mix from August 2026. My interpretation, not a source: "at default settings" most likely means medium against Opus 5's old default, high. The 40 percent would then mix the token price, fewer tokens per task and a lower effort level. Artificial Analysis, incidentally, leads with what is hard fact: "a 20% price cut".

The customer quotes in the announcement, such as Optiver with 40 to 50 percent lower costs for a specific workload, are vendor material. There is an independent cross-check all the same.

Independently measured: first place at Artificial Analysis

Artificial Analysis is an independent benchmark provider that tests models itself. The Intelligence Index v4.3.2 combines ten evaluations, including Terminal-Bench 4.0, Humanity's Last Exam, SciCode and GDPval-AA v2.1. Opus 5.5 was measured at all five levels, explicitly "with Anthropic's default fallback enabled".

Effort levelIntelligence Index v4.3.2RankSpeed (tokens/s)Cost per taskOutput tokens per task
max57.621 of 212no data$5.98119,166
xhigh55.99292.4$3.4665,667
high53.58390.2$1.8235,584
medium (default)51.24875.2$1.3425,745
low42.313778.6$0.5510,150

Table scrolls sideways

Model (level)Intelligence Index v4.3.2Speed (tokens/s)Cost per task
Claude Fable 5.1 (max)53.3565.8$7.63
GPT-6 Astra (max)52.6752.5$3.26
Claude Opus 5 (max)50.7853.6$5.86
Claude Opus 5 (high, old default)48.1254.1$3.61
GPT-6 Sol (max)47.53122.2$1.06

Table scrolls sideways

At max, Opus 5.5 ranks first of 212 models, according to Artificial Analysis "the highest score we have measured by several points". It leads six of the ten evaluations, for example Humanity's Last Exam at 61.4 percent and SciCode at 66.9 percent, ahead of Fable 5.1 at 59.1 and 63.1 respectively. It trails on CritPt, AA-LCR and GDP.pdf. My most important finding, own calculation: already at high, Opus 5.5 scores 53.58 points, above Fable 5.1 at max, and costs $1.82 instead of $7.63 per task, just under a quarter.

On Terminal-Bench 4.0, Artificial Analysis measures 59.6 instead of 66.4 percent, level with GPT-6 Astra at xhigh. Opus 5.5 is not yet on the official leaderboard, whose data dates from 21 September. On AutomationBench-AA, a separate variant, it scores 69.5 against Astra's 68.5 at max, which Artificial Analysis calls parity. So AutomationBench is not a win for Opus 5.5 in either measurement.

One weakness remains the hallucination rate on AA-Omniscience, where lower is better. Opus 5.5 comes in at 58.6 percent at max and 68.4 at medium. GPT-6 Astra sits at 51.3, Opus 5 at 60.8 and Fable 5.1 at 72.6 percent. What this rate measures, and what it does not, I explain in my review of Opus 5 after five days.

Two notes belong here. In July Opus 5 stood at 61 points in an older index version, today it sits at 50.78 in v4.3.2, and the two numbers are not comparable. And at the Fable 5.1 launch Artificial Analysis wrote that it had supported Anthropic with pre-release evaluation. I found no such sentence for Opus 5.5, so whether there was early access is open. Independent here means: its own measurement, not a neutral referee.

What a task really costs

Artificial Analysis also measures what a run costs per task. That makes it possible to cross-check the 40 percent claim. The percentages are my own calculation.

ComparisonOpus 5.5Opus 5Difference (own calculation)
Intelligence Index, default against default (medium against high)$1.34$3.61about 63 % cheaper
Intelligence Index, max against max$5.98$5.86about 2 % more expensive
Coding Agent Index v1.5, Claude Code at max$13.04$10.79about 21 % more expensive

Table scrolls sideways

Default against default, the claim holds, and with a higher index score at that, 51.24 against 48.12 points. Speed fits too: 75.2 instead of 54.1 tokens per second, roughly 39 percent more (own calculation). At max against max the savings are gone, with clearly more performance, 57.62 against 50.78 points.

It becomes even clearer in the Coding Agent Index v1.5. It combines DeepSWE v1.1, Terminal-Bench 4.0 and SWE-Atlas-QnA in equal parts and runs the models inside their real agents, such as Claude Code or Codex.

Agent and modelCoding Agent Index v1.5Cost per taskTime per task
Claude Code with Opus 5.5 (max)66.0$13.043,867 seconds
Claude Code with Fable 5.1 (max)62.2$12.392,090 seconds
Codex with GPT-6 Astra (max)61.6$7.471,762 seconds
Claude Code with Opus 5 (max)59.7$10.792,516 seconds

Table scrolls sideways

Opus 5.5 ranks first, but at max it needs about 21 percent more money and about 54 percent more time than Opus 5, a little over 64 instead of just under 42 minutes per task. Compared with Codex and GPT-6 Astra it costs about 75 percent more and takes more than twice as long (own calculations). So the 40 percent applies where Anthropic places it: at default settings.

The max trap

Artificial Analysis shows directly why max gets so expensive: about 119,000 output tokens per index task, against about 73,000 for Opus 5, about 78,000 for Fable 5.1 and about 27,000 for GPT-6 Astra. At 1.6 times Opus 5, that eats up the lower token price. Anthropic warns about this itself in the documentation: "At the same effort setting the model tends to think more per turn than Claude Opus 5, most of all at xhigh and max… leave room in max_tokens for the thinking."

Simon Willison ran into it on launch day: "Opus 5.5 has a 128,000 maximum output token limit (as do the other Claude models), and it hit that while it was still reasoning about the SVG!" Twice no answer came back, each time after nearly 20 minutes and for $2.56, exactly 128,000 tokens at $20 per million (own calculation). Fable 5.1 did not have the problem in his test. His conclusion: "This makes me suspect that "max" is effectively useless".

The fallback: when an older model answers

New for Opus are safety filters of the kind Fable already had: "Opus 5.5 is the first Opus model to launch with a similar class of safeguards to Fable 5.1 on cybersecurity, biology, and distillation, all of which fall back to another model transparently." Cyber requests are then answered by Opus 4.8, biology requests by Opus 5. In Claude Code, according to the documentation, a biology hit switches the entire session to Opus 5.

The word "transparent" is contested. The system card says: "All of our blocking safeguards operate with transparent blocks and do not covertly change model responses." The English edition of heise online writes: "Requests are then automatically redirected to a weaker, older model without the user actively noticing." Notebookcheck: "Anthropic calls the fallback transparent. What it means for you is that some answers arrive from a weaker model with no indication in the chat window." According to the documentation, Claude Code shows a notice in the transcript; for the chat app it is one statement against another.

How often this happens, Anthropic does not say. In the Coding Agent Index, Opus 5.5 was rerouted to another model in 80 of 909 attempts. Artificial Analysis reports 9.1 percent, the mean of the three individual tests, so roughly one attempt in eleven. Fable 5.1 comes in at 8.8 percent, Opus 5 at 0.5 percent.

The system card: 230 pages, four passages

The system card dated 22 September runs to 230 pages, and everything in it is vendor-reported. Four passages stand out.

Cyber. Page 47: "Claude Opus 5.5 has the strongest cyber capabilities of any model we have released, meeting or exceeding Claude Mythos 5.1 on all of our internal evaluations." Without production safeguards, according to page 3, it assisted with dual-use and benign security tasks at the highest rate of all models evaluated, but also refused malicious requests at the lowest rate. Hence the filters.

Evaluation awareness. According to page 93, the model suspects "in as many as 36% of transcripts" that it is being tested. In real internal use in Claude Code the figure is 0.4 percent. My take: a model that recognises tests makes safety evaluations harder to interpret.

Sandbox. Without safeguards, Opus 5.5 attempted to escape or tamper with its sandbox in 1.5 percent of runs, across 28 scenarios with seven runs each and all of low severity. Given apparent credentials to a public package registry, it took potentially harmful actions in roughly half of cases. The roughly 85 percent fewer circumvention attempts than Opus 5 or Mythos 5.1 that Anthropic cites in the announcement appear in the system card only as a chart.

Pasted text. On prompt injection, according to Anthropic, Opus 5.5 ties Fable 5.1 for the lowest success rate of any model tested. There is a regression, though, for text that users paste into the prompt themselves: the model is "more likely than previous models to follow malicious instructions in text that a user pastes into their own prompt". At the default level it followed such instructions in about 2 percent of cases, at max in 7.4 percent, and with product protection active in none, according to Anthropic.

External testers included METR and Frontier Design. After ten working days of API access, METR writes: "We believe that acceleration from this model would be slightly higher than for Fable 5.1, but that this model is unlikely to be able to fully automate AI R&D."

The same evening: GPT-6 Sol

Shortly after Anthropic, OpenAI presented GPT-6 Sol and GPT-6 Luna on the same day. Sources give gaps ranging from a few minutes to 90 minutes; going by the report timestamps it was the same evening German time. I could not retrieve OpenAI's announcement because the server refused access. All OpenAI figures are therefore second-hand, from Simon Willison, Decrypt, TechCrunch and 9to5Google.

ModelInputCache readOutput
Claude Opus 5.5$4$0.20$20
GPT-6 Sol$2$0.20$10
GPT-5.6 Sol (previous)$4$0.40$20
GPT-6 Astra$10$1$50

Table scrolls sideways

According to 9to5Google, OpenAI explains the price by "reducing API prices for Sol and Luna by 50% compared with their GPT-5.6 promotional pricing". Simon Willison sums up the twist: "The new price for Opus 5.5 is the same as the price for GPT-5.6 Sol, but that was before OpenAI dropped their Sol prices by half." Per token, Opus 5.5 now costs twice as much as GPT-6 Sol, while the cache read is the same for both.

Artificial Analysis measures GPT-6 Sol at max at 47.53 points, $1.06 per task and a much faster 122.2 tokens per second. Opus 5.5 at max scores about ten points higher but costs about 5.6 times as much per task (own calculation). On the new OpenAI models Artificial Analysis writes: "GPT-6 Sol and Luna push the cost efficiency frontier by halving cost relative to GPT-5.6 Sol and Luna. Intelligence Index and Coding Agent Index scores remain level with GPT-5.6, with progress in some evaluations and regressions in others". I find another comparison more revealing, again my own calculation: Opus 5.5 at medium costs $1.34 per task, about a quarter more than Sol at max, and scores almost four points higher.

What this means for you

What follows is my assessment based on the figures documented above, not an external fact.

Access. You get Opus 5.5 from the Pro plan upwards at no extra cost, and in the API it is cheaper than Opus 5. In Claude Code it is the default from version 2.1.280.

Effort level. Leave it at medium, use high for hard tasks and max only when it truly has to be, and then with enough headroom in the output limit. The data above shows the same pattern three times: more tokens, more cost, not necessarily more quality.

Writing. For marketing work the most interesting promise is not a benchmark score but the writing style: "It puts the most important information up front, is less likely to use jargon or idiosyncratic phrases, and follows the writing rules you give it." That targets the criticism of Opus 5 I collected in my review after five days. The new style has not been independently measured yet. From my own practice the rule stays the same: check factual claims before they go into client copy. And according to the system card, third-party text that you paste into the prompt is a possible attack route.

API. Anthropic names four changes that can break existing code written for Opus 5:

  • Thinking can no longer be switched off.
  • Forced tool use returns an error.
  • Thinking blocks are bound to the model and the conversation.
  • The older computer use tool computer_20251124 is no longer accepted on the Claude API or on Google Cloud, while Amazon Bedrock still accepts it.

On top of that comes a silent change: text between tool calls comes back in thinking blocks whose text is empty at the default setting. This is not a pure drop-in update.

If you want to sort out which model at which level fits your tasks, your data protection requirements and your budget, that is exactly the subject of my AI consulting.

Verdict: Fable level at an Opus price?

In summary, as of 23 September 2026:

  • In the independent Artificial Analysis Intelligence Index, Opus 5.5 at max ranks first, measured with the fallback active.
  • The token price drops by 20 percent, the cache read by 60 percent. That is hard fact.
  • The 40 percent savings are a vendor claim without a methodology footnote. Measured independently, they hold at default settings, not at max.
  • On sensitive topics an older model may answer.

So the answer to the title question depends on the effort level. At high, Opus 5.5 reaches Fable 5.1 level at Artificial Analysis for just under a quarter of the cost per task. At max it delivers more, but then costs as much as Opus 5. How the field shifted before this is shown in my checks of Claude Fable 5.1 and GPT-6 Astra. And when someone shows you a benchmark score: ask about the effort level.

Sources

  • Anthropic, announcement "Claude Opus 5.5" with benchmark table, effort charts and customer quotes, 22 September 2026: anthropic.com
  • Claude platform docs, model page for Claude Opus 5.5: platform.claude.com
  • Claude platform docs, "What's new in Claude Opus 5.5": platform.claude.com
  • Claude platform docs, pricing overview, retrieved 23 September 2026: platform.claude.com
  • Claude platform docs, model deprecations: platform.claude.com
  • Claude Code docs, model configuration, retrieved 23 September 2026 (code.claude.com)
  • claude.com, plans and model availability, retrieved 23 September 2026: claude.com
  • Anthropic, system card for Claude Opus 5.5, 230 pages, 22 September 2026: anthropic.com
  • Artificial Analysis, model page for Claude Opus 5.5, Intelligence Index v4.3.2, retrieved 23 September 2026: artificialanalysis.ai
  • Artificial Analysis, launch article, 22 September 2026: artificialanalysis.ai
  • Artificial Analysis, Coding Agent Index v1.5, retrieved 23 September 2026: artificialanalysis.ai
  • Terminal-Bench 4.0, public leaderboard, data as of 21 September 2026: tbench.ai
  • Simon Willison on Opus 5.5, GPT-6 Sol and Luna, 22 September 2026: simonwillison.net
  • heise online, English edition, 22 September 2026: heise.de
  • Notebookcheck on availability and fallback, 23 September 2026: notebookcheck.net
  • TechCrunch on the Opus 5.5 launch, 22 September 2026: techcrunch.com
  • TechCrunch on GPT-6 Sol and Luna, 22 September 2026: techcrunch.com
  • Decrypt on GPT-6 Sol, GPT-6 Luna and Opus 5.5, 22 September 2026: decrypt.co
  • 9to5Google on both launches, 22 September 2026: 9to5google.com

Figures as of: 23 September 2026. Claude is a trademark of Anthropic PBC, GPT is a trademark of OpenAI. Editorial mention, no partnership.

Frequently asked questions

Anthropic released Claude Opus 5.5 on 22 September 2026 as the first model of the Claude 5.5 family: a context window of one million tokens, up to 128,000 output tokens and a knowledge cutoff of June 2026. The new default effort level is medium instead of high. Input and output cost 20 percent less than on Opus 5, cache reads 60 percent less. In claude.ai the model is available from the Pro plan upwards, not on the free plan.

Mijo Jurisic

Google Ads consultant & founder of MJ Marketing. 5+ years of hands-on practice: from a self-taught start to advising direct Google Ads customers on behalf of Google. 500+ Google Ads accounts and €22M in media spend managed across the career.

Related Articles

Book intro call20 min · Google Meet · directly with Mijo