Skip to content
MJ Marketing
Claude Opus 5 benchmarks: Anthropic compared to Fable 5 and GPT-5.6 Sol
AI & Tools
13 min read
Mijo Jurisic

Claude Opus 5: Benchmarks, Price and Cost Paradox

Claude Opus 5 benchmarked: first place in the Intelligence Index with 61 points, half the Fable 5 price and the cost paradox of verbosity.

TL;DR

Anthropic released Claude Opus 5 in July 2026. According to Artificial Analysis the model leads the Intelligence Index with 61 points, ranked 1st of 191 listed models, while costing 5 dollars per million input tokens and 25 dollars per million output tokens, exactly half of Fable 5 and the same as its predecessor Opus 4.8. The catch is verbosity: in the index run Opus 5 produced 100 million output tokens against an average of 63 million, according to Artificial Analysis, which is why the outlet officechai puts the saving per task versus Fable 5 at 26 percent rather than 50. Factual knowledge remains the weak spot: in the AA-Omniscience ranking Opus 5 sits 3rd with 31 points.

Video for this article

Claude Opus 5 im Check: Platz 1 der Welt, halber Preis, und der Haken dabeiWatch on YouTube β†—All videos β†’
Share:

Anthropic released Claude Opus 5 in July 2026, and for the first time an Anthropic model tops the Intelligence Index from Artificial Analysis: 61 points, ranked 1st of 191 listed models. This is an independent measurement, not a vendor claim, and that is what makes the number worth something.

The price is even more interesting to me. Opus 5 costs 5 dollars per million input tokens and 25 dollars per million output tokens, exactly half of Anthropic's own flagship Fable 5 and not a cent more than its predecessor Opus 4.8. Throughout this article I keep two things strictly apart: independently measured numbers from Artificial Analysis on one side, Anthropic's own measurements from the system card and blog post on the other, because reporting tends to blend them. Which models we actually use in the agency we keep transparent on our AI in numbers page.

What Claude Opus 5 is

The most striking product decision is the absence of model variants. Where OpenAI splits the GPT-5.6 line into Luna, Terra and Sol, Opus 5 is, according to the Anthropic documentation, a single model with five effort levels: low, medium, high, xhigh and max. The default is high, in both the Claude API and Claude Code. Fortune describes that control aptly as a dial on your AI bill, because it lets companies meter effort and cost per task themselves.

The model is available through the Claude API, Amazon Bedrock, Google Vertex AI, Microsoft Foundry as well as claude.ai and Claude Code. On Claude Max it is the default model according to the-decoder, and on Claude Pro it is the strongest one available.

The second important point is context. Opus 5 has a one-million-token window, as default and maximum at the same time, with no opt-in and no beta header. Above all there is no long-context surcharge. The Anthropic pricing documentation puts it literally: "Claude 4.6 and later models include the full 1M token context window at standard pricing." A 900,000-token request is therefore billed at the same per-token rate as a 9,000-token one. Maximum synchronous output is 128,000 tokens, and the model's knowledge runs to May 2026.

First place in the Intelligence Index, measured independently

The Artificial Analysis Intelligence Index carries weight precisely because it does not come from the vendor. It bundles several individual tests, among them GDPval-AA v2, Terminal-Bench, SciCode, Humanity's Last Exam, GPQA Diamond and AA-Omniscience.

In that ranking Claude Opus 5 at max effort scores 61 points and lands 1st of 191 listed models. The xhigh level reaches 60 points and 2nd place. Behind them come Claude Fable 5 with 60, GPT-5.6 Sol with 59 and the open model Kimi K3 with 57 points. What is remarkable is less the single point of lead than the constellation: Anthropic beats its own, twice as expensive flagship with the cheaper model.

One qualification belongs here, and it comes from Anthropic itself. In the system card the company writes that Opus 5 is not overall more capable than Fable 5. So the clean phrasing is not "best model in the world" but: Opus 5 leads the Artificial Analysis Intelligence Index, as of the figures evaluated here in July 2026.

The price: half of Fable 5

The price list is the actual core of this story. Per million tokens, Opus 5 costs 5 dollars for input and 25 dollars for output, according to Anthropic. For comparison, Claude Fable 5 sits at 10 and 50 dollars, GPT-5.6 Sol at 5 and 30 dollars, Kimi K3 at 3 and 15 dollars. The predecessor Claude Opus 4.8 cost exactly the same as Opus 5, so the entire improvement arrives without a price increase.

On top of that come the usual levers: a 50 percent discount via the Batch API and 0.50 dollars for a cache hit, which is ten percent of the input price. One detail that matters more in practice than it sounds: the minimum for prompt caching drops from 1,024 to 512 tokens according to the Anthropic documentation. Prompts that were previously too short to cache can now be cached.

The cost paradox: cheaper per token, not half the price per task

This is where it gets uncomfortable, and it is the part most headlines leave out. A halved token price does not automatically mean halved cost, because the bill has two factors: price per token and number of tokens.

Artificial Analysis measured the second factor. During the Intelligence Index run, Opus 5 generated around 100 million output tokens against a model average of 63 million. Artificial Analysis accordingly calls the model "very verbose" and additionally "notably slow" as well as "particularly expensive when comparing to other models of similar price". The cost of a full index run is put at 3,835.51 dollars at max effort and 2,909.91 dollars at xhigh. On output speed, Opus 5 reaches 59.8 tokens per second, ranking 95th of 191, and Artificial Analysis lists time to first token at 66.36 seconds.

What that means for the real bill was estimated by the outlet officechai: it puts the saving per task versus Fable 5 at around 26 percent instead of the 50 percent implied by the token price. Those 26 percent are an officechai figure, not an original Artificial Analysis number, and I quote them explicitly as such.

The officechai takeaway sums the whole development up well: it is not the headline score itself but the cost per point of intelligence that labs now compete on. From my own practice I can only underline that. Anyone comparing models should not line up price lists but measure the cost of a real, typical job.

The jumps that Anthropic measures itself

The following numbers come from Anthropic's own system card and announcement blog post. They are vendor-reported. That does not devalue them, but it has to be said.

The clearest jump is on Frontier-Bench v0.1, a benchmark for agentic tasks. Anthropic reports 43.3 percent for Opus 5 there, against 21.1 percent for Opus 4.8, 33.7 percent for Fable 5 and 34.4 percent for GPT-5.6 Sol. A doubling within one model generation, at the same token price, is unusual.

Even more striking is ARC-AGI-3, a test for novel problem solving without familiar patterns. Opus 5 reaches 30.2 percent according to the system card, while GPT-5.6 Sol gets 7.8 percent and Opus 4.8 gets 1.5 percent. That is close to four times the next best model on a benchmark deliberately built to resist training.

Two further values matter more for everyday work than the record benchmarks. In GDPval-AA v2, which according to Artificial Analysis maps real knowledge work across 44 professions and 9 industries, Anthropic reports 1,861 Elo for Opus 5 versus 1,747 for Fable 5 and 1,736 for GPT-5.6 Sol. Important: the benchmark comes from Artificial Analysis, but the figure quoted here is Anthropic's own measurement. And on computer use, measured with OSWorld 2.0, Anthropic states 70.6 percent versus 66.1 percent for Fable 5 and 62.6 percent for GPT-5.6 Sol. On SWE-bench Verified, Anthropic reports 96.0 percent.

Coding Agent Index: a shared top spot at 67 points

The AA Coding Agent Index is an independent measurement again, and it rates not the model alone but the combination of agent and model. It consists of three equally weighted parts: DeepSWE, Terminal-Bench v2 and SWE-Atlas-QnA.

There, Claude Code with Opus 5 at xhigh effort reaches 67 points and shares the top spot with Codex and GPT-5.6 Sol at max effort, also 67 points. Behind them sit Opus 5 at max effort and Fable 5 with fallback at 66 points each, with Kimi K3 and Claude Opus 4.8 at 61 points each.

One detail deserves attention: the highest effort level is not the best one. Opus 5 at xhigh beats Opus 5 at max by a point, even though max costs more compute. The outlet the-decoder observes the same effect on Frontier-Bench v0.1. More effort produces a worse result there at a higher price. For practice that means picking the most expensive setting blindly is not a strategy.

The weak spot: factual knowledge and hallucinations

No model leads everywhere, and with Opus 5 the weak spot is clearly named. In the AA-Omniscience ranking, which measures factual knowledge and hallucination tendency, Opus 5 lands 3rd with 31 points, behind Claude Fable 5 (40) and Gemini 3.1 Pro Preview (33).

The remarkable part: Anthropic confirms the direction itself. The executive summary of the system card states literally: "We found a surprising number of cases in which Opus 5 confidently stated an answer about which it was in fact unsure. The model hallucinates factual claims slightly more than Opus 4.8, despite being more accurate overall." In plain terms: the model is more accurate overall, yet at the same time it makes slightly more factual claims that are wrong, and it does so with confidence.

For practice this is the most important sentence in the whole article. A model that sounds assured while being wrong a little more often is more dangerous in research work than one that visibly hesitates. Factual claims from Opus 5 need checking, especially when they end up in client texts, proposals or reports.

What this means for small and mid-sized companies

Now the part that counts day to day. This is my assessment based on the numbers documented above and on my own work with these tools, not an external fact.

For long, multi-step tasks and for coding workflows, the data speaks for Opus 5. A shared top spot in the Coding Agent Index, a strong computer-use score and the best result in the work-oriented GDPval measurement add up to a clear profile. Anyone running agents that work autonomously across many steps gets a lot of capability per euro here.

For factual research and knowledge-heavy questions I would not use Opus 5 as the only source. Third place in AA-Omniscience and Anthropic's own statement on the hallucination rate are reason enough to run a second model alongside it or to consistently demand and check sources.

For routine tasks such as summaries, text variants or preparing data, the top class is usually oversized. This is where the effort dial pays off: a low level or a smaller model handles it at a fraction of the cost. And anyone who really wants to push down cost per task should look at prompt caching and batch processing first, not at the price list.

A final point on fairness: according to Fortune, Anthropic still recommends Fable 5 for advanced autonomous projects and positions Opus 5 as the cost-efficient alternative for everyday work. A vendor that does not sell its cheaper model as a cure-all is more credible than one that does.

My take

For me, Claude Opus 5 is the clearest evidence so far that the competition has shifted. It is no longer only about who posts the highest number in a table, but about what a point of intelligence costs. Same price as the predecessor, yet first place in an independent ranking: that is the actual news.

At the same time I consider the verbosity the underrated catch. When a model produces significantly more tokens on average than others, that eats back part of the price advantage. From my own practice I therefore compare models only on real jobs, never on price lists. That is exactly where the verbosity turns up again: in CodeRabbit's code review test Opus 5 produces markedly more nitpicks than the comparison baseline, as set out in the interim review after five days.

And I keep two reservations. First the hallucinations: a confidently wrong model costs more in rework than it saves in operation. Second the short observation window, because numbers from the first few days are no substitute for weeks of real work. If you want to work out which model fits your tasks, your data protection framework and your budget, that is exactly the subject of our AI consulting. How such tools fit into Google Ads workflows in concrete terms I described in the post on AI automation in everyday Google Ads work. For comparison it is also worth looking at the benchmark and cheating debate around GPT-5.6 Sol and at the comeback of Fable 5.

Sources

As of: July 2026

Frequently asked questions

What is Claude Opus 5?

Claude Opus 5 is Anthropic's model released in July 2026. Instead of separate model variants, the Anthropic documentation lists five effort levels of a single model: low, medium, high, xhigh and max, with high as the default in the Claude API and Claude Code. The context window spans one million tokens as both default and maximum, with no long-context surcharge, and the knowledge cutoff is May 2026.

How does Claude Opus 5 perform in the benchmarks?

In the independently measured Artificial Analysis Intelligence Index, Opus 5 at max effort scores 61 points and ranks 1st of 191 models, ahead of Fable 5 (60), GPT-5.6 Sol (59) and Kimi K3 (57). In the AA Coding Agent Index, Opus 5 in Claude Code shares the top spot with 67 points alongside Codex and GPT-5.6 Sol. Anthropic's own measurements include 43.3 percent on Frontier-Bench and 70.6 percent on OSWorld 2.0.

What does Claude Opus 5 cost?

According to the Anthropic price list, Opus 5 costs 5 dollars per million input tokens and 25 dollars per million output tokens. That is exactly half of Claude Fable 5 (10 and 50 dollars) and identical to its predecessor Opus 4.8. The one-million-token context window is included in the standard price; there is no long-context surcharge according to the documentation.

Is Claude Opus 5 really half the price of Fable 5?

Per token yes, per task no. Artificial Analysis calls Opus 5 very verbose and notably slow: 100 million output tokens in the index run against an average of 63 million. The outlet officechai therefore calculates around 26 percent saving per task versus Fable 5 instead of the 50 percent the token price suggests.

Where is Claude Opus 5 weak?

On factual knowledge. In the AA-Omniscience hallucination benchmark from Artificial Analysis, Opus 5 ranks 3rd with 31 points, behind Fable 5 (40) and Gemini 3.1 Pro Preview (33). Anthropic confirms the direction in its own system card, writing that the model surprisingly often stated answers confidently while actually being unsure. Factual claims therefore still need to be verified.

Mijo Jurisic

Google Ads consultant & founder of MJ Marketing. Five-plus years of hands-on practice: from a self-taught start to the Google Premier Partner programme with 500+ direct Google Ads clients and €20M+ in managed media spend.

Share this article:

Share: