Skip to content
MJ Marketing
Claude Opus 5 experience report: effectively tied in the Intelligence Index, contested in daily practice
AI & Tools
13 min read
Mijo Jurisic

Claude Opus 5 After Five Days: Experience and Criticism

Claude Opus 5 after five days: effectively tied in the Intelligence Index, a 50 percent hallucination rate and rank 5 in the blind user vote.

TL;DR

Claude Opus 5 has been public since 24 July 2026, and this article takes stock after five days, with all figures as of 28 July 2026. Artificial Analysis measures 61 points for Opus 5 and 60 for Fable 5 and calls the result effectively tied itself, while also stating that it supported Anthropic in evaluating the model ahead of release. The same evaluation reports a 50 percent hallucination rate alongside 7 points more factual knowledge than Opus 4.8. In the blind user vote of the Arena leaderboard, Opus 5 sits at rank 5 in the overall text ranking, with a strikingly wide confidence interval. In CodeRabbit's code review test the model finds fewer known bugs than the comparison baseline yet is right more often on actionable findings. And Anthropic recommends in its own documentation that certain verification instructions be deleted from old prompts rather than new ones added.

Video for this article

Opus 5 nach fünf Tagen: Was die Zahlen wirklich sagenWatch on YouTubeAll videos
Share:

Claude Opus 5 has been public since 24 July 2026. This article takes stock after five days, with all figures as of 28 July 2026 and the supporting screenshots taken on 29 July. Five days is too short for a verdict and long enough for a mood check. The release benchmarks are explicitly not the subject here; they are covered in the benchmark article on the release. This one is about the rift between measurement and daily work.

That rift can be shown in a single post. Tester Claire Vo ran seven models blind against each other and put Opus 5 in first place, ahead of all six others. In the same post on X dated 24 July she writes, in the original: "I hate working with it." Two lines further down, in the same post: "in a blind taste test, I ranked it above every other model (even Fable and my beloved GPT-5.6)". That contradiction is the finding after five days, not the leaderboard.

Five days, not even a week

Before the numbers, the method. I sort every source in this review into one of three boxes, because the labels keep slipping in the reporting.

Vendor-reported covers Anthropic's own publications: announcement, system card and documentation. Independently measured covers figures the vendor was not involved in producing, here the blind user vote of the Arena leaderboard, Claire Vo's blind test and CodeRabbit's code review test. And vendor-adjacent is a third category I need in this case: Artificial Analysis writes in its own article that it supported Anthropic in evaluating the model ahead of release. Those figures are therefore neither a vendor claim nor an independent cross-check. In the benchmark article on the release I still listed Artificial Analysis as an independent measurement. That was too generous a phrasing, and I am correcting it here.

Which models I actually use in daily work, and how much, I keep transparent on the AI in numbers page. The same review exists as a video; it sits at the top of this page and in the video overview.

61 to 60, and who measured it

Artificial Analysis measures 61 points in the Intelligence Index for Claude Opus 5 at the highest effort level and 60 points for Claude Fable 5. What matters is how the source itself frames that. It calls the result, literally, "effectively tied with Claude Fable 5 (max, 60)". Effectively tied, then, and explicitly not a win. Anyone building a ranking out of a single point of separation is building it against the wording of the measuring body.

The second sentence from the same source belongs with it: "We supported Anthropic to evaluate Claude Opus 5 ahead of release". That does not make the numbers wrong. It does change their label, which is why this article says vendor-adjacent everywhere Artificial Analysis is the source. How quickly a benchmark number turns into a debate about the measurement itself, I described in the benchmark and cheating debate around GPT-5.6 Sol.

Incidentally, the constellation remains remarkable, because Fable 5 is the more expensive model from the same house. What became of the comeback of Fable 5 will only show over weeks, not over five days.

The 50 percent hallucination rate, and what it does not mean

The same evaluation reports a hallucination rate of 50 percent, 14 points higher than for Claude Opus 4.8. This number is misread almost everywhere, so here is the definition first.

It does not mean that every second answer is wrong. Artificial Analysis defines the rate on the AA-Omniscience knowledge test as incorrect divided by incorrect plus partial plus not attempted. Only the non-correct answers are counted, and half of those are wrong guesses rather than skipped questions. This is about behaviour under uncertainty, not about the overall hit rate.

The counter-direction is in the same text and belongs there just as much: Opus 5 achieves 7 points more factual knowledge than Opus 4.8. Taken together, that gives a clear behavioural profile. The model knows more and answers more often, even when it is unsure. For practice this means the same as it did at release: factual claims need checking before they end up in client texts, proposals or reports.

Rank 5 in the blind user vote

The second data point is a blind user vote, the Arena leaderboard. In the overall text ranking, Claude Opus 5 at max effort sits at rank 5 with 1495 points. Ahead of it are Claude Fable 5 with 1508, Claude Opus 4.6 Thinking with 1505, Claude Opus 4.7 Thinking with 1502 and Claude Opus 4.6 with 1497. All four come from Anthropic itself, so in the user vote the newest model sits behind three of its own predecessors.

Then the qualification, without which the rank says little. The confidence interval is plus minus 12, while the established models on the list sit at plus minus 4 to 6. The value is therefore considerably less certain than the ones above it, and Arena flags it in a tooltip as being based on pre-release testing. The page states no data cutoff; my own retrieval was on 29 July 2026.

First place in the blind test, and still no praise

The sharpest single finding comes from Claire Vo. She tested seven models blind against each other across six task types, published on 25 July on ChatPRD. Opus 5 wins that test, ahead of all six other models, including Fable 5.

In the same text she calls the output "Claude Slop" and the model "neurotic" and "timid". The page says, in substance, that it is her most loathed colleague and still delivers the best work. The sharper original wording is not there but in the X post from 24 July quoted at the top, in which both statements sit in a single post.

That is not a contradiction in the tester's thinking but a pointer to two different quantities. With this model, output quality and working experience come apart. Anyone who only reads leaderboards misses that entirely.

CodeRabbit: fewer bugs found, right more often

The same rift becomes measurable at CodeRabbit, a commercial code review provider that has published its methodology. The commercial background has to be stated, and the numbers are still worth reading closely.

In that test Opus 5 at x-high effort finds 55.2 percent of the known bugs, while the comparison baseline sits at 61.1 percent. What matters is what that baseline is: not a competing model but CodeRabbit's established production mix, that is, a system tuned over a long period. The comparison is therefore harsher than it looks at first glance.

On precision the picture turns, and it turns in both directions. On actionable findings Opus 5 is right more often, 39.3 against 35.2 percent. Measured across the entire output it flips back, 28.6 against 32.8 percent. On top of that come 92 nitpicks instead of 23 and, according to the same results table, around 60,500 instead of around 40,500 input tokens. The provider's conclusion is, in substance: a specialist, not the sole reviewer. The verbosity that was already noticeable at release turns up a second time here, this time in a measurement the vendor was not involved in. That is carried by the 92 nitpicks instead of 23; the higher input token count belongs to the input side and says nothing about the size of the output.

For Claude Code users: delete instructions instead of adding them

Now the part that changes something fastest in daily work. In its own documentation on prompting Opus 5, Anthropic recommends removing certain verification instructions from old prompts rather than adding new ones. Named among others are formulations such as double-check your answer, re-verify before responding, include a final verification step for any non-trivial task, or use a subagent to verify. The documentation's reasoning: Opus 5 verifies its own work, and duplicate instructions to do so only create additional effort.

From my own practice, this is the point where old project files bite back most easily. Anyone who has written verification lines into a prompt or project file over months now has a clean-up task instead of an extension task. That is unfamiliar, because model changes usually go the other way.

You could read a migration shock into this. The documentation does not support that. It states explicitly that the model performs well with existing Claude Opus 4.8 prompts. What the same documentation does confirm is exactly the criticism coming from the community: longer responses, longer written files, more self-commentary and an expanded task scope.

The proportions: loud is not a majority

Volume needs a sense of scale, otherwise every angry post becomes a trend. The announcement thread for Opus 5 on Hacker News sits at around 1,700 points, while the thread calling Opus 5 a really bad model sits at just over 50.

No majority picture follows from that, in either direction. Points on an announcement are not a vote of approval for the model but an attention score. And a small critique thread neither refutes the criticism nor proves it. Both are mood indicators, not measurements, and that is exactly how I treat them here.

What I did not verify

Honesty includes the limits of one's own research. Three of those limits matter enough to name.

First, I only saw the Reddit sentiment through an aggregator and took no upvote figure from there. Second, from the Arena leaderboard I evaluated only the overall text ranking, not the sub-leaderboards. A model can sit quite differently in individual categories, and I make no claim about those categories. Third, five days is simply not an observation period, and several of the reports quoted here rest on a few days of use.

My take after five days

What follows is my assessment based on the figures documented above and on my own work with these tools, not an external fact.

The most stable pattern across all sources is the split by task type. On narrowly scoped tasks the data reads well: Snorkel AI reports the best result for Opus 5 on bug and performance investigation, Claire Vo's blind test puts it ahead on frontend work, and on actionable findings in the CodeRabbit test it is right more often than the comparison baseline. On long, open-ended tasks the reports read mixed, and the criticisms repeat strikingly: more text, more scope, more initiative than ordered. The outlet leadwithai likewise describes longer pieces of work as less ambitious.

In practice that means two things to me. First, it pays to cut tasks more narrowly instead of handing over one long assignment in a single go. Second, it pays to look into your own prompt and project files, because they carry instructions this model no longer needs. Both are a manageable amount of work, and from my own practice the difference shows up quickly.

And I stick to the restraint this article started with. One point of separation in a vendor-adjacent ranking is not a win, a fifth place with a wide confidence interval is not a defeat, and a single angry thread is not a consensus. Anyone selling a final verdict after five days is selling an opinion as a measurement. If you want to work out which model fits which tasks, which data protection framework and which budget in your company, that is exactly the subject of my AI consulting. For a look at the open-weight competition, the post on Kimi K3 is also worth reading.

Sources

Figures as of: 28 July 2026, screenshots retrieved 29 July 2026

Frequently asked questions

How good is Claude Opus 5 according to the first hands-on reports?

Mixed, and that is the actual finding. Artificial Analysis measures 61 points in the Intelligence Index against 60 for Fable 5 and calls that effectively tied itself. In the blind user vote of the Arena leaderboard, Opus 5 sits at rank 5 in the overall text ranking. Tester Claire Vo ranked the model first across six task types in a blind test, ahead of all six comparison models, and sharply criticises its output in the same text. Five days is too short for a final verdict.

What does the 50 percent hallucination rate for Claude Opus 5 mean?

Not that every second answer is wrong. Artificial Analysis defines the rate on the AA-Omniscience knowledge test as the share of incorrect answers among all non-correct answers, that is incorrect divided by incorrect plus partial plus not attempted. Of the non-correct answers, the model therefore guesses wrongly in half the cases instead of skipping the question. The value is 14 points above that of Opus 4.8, while the same evaluation reports 7 points more factual knowledge.

Why is Claude Opus 5 only at rank 5 in the Arena user vote?

In the overall text ranking, four models sit ahead of Opus 5 according to the Arena leaderboard, and all four come from Anthropic itself: Claude Fable 5 with 1508 points, Claude Opus 4.6 Thinking with 1505, Claude Opus 4.7 Thinking with 1502 and Claude Opus 4.6 with 1497, against 1495 for Opus 5 at max effort. The uncertainty matters: the confidence interval is plus minus 12, while established models sit at plus minus 4 to 6. Arena flags the value in a tooltip as being based on pre-release testing.

Do I have to adapt my prompts for Claude Opus 5?

Not necessarily according to the Anthropic documentation, which states explicitly that the model performs well with existing Claude Opus 4.8 prompts. What is recommended, though, is unusual: removing certain verification instructions from old prompts instead of adding new ones, for example double-check your answer, re-verify before responding, or use a subagent to verify. The documentation reasons that Opus 5 verifies its own work and that duplicate instructions to do so only create extra effort.

Can Claude Opus 5 be judged after just five days?

No, and that is exactly my position. Five days is enough for a mood check and not enough for a verdict. The most stable pattern across all sources evaluated here is this: strong on narrowly scoped tasks, mixed on long and open-ended ones. Anyone shaping a final verdict from that right now is selling an opinion as a measurement.

Mijo Jurisic

Google Ads consultant & founder of MJ Marketing. Five-plus years of hands-on practice: from a self-taught start to the Google Premier Partner programme with 500+ direct Google Ads clients and €20M+ in managed media spend.

Share this article:

Share: