On 1 September 2026 Anthropic released Claude Fable 5.1, along with a twin called Mythos 5.1 that almost nobody is allowed to use. The announcement carries a table with nine benchmarks, and the new model leads in all nine rows. That table is the reason the model is currently being called the new number one everywhere. It is also the reason a second look pays off.
This article keeps three classes of statement apart: vendor-reported figures from Anthropic's own post and product documentation, independent measurements from Artificial Analysis and the public Terminal-Bench leaderboard, and my own assessment, which I label as such. The video version of this analysis sits at the top of this page and on the video page. Which AI tools I actually use in my own work is documented openly on the AI in numbers page. All figures in this article are as of 8 September 2026.
The facts from the announcement
First the verifiable part, all of it from Anthropic's announcement and from the Claude platform model overview.
| Attribute | Claude Fable 5.1 |
|---|---|
| Released | 1 September 2026 |
| Model ID | claude-fable-5-1 |
| Context window | 1 million tokens |
| Maximum output | 128,000 tokens |
| Knowledge cutoff | June 2026, labelled "reliable knowledge cutoff" by Anthropic |
| Platforms | Claude API, Amazon Web Services, Google Cloud, Microsoft Azure |
| In claude.ai | Max and Enterprise, on Pro only via usage credits, not in the Team plan |
Two points from that table get overlooked in most reports. The one-million-token context window is shared with Opus 5 and Sonnet 5, so it is not a differentiator. And inside the Claude app the model is far less available than the headlines suggest: the Team plan does not carry it at all, and on Pro it only runs through usage credits.
Anthropic's table: nine rows, nine wins
The values below are vendor-reported, measured by Anthropic itself. The only competitor listed is GPT-5.6 Sol. GPT-6 Astra is missing, which is plausible because it only appeared two days after Fable 5.1.
| Benchmark | Fable 5.1 | Fable 5 | Opus 5 | GPT-5.6 Sol |
|---|---|---|---|---|
| Terminal-Bench-Science 0.1 | 52.6 % | 24.7 % | 29.0 % | 22.4 % |
| Terminal-Bench 4.0 | 55.8 % | 42.0 % | 52.3 % | 37.3 % |
| GDPval-AA v2 (Elo) | 1853 | 1723 | 1824 | 1711 |
| OSWorld 2.0, partial | 77.9 % | 72.9 % | 75.4 % | no value |
| OSWorld 2.0, strict | 41.7 % | 36.1 % | 39.6 % | no value |
| Humanity's Last Exam, no tools | 60.9 % | 57.8 % | 56.6 % | no value |
| Humanity's Last Exam, with tools | 65.0 % | 63.8 % | 63.6 % | no value |
| AutomationBench | 31.4 % | 17.1 % | 26.9 % | 19.6 % |
| CursorBench 3.2.0 | 73.4 % | 70.5 % | 70.0 % | 67.2 % |
Table scrolls sideways
Anthropic itself notes two caveats below the table. First, Fable 5.1 ran with production safeguards enabled; wherever those safeguards intervened the task scored zero, which according to Anthropic likely depresses its own OSWorld 2.0 and AutomationBench results. Second, the OSWorld figures rest on the August 2026 task release, which is why no competitor score is shown there at all. That is disclosed and fair. It does not change the fact that all nine rows come from a single setup.
Terminal-Bench-Science: the number everyone quotes
The headline from that table is 52.6 percent on Terminal-Bench-Science 0.1, more than double the predecessor at 24.7 percent. The benchmark measures whether an AI agent can do real research work in a terminal: analysing data, running simulations, fitting models. It is run by the Terminal-Bench project, and the leaderboard is public.
That is precisely where Fable 5.1 does not appear as of 8 September 2026. Listed there are Opus 5 at 30.0 percent, GPT-5.6 Sol at 22.4 and Fable 5 at 21.4 percent. Anthropic discloses this in the footnote to its table, quotes both leaderboard values, puts its own reproduction next to them (29.0 for Opus 5 and 24.7 for Fable 5) and states a standard error of 3.5 to 4.5 percentage points per model.
That is clean work, and it does not make the 52.6 percent wrong. It makes them a vendor-reported figure from an in-house setup rather than an independently reproduced result. As a side note, because it shows the pattern: OpenAI reports 64.6 percent for GPT-6 Astra on the same test against 52.6 for Fable 5.1. That too is a vendor-reported figure, just from the other house.
The independent cross-check: Terminal-Bench 4.0
There is one benchmark where vendor claim and independent measurement sit side by side: Terminal-Bench 4.0, agentic coding in the terminal. These are the public leaderboard values, retrieved on 8 September 2026.
| Model (agent) | Resolution rate | Tokens | Cost per run |
|---|---|---|---|
| GPT-6 Astra (Codex) | 58.2 % (± 2.8) | 1.5 bn | around 3,300 dollars |
| Claude Fable 5.1 (Claude Code) | 57.9 % (± 3.8) | 2.7 bn | around 6,200 dollars |
| Claude Opus 5 (Claude Code) | 51.8 % (± 3.4) | 6.5 bn | around 6,000 dollars |
| Claude Fable 5 (Claude Code) | 44.5 % (± 3.8) | 3.8 bn | around 7,300 dollars |
Table scrolls sideways
Two findings sit in there. First, Anthropic's numbers survive the cross-check: the vendor reports 55.8 percent for Fable 5.1, the measurement says 57.9. For Opus 5 Anthropic reports 52.3 against a measured 51.8, for Fable 5 42.0 against 44.5, and for GPT-5.6 Sol both say 37.3. Everything sits inside the error bars.
Second, Fable 5.1 is not in first place here. GPT-6 Astra leads at 58.2 percent, using roughly half as many tokens and about half the money. The 0.3 percentage point gap is effectively a tie, the cost gap is not.
Artificial Analysis: level at the top, split at the checkout
The second independent source is Artificial Analysis. In version 4.3 their Intelligence Index combines ten evaluations, among them Terminal-Bench v4.0, SciCode, Humanity's Last Exam and GDPval-AA v2, measured on dedicated hardware.
| Model (variant) | Intelligence Index v4.3 | Output speed (tokens/s) | Cost per index task |
|---|---|---|---|
| Claude Fable 5.1 (max with fallback) | 53.37 | 69.54 | 7.63 dollars |
| GPT-6 Astra (max) | 52.81 | 61.52 | 3.26 dollars |
| Claude Opus 5 (max) | 50.70 | 54.26 | 5.86 dollars |
| Claude Fable 5 (with fallback) | 49.70 | 61.66 | 8.75 dollars |
Table scrolls sideways
Fable 5.1 therefore holds first place, by a margin of 0.56 index points. In the Artificial Analysis chart both leading models round to 53.
Two notes belong with this, otherwise you end up quoting the wrong number. First, the index has been rebuilt twice since launch, to v4.2 on 4 September and to v4.3 on 7 September, among other things moving to Terminal-Bench v4 and adding AutomationBench-AA. The figure 66 that appears in many reports from 1 September comes from the older index version and is not comparable to today's. Second, Artificial Analysis states itself that it supported Anthropic with pre-release evaluation. These are their own runs on their own hardware, but fully independent is something else.
Why a single task costs that much
The most interesting column is not the index, it is the price per task. One index task averages 7.63 dollars on Fable 5.1, 5.86 on Opus 5 and 3.26 on GPT-6 Astra. Among the current frontier models, Fable 5.1 is the most expensive route to a result. With one caveat that belongs in the sentence: its own predecessor Fable 5 sits higher still in the same measurement at 8.75 dollars. So Fable 5.1 is expensive, but cheaper than the model it replaces.
The reason is verbosity. For the index run Fable 5.1 generated 190 million tokens according to Artificial Analysis, against a median of 88 million for comparable models. The full run cost 13,128.86 dollars. Artificial Analysis summarises the model itself as among the leading models in intelligence, particularly expensive compared to other models of similar price, slower than average and very verbose.
The Mythos 5.1 twin, explained the right way round
The naming suggests Mythos is the more heavily guarded variant. It is the other way round. Anthropic writes that both are the same model with different levels of safeguards. Fable 5.1 is the generally available version with the stricter standard safeguards. Mythos 5.1 carries the more permissive ones, and access goes only to vetted parties: cyberdefenders through the Cyber Verification Program and life scientists through the Life Sciences Verification Program, currently US organisations only.
You can see the difference in Anthropic's own table: on Terminal-Bench 4.0 the company reports 60.9 percent for Mythos 5.1 against 55.8 for Fable 5.1. Anthropic explains the gap with cyber safeguards that intervened on some Fable tasks.
Practically relevant for development work: Fable 5.1 is now allowed to identify software vulnerabilities, that is, to do defensive security work. Penetration testing, exploit generation and binary-based vulnerability scanning are still routed to the Opus models, as is research and development in the life sciences. Anthropic puts the effect of the new, more precise rules at around 60 percent fewer interventions per Claude Code session. That too is a vendor-reported figure.
Pricing: one single line changed
All values per million tokens, vendor list prices, as of 8 September 2026.
| Model | Input | Output | Cache read |
|---|---|---|---|
| Claude Fable 5.1 | 10 dollars | 50 dollars | 0.25 dollars |
| Claude Fable 5 | 10 dollars | 50 dollars | 1.00 dollars |
| Claude Opus 5 | 5 dollars | 25 dollars | 0.50 dollars |
| GPT-6 Astra | 10 dollars | 50 dollars | 1.00 dollars |
Table scrolls sideways
List pricing for input and output is unchanged against Fable 5 and still twice that of Opus 5. The only real change sits in the cache column. When the model reads context it has already processed, that used to cost one dollar per million tokens. Now it costs 25 cents, 75 percent less, and for the first time less than Opus 5 at 50 cents. GPT-6 Astra carries exactly the same list prices of 10 and 50 dollars but still charges one dollar for cache reads.
Why that is the most important line: an agent working for hours on the same task reads the same context over and over. On runs like that the bill is mostly cache reads. Anthropic calculates that typical workloads therefore come out around 25 percent cheaper than on Fable 5, and heavily agentic ones up to 45 percent. Important context: those percentages are a model calculation on Anthropic's own usage mix from August 2026, not a discount that shows up on your invoice. The hard fact is the cache price, one dollar down to 25 cents. Anyone able to batch tasks gets another 50 percent off input and output through batch processing.
Looking back: what became of the leak claims
At the end of July I checked the leak claims about Claude Fable 5.1 that were circulating at the time. The finding was: no artifact, no model ID, no entry in the documentation, so it stays a claim. The artifact is on the table now, so the accounts can be settled.
The name Fable 5.1 was right. The date was not: the announced August turned into 1 September. The claim that prices stay unchanged was half right, because list pricing did hold, but the actual price change, the cache discount, appeared in none of the circulating claims. And the assumption that Anthropic was holding the model back for the next OpenAI release? Fable 5.1 shipped on 1 September, GPT-6 Astra on the 3rd. The timing pattern fits, proof of a causal link it is not, and Anthropic has never commented on it.
The lesson is the part I find most interesting: the rumour mill half guessed the name and completely missed the only change that matters in practice.
Who Fable 5.1 is worth it for, and when Opus is enough
What follows is my assessment based on the documented figures above and my own work with these models, not an external fact.
Anthropic phrases the recommendation in its own model overview like this: if you are unsure, start with Claude Opus 5; Fable 5.1 is meant for demanding reasoning and long-horizon agentic work, or for cases where your evals on Opus 5 at higher effort still fall short. From my own practice that matches what I see. For copy, analysis and routine automation, Opus 5 at half the token price is the sensible choice, and I wrote up its benchmark numbers in the Claude Opus 5 analysis. You reach for Fable 5.1 when an agent works for hours on the same large context, because that is where the new cache price takes effect. Whether Claude Opus 5.5 actually delivers the Fable 5.1 level Anthropic promises at lower cost is what I examine in the Claude Opus 5.5 check.
One detail is missing from almost every calculation: the effort level. Fable 5.1 has five of them, and according to the Artificial Analysis evaluation the gap between the lowest and the highest is elevenfold in output tokens. Set that level deliberately, or you pay maximum effort for tasks that do not need it. If you want to sort out which models fit your tasks, your data protection requirements and your budget, that is exactly the subject of my AI consulting.
Sources
- Anthropic, announcement "Introducing Claude Fable 5.1 and Claude Mythos 5.1", 1 September 2026: anthropic.com
- Claude platform docs, model overview with model IDs, context windows and knowledge cutoffs: platform.claude.com
- Claude platform docs, pricing overview including cache and batch: platform.claude.com
- claude.com, plans and model availability, retrieved 8 September 2026: claude.com/pricing
- Anthropic, system card for Fable 5.1 and Mythos 5.1: anthropic.com
- Artificial Analysis, model page for Claude Fable 5.1, Intelligence Index v4.3, retrieved 8 September 2026: artificialanalysis.ai
- Artificial Analysis, article of 1 September 2026, index version before v4.2: artificialanalysis.ai
- Terminal-Bench 4.0, public leaderboard, retrieved 8 September 2026: tbench.ai
- Terminal-Bench-Science 0.1, public leaderboard, retrieved 8 September 2026: tbench.ai
- OpenAI, GPT-6 Astra announcement with its own comparison against Fable 5.1: openai.com
- OpenAI, API pricing for GPT-6 Astra: developers.openai.com
Figures as of: 8 September 2026




