devtake.dev

Six Claude 5 models in twelve weeks. Fable 5.1 is the first to watermark its output.

Anthropic's Fable 5.1 doubles a science benchmark and cuts cache reads 75%. Independent tests measured tasks costing up to 3x more than Fable 5.

Dieter Morelli · · 14 min read · 12 sources
Anthropic's launch card for Claude Fable 5.1 and Claude Mythos 5.1, white serif text over a pale blue sky with clouds and a daytime moon.
Image: Anthropic · Source

Anthropic shipped Claude Fable 5.1 and Mythos 5.1 on September 1. They’re the fifth and sixth models in a generation that opened twelve weeks earlier. The benchmark numbers went up again. What a finished task costs, and who’s allowed to run one, moved in a less tidy direction.

Anyone running Claude in production now has four live tiers to choose between, five effort levels on the flagship, two vetting programs gating the strongest variant, and a fresh watermark in every reply. Anthropic’s own figures say Fable 5.1 is roughly 25% cheaper than Fable 5 for typical work. Two independent measurements put the per-task bill higher, not lower. Both claims rest on real arithmetic, and reconciling them is most of the work of deciding what to run.

Six models in twelve weeks

The Claude 5 generation opened on June 9, 2026 with Fable 5 and Mythos 5. Sonnet 5 arrived June 30, Opus 5 on July 24, and the 5.1 refresh of both top-tier models on September 1, per Hidekazu Konishi’s release timeline. No Haiku 5 exists yet.

ModelReleasedWho can run it
Fable 5June 9General availability
Mythos 5June 9Project Glasswing participants
Sonnet 5June 30Default on Free and Pro
Opus 5July 24Default on Max, top model on Pro
Fable 5.1September 1General availability, Max and credits
Mythos 5.1September 1Vetted US organizations only

The older tier names were all forms of writing. Haiku, Sonnet, Opus. Mythos breaks that pattern because it isn’t a form of writing at all. It’s a capability class, and Fable is the model that lives inside it, roughly the way a product lives inside a product line. Developers Digest read the split as deliberate: separating the class name from the model name means Anthropic can ship a second Mythos-class model without having to call it Fable 5.2.

Mythos 5.1 and Fable 5.1 are the same underlying model with different safeguards. Anthropic says so directly. The Mythos variant runs with cybersecurity and life-sciences restrictions lifted, and it reaches customers through two vetting programs rather than the open API. Same weights, different guardrails, different paperwork. Anthropic started routing Mythos 5 to cyberdefenders in June, the same day Fable 5 launched as the first publicly available Mythos-class model.

That generation has not been a smooth run. Fable 5 shipped June 9, got suspended June 12 under US export controls, and came back on July 1 with a new safety classifier after Amazon researchers found a jailbreak that got the model to identify and then exploit a specific vulnerability. Twelve weeks, six models, one full withdrawal and redeployment.

What 5.1 actually changed

The headline gain is on scientific agentic work. Fable 5.1 scores 52.6% on Terminal-Bench-Science 0.1 against 24.7% for Fable 5, 29.0% for Opus 5 and 22.4% for OpenAI’s GPT-5.6 Sol. That’s a doubling over its own predecessor on a benchmark three months old. The rest of the table is a steadier climb.

BenchmarkFable 5.1Fable 5Opus 5GPT-5.6 Sol
Terminal-Bench-Science 0.152.6%24.7%29.0%22.4%
Terminal-Bench 4.055.8%42.0%52.3%37.3%
AutomationBench31.4%17.1%26.9%19.6%
GDPval-AA v21853172318241711
OSWorld 2.0 (strict)41.7%36.1%39.6%not tested
CursorBench 3.2.073.4%70.5%70.0%67.2%
Humanity’s Last Exam (no tools)60.9%57.8%56.6%not tested

Mythos 5.1, the restricted twin, scores higher still on the one benchmark Anthropic published for it: 60.9% on Terminal-Bench 4.0 against Fable 5.1’s 55.8%.

Read that column-by-column and the pattern is clear. Terminal-Bench-Science and AutomationBench roughly double. Everything else moves two to four points. If your workload looks like long-running agentic research, 5.1 is a different model. If it looks like ordinary coding, it’s an increment on Opus 5, which is itself only a few points back at half the sticker price.

Anthropic backed the launch with named customers, and those accounts are worth reading with the usual discount applied. A vendor picks its testimonials. Craig Falls, head of quantitative research at Jane Street, said “Claude Fable 5.1 solves more of our coding problems than Fable 5 or Opus 5.” Cognition co-founder Walden Yan went further: “We’re moving our Opus 5 traffic in Devin to Claude Fable 5.1 on launch day.” A senior portfolio manager at Millennium described the model disassembling an external vendor library, matching it against a core dump, and tracing a rare crash that had gone uninvestigated for four to five years. Red Hat distinguished engineer Josh Boyer said it “correctly identified the root cause of every broken build we tested.” Browserbase reported 82% task completion on its hardest benchmark against 74% for Opus 5.

The pattern in those accounts is consistent, and it’s the same pattern the benchmarks show. Every example is a long, messy, multi-step investigation rather than a quick edit. Nobody’s testimonial is about writing a function faster.

The other real change is that the model refuses less. Anthropic says its newest cybersecurity safeguards “block 60% fewer false positives than before,” which translates to roughly 60% fewer interventions per Claude Code session. Biology safeguards fire about 85% less often on benign elementary biology and medical questions. Fable 5.1 is now permitted to discover software vulnerabilities, though still not to develop exploits for them, and life-sciences R&D requests get redirected to Opus models rather than blocked outright. TechCrunch framed the release as “cheaper, less restrictive,” which is accurate and also the whole tension of the launch in three words.

The science claims, and who checks them

Anthropic put three research results in the announcement, and they’re the boldest thing in it. Two of the three came from Mythos 5.1, the variant almost nobody can run.

The protein-design result is the one with independent verification. Adaptyv Bio tested Claude’s designs in its automated wet lab, anonymized, measuring binding affinity by surface plasmon resonance across five target concentrations in duplicate. Of 1,320 designs, 354 bound their targets, a 26.8% overall hit rate, with binding achieved on 14 of 15 targets. On 15-PGDH the best Claude binder came in at 33.4 nM against a previous competition best of 1.7 µM. On RBX1 it hit 3.9 nM against 25.7 nM. Anthropic’s summary of the same work reports affinities 10 times higher than the best competition submissions on three targets, and a hit rate near 50% across 12 targets where 10% to 15% is typical.

Nine orange ribbon diagrams of protein structures arranged to spell the word ANTHROPIC, each labelled with its target protein and binding affinity.

Nine of Claude’s experimentally confirmed de novo binders, each labelled with its target and measured affinity. TREM2 at 4 nM and IL-7Ralpha at 3.6 nM sit at the strong end, EGFR at 1.7 µM at the weak end. Image: Anthropic. Source

Adaptyv is careful about what that means. These were “not real therapeutics,” only initial binding validation. One target produced no binders at all, and another was thrown out for low-quality measurements because the protein aggregated. Claude couldn’t beat the best binder on the Nipah virus target.

Anthropic figure card for the Nipah G target, showing an orange two-helix designed binder docked against a pale grey target protein surface, labelled with a 20 angstrom scale bar.

Nipah G, the target where the tension shows. Claude hit 18 of 30 designs, a 60% rate, and still didn’t beat the best binder from Adaptyv’s public competition. Image: Anthropic. Source

The bigger caveat is methodological. The whole thing ran open-loop, with no iterative learning from results, which is how a human protein engineer would actually work. Those caveats come from the lab that ran the assays, not from a critic.

The other two results have no outside check yet. Fable 5.1 trained a neural network to build an elevation map covering a third of Venus, using NASA Magellan radar data that’s been sitting around for 30 years. The new map resolves features at two to three kilometers instead of 10 to 20, with heights up to 25% more accurate. Anthropic is releasing it under a Creative Commons license ahead of the NASA VERITAS and ESA EnVision missions.

Mythos 5.1 also sped up seven open-source deep learning models by as much as 2.5x and cut estimated GPU costs 30% to 60%. Both results are checkable in principle. Neither has been checked in public.

What it costs, and who gets it

The commercial headline is a 75% cut to cache reads, from $1.00 to $0.25 per million tokens. Base rates hold at $10 in and $50 out. Anthropic puts the practical saving at around 25% for typical workloads and up to 45% for heavily agentic ones, which is a real change to how prompt caching economics work on long agent runs.

Then the output side moved the other way. Artificial Analysis measured Fable 5.1 using roughly 1.7x the output tokens of Fable 5 on the same evaluations, and calculated a total net per-task cost increase of 20% once the cache savings were netted off. Their run also flagged a measurement wrinkle worth knowing about: about 4% of output tokens across the evals came from fallback models rather than Fable 5.1 itself, which muddies attribution on any benchmark where safeguards trigger a handoff. Latent.Space’s roundup collected both figures on launch day.

An independent benchmark run found the same direction and a steeper slope. MineBench put both models through 15 Minecraft builds at maximum thinking effort. Fable 5.1 cost $147.55 against Fable 5’s $54.93, close to 3x, with average inference time going from 18 minutes 4 seconds to 40 minutes 12 seconds. Output size barely budged, 34.07 MiB average against 30.65 MiB, so the extra money bought thinking rather than artifacts. As MineBench’s maintainer put it on r/ClaudeAI, “Despite no change in API pricing, Fable 5.1 was nearly 3x as expensive as Fable 5 on MineBench.”

Failures compound that. Eight of 23 attempts died in the harness, which works out to $9.84 per finished build. On a reasoning-heavy model a dead run is close to a total write-off, because you pay for the thinking either way.

Access tightened alongside. Fable 5.1 is on Max plans and credits, not on Pro or Free, and the r/ClaudeAI thread asking Anthropic to put Fable back on Pro drew 488 upvotes in two days. Mythos 5.1 sits behind the Cyber Verification Program and the Life Sciences Verification Program, the latter run in partnership with the US government, and it’s US-only for now.

Enterprises have been voting on all of this with their budgets, and the answer so far is restrained. Ramp’s Economics Lab put Anthropic at 43.5% of businesses with paid AI subscriptions in July, ahead of OpenAI’s 39.7%. Fable 5, a month after launch, accounted for 6% of Anthropic tokens and 11.4% of Anthropic spend. GPT-5.6 Sol took 25% of OpenAI’s tokens over the same period. Ramp’s read is that businesses are hitting a ceiling on AI spend, and that Fable 5 marked the upper bound of what they’ll pay. Anthropic is building an in-house chip team partly to push that bound down from the supply side.

The effort dial is the real lever

Fable 5.1 exposes five reasoning effort levels: low, medium, high, xhigh and max. There’s no way to turn reasoning off. That single setting now swings cost more than any pricing change Anthropic announced.

Simon Willison ran the same prompt across all five and published the numbers. At low and medium, his SVG request produced no visible reasoning at all and landed near 1,980 output tokens. High produced 2,612 tokens. Then xhigh jumped to 36,767 tokens over 7 minutes 51 seconds for $1.83, and max reached 65,927 tokens in nearly 14 minutes for $3.30. Same prompt, same model, a 25x spread in output tokens between high and max, and 33x measuring up from medium.

Willison rated the max result “the best pelican I’ve seen from any of Anthropic’s models,” while noting he “didn’t ask for flair” and that his pelican test’s correlation with real task performance “didn’t seem to hold as strongly” this year. He finds the within-family effort comparison more useful than the cross-model one, which is the right instinct here.

Anthropic’s own claim is that 5.1 at lower effort matches or beats Fable 5 at high effort, and that’s the part most people are getting wrong. The defaults do not split the difference for you. Claude Code defaults to high, while Claude Cowork and Claude.ai default to medium. Move a Claude Code workload from Fable 5 to Fable 5.1 and leave the effort setting alone, and you’ve bought a longer, more expensive run for a modest quality gain. That’s the 20% net increase Artificial Analysis measured.

Every output is watermarked now

Fable 5.1 and Mythos 5.1 are the first Claude models to embed an invisible watermark in generated text. Anthropic committed to marking any model launched on or after August 2, 2026, tracking the transparency requirements in Article 50 of the EU AI Act, and it applies the marks “wherever Claude is offered, worldwide” rather than only in Europe.

The mechanism biases word selection when several plausible options exist, so the mark rides along in the text itself. Anthropic says it carries no information about the user, the organization or the conversation, and doesn’t change response quality. Generated files in formats like .svg, .png and .jpg get signed C2PA provenance metadata instead. There’s no opt-out.

The honesty about its limits is the useful part. Anthropic’s own documentation states that “a detected mark provides a signal that content was processed by Claude, but is not fully conclusive,” and that a “lack of a detected mark doesn’t mean the content wasn’t AI-generated or processed.” False negatives show up when text is heavily edited, paraphrased, translated or simply very short. Detection stays in private preview for regulators, law enforcement, media, fact-checkers and researchers, plus enterprises with compliance obligations. Everyone else gets a file-credential checker.

Running alongside it is a new option called Enterprise Frontier Safeguards, rolling out this autumn across Claude Code, Claude Enterprise, the APIs and the major clouds. It keeps safety data on infrastructure the customer controls, under the customer’s own encryption keys, so an organization can hold zero data retention and still let Anthropic’s misuse detection run. Anthropic paired the announcement with a flat commitment: it “has never trained on enterprise data without explicit permission, and never will.”

The quieter change in the same release will break more code. API accounts created after August 31, 2026 at 00:00 UTC can no longer edit Claude’s prior context in a multi-turn conversation while keeping the transcript of its earlier thinking. That combination was a documented route to extracting Claude’s reasoning for distillation, and Anthropic closed it. The API now validates that thinking blocks come back with identical system prompts, tools and preceding messages, and raises an error if any of them changed. You can opt into a non-strict mode where the request succeeds and the affected thinking blocks get dropped from what the model sees, with the response telling you which ones went.

The blast radius is narrow today and won’t stay that way. It covers new Claude Platform, Bedrock, Vertex AI and Azure Foundry accounts, applies only to Fable 5.1, and leaves Claude Code, Claude Cowork and Claude.ai alone. Anthropic says future releases will extend it to everyone.

The risk rating Anthropic moved on itself

The September 1 system card contains the most newsworthy sentence of the launch, and it isn’t about capability. Anthropic now assesses the risk of catastrophic harm from alignment failures as low rather than very low, citing its own increased uncertainty following recent incident disclosures about model behavior.

Companies don’t usually revise their own safety ratings downward in a launch document. One notch is still a long way from alarming, and the same card keeps automated AI R&D risk at low on the grounds that the model stays well below Anthropic’s own researchers and engineers, a finding METR’s external testing supported. But the direction of travel is worth marking, because every other number in the announcement moved up.

The card also says Fable 5.1 and Mythos 5.1 have the strongest cyber capabilities of any model Anthropic has released, meeting or exceeding Mythos 5. And it notes that Mythos 5.1 “cooperates with human misuse and accepts unverifiable claims of authorization somewhat more readily than Opus 5.” Set that next to the 60% reduction in cybersecurity false positives and the newly permitted vulnerability discovery, and you can see the trade Anthropic made: fewer wrong refusals, a model that pushes back less on who you claim to be. Knowledge cutoff for both models is June 2026.

What I’d actually set

Anthropic’s own guidance is that Fable 5.1 at low or medium effort lands near Fable 5 at high. That single claim, if it holds for your workload, is worth more than the 75% cache discount. Here’s the order I’d work through it in.

  • Don’t lift and shift. Moving a Fable 5 workload to Fable 5.1 without touching the effort setting is the specific mistake that produces a 20% to 200% cost increase for a few benchmark points.
  • Try medium first. Anthropic’s claim that 5.1 at low or medium matches Fable 5 at high is the cheapest version of this upgrade, and it’s the one nobody tests because the defaults hide it.
  • Reserve max for research. Terminal-Bench-Science doubled and AutomationBench nearly doubled. Ordinary coding gained two to four points, which max effort will not pay back.
  • Budget for failures. A third of MineBench’s runs died in the harness. On a reasoning-heavy model a dead run costs nearly as much as a live one, so pad the estimate rather than the retry logic.
  • Check your API account date. If you opened one after August 31 and you edit context between turns, you’ll hit the new thinking-block validation error before you hit anything else.

What’s still open

Nobody outside Anthropic has verified the Venus map or the GPU speedups, and both are the kind of claim an independent group could check in a few weeks. The protein work already has a wet lab behind it, and Adaptyv’s own caveat about open-loop design is the honest limit on how far to read it. There’s still no Haiku 5, which leaves a gap at the cheap end of a lineup where every tier now carries a 1M-token context window.

The watermark is the thing I’d watch through the autumn. Anthropic is the first major lab to ship marked text output globally under Article 50, its detection API is gated, and its own documentation says a negative result proves nothing. If a regulator or a newsroom tries to use it in an actual dispute this year, we’ll learn quickly whether statistical watermarking survives contact with people who want to remove it.

Share this article

Quick reference

prompt caching
A server-side cache of the tokens already processed in a conversation, so repeated context is not re-billed or recomputed on each new turn. Changing the tool set usually invalidates it.
reasoning effort
A per-request setting for how many tokens a model spends thinking before it answers. Higher levels raise answer quality and cost together.
watermark
Statistical marks embedded in generated text by biasing word choice among plausible options. Invisible to readers, it survives copy-paste but not heavy paraphrasing.

Sources

Mentioned in this article