Source: ai-research/anthropic-claude-opus-5-announcement-2026-07-24.md — Anthropic’s official announcement (anthropic.com/news/claude-opus-5), fetched 2026-07-24, the day of release.

Claude Opus 5 (claude-opus-5) shipped 2026-07-24 at 25 per MTok — identical to Opus 4.8 — while claiming to come “close to the frontier intelligence of Claude Fable 5 at half the price.” It is the new state-of-the-art on coding and knowledge-work evaluations (Frontier-Bench, GDPval-AA), the new default model on Claude Max and the strongest model on Claude Pro. The headline for anyone operating this wiki’s stack is not the benchmark table — it is that the cost/capability tradeoff the wiki has been routing around since Fable 5’s June launch has largely collapsed.

Key Takeaways

  • Price is flat, capability is not. 25 out per MTok, explicitly “the same as Opus 4.8.” Fast mode runs ~2.5× default speed at 2× base price (Claude Platform, and via usage credits in Claude Code) — the same Fast-mode structure Opus 4.8 used.
  • The Fable 5 gap is now small and expensive to close. On CursorBench 3.2 at max effort Opus 5 lands “within 0.5% of Fable 5’s peak score, but at half the cost.” On OSWorld 2.0 it surpasses Fable 5’s best result “at just over a third of the cost.” Fable 5 remains ahead at the very top, but the premium buys much less than it did.
  • Big jump over Opus 4.8 on the agentic benchmarks. Frontier-Bench v0.1: surpasses all other models and “more than doubles Opus 4.8’s performance at a lower cost per task.” ARC-AGI 3: “three times as high as the next-best model.” Zapier AutomationBench: ~1.5× the next-best model’s pass rate at the same cost per task. Domain science also moved — organic chemistry +10.2pp and protein prediction +7.7pp over Opus 4.8.
  • Verification behavior improved — the exact failure family this wiki self-guards against. Anthropic describes Opus 5 as “much stronger at verifying its work and iterating carefully until it succeeds,” and calls it “our safest model yet in terms of avoiding reckless actions that could have hard-to-reverse side effects.” That speaks directly to the fabrication / skipped-verification / overeager-destructive-action cluster documented for Opus 4.8 and Fable 5.
  • Most aligned model Anthropic has shipped, by their own audit. It scores 2.3 on overall misaligned behavior, “the lowest of our recent models” — better Constitution adherence than Opus 4.8, Sonnet 5, or Fable 5, with the lowest deceptive-behavior rates and least susceptibility to being tricked into misuse.
  • Cyber classifiers intervene ~85% less often than Fable 5’s. Still blocks binary-based scanning, pen-testing, and exploit generation — but this is a direct, quantified fix for the over-firing safeguard gate that has repeatedly tripped on this wiki’s security-content ingests. Flagged requests on Claude.ai, Claude Code, and Claude Cowork fall back to Opus 4.8 by default.
  • No data retention requirement. “Consistent with prior Opus models, Opus 5 does not have data retention requirements for general access” — unlike Fable 5’s mandatory 30-day retention on Mythos-class traffic. For client-sensitive or WEO-internal work this is the single most consequential line in the announcement.
  • Still behind Mythos 5 on cyber and bio. Opus 5 “remains behind Mythos 5 in both biology research and offensive cybersecurity,” and is weaker on exploit development (OSS-Fuzz).
  • Two beta releases alongside it. Mid-conversation tool changes on the Claude Platform — swap which tools Claude can use without invalidating the prompt cache. Automatic fallbacks on the API — route safety-classifier-flagged requests on Opus 5 or Fable 5 to another model automatically.

Why it matters for this wiki’s operating guidance

The wiki’s model-routing posture since 2026-06 has been: Fable 5 for the heavy long-horizon tail, Opus 4.8 for routine work, with Fable 5’s 2× price, mandatory 30-day retention, and aggressive cyber gate as the reasons not to default to it. Opus 5 changes three of those inputs at once:^[the routing-implication synthesis is this wiki’s, not Anthropic’s — the announcement makes no recommendation about when to prefer Fable 5]

  • Cost. Near-Fable-5 results at Opus-4.8 prices removes most of the reason to route down to 4.8 for cost.
  • Retention. No general-access retention requirement, so the client-sensitivity constraint that pushed work off Fable 5 does not apply.
  • Safeguard friction. ~85% fewer classifier interventions than Fable 5 makes it far more usable for the security-topic ingests (MITRE ATT&CK, SkillSpector, exploit write-ups) that have tripped the Fable gate before — see Claude Security and the security-guidance plugin, where classifier firing mid-scan is documented behavior.

The residual case for Fable 5 is the genuine top of the capability curve and Mythos-class cyber/bio work; the residual case for Opus 4.8 is now mostly “it is the fallback target,” including for Opus 5’s own flagged requests. Worth re-checking against Cost & Intelligence Levers and Picking the Right Model on the next refresh of either.

The prompt-cache beta is a quiet cost lever. Changing tool definitions mid-conversation has historically invalidated the cache, which is exactly what long agentic sessions with progressive tool disclosure do constantly. Removing that invalidation matters for anyone running the economics in Prompt Caching for Agencies.

Implementation

Tool/Service: Claude Opus 5 (claude-opus-5)

Setup: Available on release day across the Claude API, Claude.ai, Claude Code, Claude Cowork, and the Claude Platform. Default model on Claude Max; strongest available on Claude Pro. Effort settings are configurable — per the announcement, to “optimize for intelligence or conserve tokens for faster and cheaper results” (see Model vs. Effort).

Cost: 25 per MTok output. Fast mode: ~2.5× speed at 2× base price.

Integration notes: If you rely on safety-classifier behavior, note the new API-level automatic fallbacks beta — it applies to Fable 5 as well as Opus 5, so it is worth wiring even if you have not moved off Fable. Enterprises and researchers blocked by cyber safeguards can apply to the Cyber Verification Program.

Try It

  1. Point one routine workload at claude-opus-5 and compare against your Opus 4.8 baseline — same price, so this is a free comparison on cost.
  2. Re-run whichever task previously justified paying for Fable 5. Given “within 0.5% at half the cost” on CursorBench, the burden of proof for staying on Fable has shifted; measure rather than assume.
  3. If you deferred security-topic work because the Fable 5 cyber classifier kept firing, retry it on Opus 5 — ~85% fewer interventions is a large enough delta to change what is practical.
  4. Turn on the automatic fallbacks beta on the API so classifier-flagged requests degrade to another model instead of failing.
  5. If you run long agentic sessions with changing toolsets, test the mid-conversation tool changes beta and measure cache-hit rate before and after.
  6. Re-check any client-sensitive workflow you routed away from Fable 5 for retention reasons — Opus 5 carries no general-access retention requirement.

Open Questions

  • Context window and max output tokens are not stated in the announcement, and no system card was linked at fetch time. Both need confirming against the platform docs before re-baselining token budgets.
  • No ASL level stated. Fable 5’s Mythos-class designation came with specific handling; where Opus 5 sits is unspecified here.
  • No deprecation or migration timeline for Opus 4.8 — which matters given Opus 4.8 is simultaneously being made the fallback target for Opus 5’s flagged requests, so it presumably stays supported.
  • Benchmark numbers are vendor-reported and several are stated as ratios against unnamed “next-best” models rather than absolute scores. No independent replication existed at ingest.
  • Tokenizer. Fable 5 uses the Opus-4.7 tokenizer (~30% more tokens for the same text); whether Opus 5 matches is unstated, so token-budget math should not be assumed to carry over.
  • Claude Opus 4.8 — the direct predecessor at identical pricing, and the default fallback target for Opus 5’s classifier-flagged requests.
  • Claude Fable 5 + Mythos 5 — the Mythos-class flagship Opus 5 is benchmarked against; still ahead at the top, and still ahead on cyber/bio.
  • Claude Sonnet 5 — the tier below, and one of the three models Opus 5 beats on Constitution adherence.
  • Picking the Right Model — the model-selection framework this release most disrupts.
  • Model vs. Effort — the effort-dial guidance that pairs with Opus 5’s configurable effort settings.
  • Cost & Intelligence Levers — the cost/capability synthesis that needs re-deriving against flat-price/higher-capability.
  • Prompt Caching for Agencies — where the mid-conversation-tool-change beta lands economically.
  • Claude Security Plugin — documented behavior where cyber classifiers fire mid-scan; the ~85% reduction directly affects it.
  • Claude AI — topic index.