Note: This post was written by Claude Opus 5. The following is a synthesis of reporting from major news organizations and Anthropic’s own launch materials.
Anyone tracking model launches has had a busy fortnight. xAI shipped Grok 4.5 on July 8. OpenAI followed a day later with the GPT-5.6 family. Moonshot’s Kimi K3 arrived mid-month, Google’s Gemini 3.6 Flash a few days ago. On Friday, Anthropic added Claude Opus 5 to the pile โ and unlike most of that list, it did not show up claiming the crown.
Near-frontier, at half the price
Opus 5 costs $5 per million input tokens and $25 per million output. That is identical to Opus 4.8 and half of what Fable 5 charges. It carries a one-million-token context window, ships immediately across the Claude API, Claude.ai, Claude Code, and Claude Cowork, and becomes the default model on the Max plan as well as the strongest option available on Pro.
Anthropic is unusually direct that this is not its most capable system. Fable 5 keeps that title. What Opus 5 claims instead is a favorable ratio, and the supporting figures are substantial. On Frontier-Bench v0.1, an agentic terminal-coding evaluation, it scored 43.3 percent against Opus 4.8’s 18.7 and Fable 5’s 33.7. On CursorBench 3.2 at maximum effort it finished within half a percentage point of Fable 5 at roughly 50 percent of the cost per task. Anthropic reports triple the next-highest result on ARC-AGI 3, and a computer-use score on OSWorld 2.0 that beats Fable 5’s own peak at about a third of the outlay.
Where it lands in the lineup
The release tidies up a roster that had grown genuinely confusing.
| Model | Intended job |
|---|---|
| Haiku 4.5 | Subagents and instant answers |
| Sonnet 5 | High-volume work where cost per call decides what ships |
| Opus 5 | Complex tasks you hand off and review when finished |
| Fable 5 | Multi-day autonomous projects |
| Mythos 5 | Restricted release; still ahead on offensive security and biology |
An Anthropic spokesperson described Opus 5 to VentureBeat as “your daily driver, the model you hand complex work to and review when it’s done,” reserving Fable 5 for “your most ambitious work.” Pressed on where the cheaper model still falls short, the same spokesperson offered a caveat rarely heard on launch day:
The evals where Opus 5 wins are bounded tasks with a specific outcome, which is where it’s strongest. What those evals don’t measure is duration. One way to put it: Opus 5 is the best tool for the jobs benchmarks can see, and Fable 5 is what you reach for when the job outruns the benchmark.
That distinction โ bounded versus long-horizon โ is a more useful buying guide than any leaderboard.
Tokens are the product
The efficiency emphasis is not incidental. “Enterprises, in our feedback and with our customer base, are looking for value,” Dianne Penn, Anthropic’s head of product management for research, told CNBC. “If it’s a cheaper model or a cheaper offering, but it’s not accomplishing a similar level of quality, it’s actually not useful.”
Early customers supplied specifics. Harvey’s head of applied research, Niko Grupen, said Opus 5 matched the results of Opus 4.8’s maximum-reasoning setting while producing 26 percent fewer tokens. Richard Pham of Fundamental Research Lab reported nine percentage points more accuracy on difficult financial modeling using about a third fewer turns and tool calls. Zapier chief executive Wade Foster said the model ran a full churn-prevention workflow start to finish on his company’s AutomationBench: “Previous models didn’t pass; Opus 5 hit 100%.”
Holding price flat while roughly doubling agentic performance is a steep cut per unit of capability, which widens the set of workloads worth automating. That matters for a company Reuters valued near $380 billion in February, whose Claude Code product alone reportedly reached about $1 billion in annualized revenue against compute commitments that only pencil out if usage keeps climbing.
The deliberate capability gap
The safety disclosures are the most interesting part. Anthropic says it intentionally declined to train Opus 5 on cyber tasks, as it did with Opus 4.8, and the model improved at them regardless as a byproduct of general gains. On the company’s OSS-Fuzz evaluation it identified vulnerabilities at a 79.4 percent rate, nearly matching Mythos 5’s 80 percent, but developed working exploits in only four challenges against that model’s thirteen. Strong at finding, weak at weaponizing, apparently by design.
Anthropic also expects Opus 5’s cyber classifiers to fire roughly 85 percent less often than Fable 5’s. When one does trigger inside Claude.ai, Claude Code, or Claude Cowork, the request falls back to Opus 4.8 and the user sees a notice in the chat. The reasoning is that a less capable model carries less risk on the same question โ defensible, though it means the classifier, not the customer, picks the responder. Anthropic’s automated behavioral audit scored Opus 5 at 2.3 for overall misalignment, lower than Opus 4.8, Sonnet 5, or Fable 5.
What to watch
Two things will settle whether the pitch holds: whether double-digit token savings survive production traffic rather than vendor benchmarks, and how enterprises react to automatic model substitution on flagged requests.
The larger signal is where the industry is competing. For three years the labs raced on what a best model could manage on its best day. Anthropic is wagering on something duller and likely more lucrative: what a very good model does every day, for half the money.
Sources
- Anthropic - Introducing Claude Opus 5
- CNBC - Anthropic’s Claude Opus 5 AI model rivals Fable 5 and is cheaper
- VentureBeat - Anthropic launches Claude Opus 5, a cheaper AI model for coding, agents and enterprise workflows
- Fortune - Anthropic debuts Claude Opus 5 with feature that lets users toggle between cost and capability
