Anthropic released Claude Opus 5 on Friday, positioning its new flagship as a model that approaches the intelligence of the company's frontier system, Fable 5, at roughly half the price. It is available immediately across Anthropic's platforms and priced identically to its predecessor, Opus 4.8 — $5 per million input tokens and $25 per million output tokens — a pitch built around holding cost flat while raising performance. Opus 5 becomes the default model on the Claude Max subscription and the strongest option on Claude Pro.
The framing matters as much as the numbers. Rather than chase an outright performance crown, Anthropic is selling Opus 5 as the model to reach for every day: near-frontier quality, but efficient enough to run at scale without the token bill of a top-tier system.
What the benchmarks show
By Anthropic's own measurements, the generational jump over Opus 4.8 is substantial. On Frontier-Bench v0.1, a test of agentic terminal coding, Opus 5 scores 43.3% — more than double Opus 4.8's 21.1% and ahead of both Fable 5 (33.7%) and OpenAI's GPT-5.6 Sol (34.4%). It leads the company's knowledge-work index GDPval-AA v2 (1861, against 1593 for Opus 4.8), and posts an outsized result on ARC-AGI-3, a novel-problem-solving test, where its 30.2% dwarfs the 7.8% of the next-best model. On the computer-use benchmark OSWorld 2.0 it reaches 70.6%, and on Zapier's AutomationBench, which grades whether a model can carry a business task from start to finish, it hits 26.0% — comfortably above every rival listed.
It is not a clean sweep, and to Anthropic's credit the company's own comparison table shows where Opus 5 falls short. GPT-5.6 Sol leads on the DeepSWE v1.1 coding benchmark (72.7% to Opus 5's 68.8%). Fable 5 narrowly edges it on Humanity's Last Exam without tools and on a held-out legal-agent test. And on health and on the "human-solved" tier of a biology benchmark, the specialized Mythos 5 remains ahead. The picture is of a model that tops most categories decisively while conceding a handful to more specialized or more expensive systems — which is roughly what the "near-frontier at lower cost" positioning would predict.
As with any launch, these are vendor-reported figures generated on the company's own harnesses, and independent replication will take time. The comparisons are most useful as a statement of where Anthropic believes its model stands, not as settled fact.
Agentic behavior is the selling point
The through-line in Anthropic's pitch is not raw scores but judgment — the model's willingness to verify its own work and keep going until a task actually succeeds. The company highlights cases from testing: on one Frontier-Bench task, given a drawing it was deliberately prevented from viewing, Opus 5 wrote its own computer-vision pipeline to extract the geometry from raw pixels and rebuild a 3D model, something no competing model managed in five attempts. In another, handed a real bug in an open-source package manager, it traced the root cause and fixed an edge case the community's own patch had missed, where a rival model addressed only the surface symptom and declared victory. A trading-firm engineer described using it to build a market-data feed in a single session, with the model constructing its own test harness when no live feed was available to check against.
Early-access customers quoted by Anthropic — including Cognition, Cursor, Zapier and others — echo the theme, describing gains concentrated on longer, vaguer, multi-step work and, notably, steadier results from run to run. Those are testimonials solicited for a launch, so they warrant the usual skepticism, but the consistency of the "judgment and reliability" framing across them is the clearest signal of what Anthropic thinks it has built.
Safety and the deliberate limits
Anthropic pairs the release with an unusually prominent safety argument. It says an automated behavioral audit rated Opus 5 its most aligned model to date, with the lowest measured rate of deceptive behavior and the best adherence to the company's published model constitution — a score of 2.3 on overall misaligned behavior, the lowest among its recent models. Again, that audit is Anthropic's own.
More concretely, the company stresses what Opus 5 deliberately does not do. It says the model does not advance the frontier of dangerous dual-use capability: it remains behind Mythos 5 on both offensive cybersecurity and biology research. Anthropic notes it did not train Opus 5 on cyber tasks, and that while the model has become nearly as good as Mythos 5 at finding software vulnerabilities, it lags far behind at exploiting them — the step that turns a flaw into an actual threat. Cyber safeguards allow vulnerability-finding in source code but block binary scanning, penetration testing and exploit generation, with flagged requests falling back to Opus 4.8. On biology, Opus 5 is now Anthropic's most capable generally available model for scientific research, though the company says it still shows meaningful limits on the long-running autonomous work it considers highest-risk.
Availability
Opus 5 is live today on all of Anthropic's platforms, callable as claude-opus-5 on the company's API, with a Fast mode that runs about 2.5 times quicker at twice the base price. Anthropic also shipped two developer features in beta alongside it: the ability to change which tools a model can use mid-conversation without invalidating the prompt cache, and automatic fallbacks that route safety-flagged API requests to another model rather than blocking them outright.
Whether Opus 5 lives up to the "near-frontier at half the cost" billing will be decided by independent testing and real-world use over the coming weeks. For now, it is a clear statement of Anthropic's strategy: compete less on the absolute top of the leaderboard and more on the balance of capability, cost and reliability that determines which model people actually run all day.



