Tech·July 24, 2026·5 min read

Anthropic Releases Claude Opus 5, Pitching Near-Frontier Performance at a Lower Price

Anthropic's new Opus 5 lands today at the same price as its predecessor, with company benchmarks showing large gains in agentic coding, computer use and knowledge work. It tops most of Anthropic's charts but not all of them, and still trails the specialized Mythos 5 on the riskiest cyber and biology tasks.

By Joseph Cooper

Anthropic Releases Claude Opus 5, Pitching Near-Frontier Performance at a Lower Price

Anthropic released Claude Opus 5 on Friday, positioning its new flagship as a model that approaches the intelligence of the company's frontier system, Fable 5, at roughly half the price. It is available immediately across Anthropic's platforms and priced identically to its predecessor, Opus 4.8 — $5 per million input tokens and $25 per million output tokens — a pitch built around holding cost flat while raising performance. Opus 5 becomes the default model on the Claude Max subscription and the strongest option on Claude Pro.

The framing matters as much as the numbers. Rather than chase an outright performance crown, Anthropic is selling Opus 5 as the model to reach for every day: near-frontier quality, but efficient enough to run at scale without the token bill of a top-tier system.

What the benchmarks show

By Anthropic's own measurements, the generational jump over Opus 4.8 is substantial. On Frontier-Bench v0.1, a test of agentic terminal coding, Opus 5 scores 43.3% — more than double Opus 4.8's 21.1% and ahead of both Fable 5 (33.7%) and OpenAI's GPT-5.6 Sol (34.4%). It leads the company's knowledge-work index GDPval-AA v2 (1861, against 1593 for Opus 4.8), and posts an outsized result on ARC-AGI-3, a novel-problem-solving test, where its 30.2% dwarfs the 7.8% of the next-best model. On the computer-use benchmark OSWorld 2.0 it reaches 70.6%, and on Zapier's AutomationBench, which grades whether a model can carry a business task from start to finish, it hits 26.0% — comfortably above every rival listed.

It is not a clean sweep, and to Anthropic's credit the company's own comparison table shows where Opus 5 falls short. GPT-5.6 Sol leads on the DeepSWE v1.1 coding benchmark (72.7% to Opus 5's 68.8%). Fable 5 narrowly edges it on Humanity's Last Exam without tools and on a held-out legal-agent test. And on health and on the "human-solved" tier of a biology benchmark, the specialized Mythos 5 remains ahead. The picture is of a model that tops most categories decisively while conceding a handful to more specialized or more expensive systems — which is roughly what the "near-frontier at lower cost" positioning would predict.

As with any launch, these are vendor-reported figures generated on the company's own harnesses, and independent replication will take time. The comparisons are most useful as a statement of where Anthropic believes its model stands, not as settled fact.

Agentic behavior is the selling point

The through-line in Anthropic's pitch is not raw scores but judgment — the model's willingness to verify its own work and keep going until a task actually succeeds. The company highlights cases from testing: on one Frontier-Bench task, given a drawing it was deliberately prevented from viewing, Opus 5 wrote its own computer-vision pipeline to extract the geometry from raw pixels and rebuild a 3D model, something no competing model managed in five attempts. In another, handed a real bug in an open-source package manager, it traced the root cause and fixed an edge case the community's own patch had missed, where a rival model addressed only the surface symptom and declared victory. A trading-firm engineer described using it to build a market-data feed in a single session, with the model constructing its own test harness when no live feed was available to check against.

Early-access customers quoted by Anthropic — including Cognition, Cursor, Zapier and others — echo the theme, describing gains concentrated on longer, vaguer, multi-step work and, notably, steadier results from run to run. Those are testimonials solicited for a launch, so they warrant the usual skepticism, but the consistency of the "judgment and reliability" framing across them is the clearest signal of what Anthropic thinks it has built.

Safety and the deliberate limits

Anthropic pairs the release with an unusually prominent safety argument. It says an automated behavioral audit rated Opus 5 its most aligned model to date, with the lowest measured rate of deceptive behavior and the best adherence to the company's published model constitution — a score of 2.3 on overall misaligned behavior, the lowest among its recent models. Again, that audit is Anthropic's own.

More concretely, the company stresses what Opus 5 deliberately does not do. It says the model does not advance the frontier of dangerous dual-use capability: it remains behind Mythos 5 on both offensive cybersecurity and biology research. Anthropic notes it did not train Opus 5 on cyber tasks, and that while the model has become nearly as good as Mythos 5 at finding software vulnerabilities, it lags far behind at exploiting them — the step that turns a flaw into an actual threat. Cyber safeguards allow vulnerability-finding in source code but block binary scanning, penetration testing and exploit generation, with flagged requests falling back to Opus 4.8. On biology, Opus 5 is now Anthropic's most capable generally available model for scientific research, though the company says it still shows meaningful limits on the long-running autonomous work it considers highest-risk.

Availability

Opus 5 is live today on all of Anthropic's platforms, callable as claude-opus-5 on the company's API, with a Fast mode that runs about 2.5 times quicker at twice the base price. Anthropic also shipped two developer features in beta alongside it: the ability to change which tools a model can use mid-conversation without invalidating the prompt cache, and automatic fallbacks that route safety-flagged API requests to another model rather than blocking them outright.

Whether Opus 5 lives up to the "near-frontier at half the cost" billing will be decided by independent testing and real-world use over the coming weeks. For now, it is a clear statement of Anthropic's strategy: compete less on the absolute top of the leaderboard and more on the balance of capability, cost and reliability that determines which model people actually run all day.

Share this story

Comments

Loading comments…

Get the Consensus Digest

The day’s most important stories, briefed and delivered to your inbox. No spam — just the news that matters.

No spam. Unsubscribe anytime.

More in Tech

A Nikon Z9 professional mirrorless camera body fitted with a Nikkor Z 24-70mm f/2.8 S lens, shown front-on.
Tech·August 16, 2026·5 min read

Nikon Has Announced Zero Cameras in 2026. Its Own Filing Names the Culprit Twice, and It Isn't the Camera Market.

Nearly everything on Nikon's 2026 roadmap is slipping toward 2027, and the company has cut its full-year forecast by 50,000 bodies and 50,000 lenses. Buried in the Q1 filing is the reason: 'higher memory prices.' DRAM spot prices are up roughly 700 percent in a year because AI data centres are buying the supply. That is now showing up in the camera aisle.

Read →
An aerial view at dusk of a vast industrial semiconductor facility beside water.
Tech·August 5, 2026·4 min read

SpaceX Is Building a $119 Billion Chip Fab in Rural Texas — and Its Own Power Plants, Because the Grid Won't Carry It

Terafab is now confirmed for Grimes County: 100 million square feet, a $55 billion first phase, up to $119 billion total, and a target of one terawatt of output a year. SpaceX says it will generate its own electricity rather than draw from ERCOT. Nearly 900 residents have signed a petition asking for protections — after the agreement was already signed.

Read →