Anthropic releases Claude Opus 5 at same price as Opus 4.8

The model becomes the default for Claude Max and the strongest option on Claude Pro, with new controls governing cybersecurity requests and safety fallbacks.

ETIH illustration representing cloud-based AI, connected digital services and enterprise technology. Claude Opus 5 is now available through Anthropic’s platforms and API.

Anthropic has made Claude Opus 5 available across all platforms at the same base API price as Opus 4.8

Anthropic has released Claude Opus 5 across all platforms, keeping the same base API pricing as its Opus 4.8 predecessor while claiming substantial gains across coding, knowledge work, computer use and scientific research.

Opus 5 costs $5 per million input tokens and $25 per million output tokens. Developers can access it through the Claude API using claude-opus-5, while the model is now the default on Claude Max and the strongest model available through Claude Pro.

The company positions Opus 5 as an everyday alternative to its more capable Claude Fable 5 model, claiming it approaches Fable 5 performance at half the price. That comparison varies by task. Anthropic acknowledges that Opus 5 remains behind Fable 5 on some benchmarks and trails its Mythos 5 model in offensive cybersecurity and long-running biological research.

Customers can also adjust the model’s effort setting, trading additional reasoning and token use against speed and cost. Anthropic says this allows organizations to reserve higher effort levels for complex work rather than paying the maximum cost for every request.

Anthropic claims gains in coding and business tasks

The strongest performance claims center on agentic work, where an AI model must complete a sequence of actions rather than produce a single response.

On Frontier-Bench v0.1, Anthropic reports that Opus 5 recorded a score of 43.3%, compared with 33.7% for Fable 5 and 21.1% for Opus 4.8. The company’s published comparison also places GPT-5.6 Sol at 34.4%.

The Frontier-Bench figures come from an internal run using the mini-SWE-agent harness and a Google Kubernetes Engine backend. Results are based on a mean reward across five attempts per task. Opus 4.8 was used as a fallback when the safety classifiers for Opus 5 or Fable 5 refused a request, meaning the headline score does not represent Opus 5 operating alone in every case.

Cost also forms part of Anthropic’s performance argument. On CursorBench 3.2, Opus 5 at its maximum effort setting reportedly came within 0.5% of Fable 5’s peak result while costing half as much per task. Anthropic further claims Opus 5 produced better performance at a given cost than competing models at its high, xhigh and maximum effort settings.

Results supplied by the company place Opus 5 ahead on several knowledge-work and computer-use evaluations. It scored 1,861 on GDPval-AA v2, against 1,747 for Fable 5, 1,593 for Opus 4.8 and 1,736 for GPT-5.6 Sol.

On AutomationBench, which tests whether a model can complete business workflows from beginning to end, Opus 5 recorded a 26% pass rate. Fable 5 reached 17.4%, Opus 4.8 scored 17% and GPT-5.6 Sol recorded 18.1%.

Anthropic says Opus 5’s AutomationBench pass rate was around 1.5 times that of the next-best model at the same cost per task. It also claims that the model surpassed every competitor even when operating at its lowest effort setting.

The advantage was less consistent elsewhere. Opus 5 scored 68.8% on DeepSWE v1.1, behind Fable 5 at 69.7% and GPT-5.6 Sol at 72.7%. On Humanity’s Last Exam without tools, it scored 56.3%, narrowly below Fable 5’s 56.5%.

Anthropic describes Opus 5 as more likely to verify its work and continue iterating until a task succeeds. Its examples include reconstructing a machine part as a 3D FreeCAD model after building a computer vision pipeline to extract geometry from raw pixels, and fixing the underlying cause of a bug in an open-source package manager after another model addressed only the immediate symptom.

A further early-access example involved an engineer using the model to build a market data feed for a new exchange. Anthropic says Opus 5 created a test harness when no live feed was available for validation. The company presents these as examples from evaluations and early-access use, separate from the standardized benchmark results.

Research performance improves, with stated limits

Opus 5 outperformed Opus 4.8 across Anthropic’s life sciences evaluations, which cover structural biology, organic chemistry and bioinformatics.

The largest gains included a 10.2 percentage point improvement on an internal task involving the inference of molecular structures from spectroscopy data. Opus 5 also scored 7.7 percentage points higher on tasks that asked the model to predict how changes in a protein sequence would affect its function.

Anthropic describes Opus 5 as its most capable generally available model for scientific research, but draws a boundary around autonomous biological work. The company says the model retains important limitations on long-running research tasks and that Mythos 5 remains stronger for this type of work.

Biology-related requests blocked by Fable 5 will now route to Opus 5 instead of Opus 4.8. The broader safeguard system remains similar to that used for Opus 4.8.

Cyber controls separate detection from exploitation

Anthropic says Opus 5 does not move the frontier forward in risky dual-use capabilities. In evaluations conducted with private-sector and government partners, it remained behind Mythos 5 in biology research and offensive cybersecurity.

The distinction is clearest in OSS-Fuzz, Anthropic’s evaluation of whether models can identify and exploit vulnerabilities without extensive human guidance. Opus 5 passed 79.4% of vulnerability-identification challenges, close to the 80% result recorded by Mythos 5 and ahead of Opus 4.8 at 61.5%.

Its ability to turn those vulnerabilities into working exploits was considerably lower. Mythos 5 completed 13 exploitation challenges, compared with four for Opus 5 and none for Opus 4.8.

Anthropic says it intentionally avoided training Opus 5 specifically on cyber tasks. The model nevertheless improved as its wider capabilities increased, prompting stronger controls for a narrow range of cybersecurity requests.

Opus 5 can be used to find vulnerabilities in source code, but its classifiers block binary-based vulnerability scanning, penetration testing and exploit generation. Anthropic expects these controls to intervene around 85% less often than those attached to Fable 5.

Flagged requests in Claude.ai, Claude Code and Claude Cowork fall back to Opus 4.8 by default. API customers can also enable automatic fallbacks, one of two beta updates released alongside the model. When enabled, a request flagged by Opus 5 or Fable 5 is routed to another available model instead of being blocked.

Existing members of Anthropic’s Cyber Verification Program have immediate access to an Opus 5 version with fewer security restrictions for approved cybersecurity work.

A separate behavioral audit produced an overall misaligned-behavior score of 2.3, which Anthropic says is the lowest among its recent models. The company reports lower rates of deceptive behavior and fewer reckless actions with potentially difficult-to-reverse consequences than Opus 4.8, Sonnet 5 or Fable 5.

The second beta update allows developers to change the tools available to Claude during a conversation without invalidating the prompt cache.

Fast mode runs Opus 5 at around 2.5 times its standard speed. It costs twice the model’s base price on the Claude Platform and can be accessed through usage credits in Claude Code. General access to Opus 5 carries no data-retention requirements.

Previous
Previous

U.S. Commerce plans $169 million investment across six Tech Hubs

Next
Next

OpenAI opens four-week Codex studio for faculty and researchers