AI Consulting
Claude Opus 5: What Anthropic's New Model Can Do, What It Costs — and Where the Record Numbers Demand a Second Look
Synthetic voice · AI-generated (text-to-speech).
On July 24, 2026, Anthropic released Claude Opus 5 — by its own account the most capable and, at the same time, best-aligned model the company has ever built. The first of those two promises is unusually well supported: this time the record numbers do not appear only in the marketing but are broadly confirmed by independent evaluators. The second is contested. And, of all things, the metric that matters most in day-to-day use — reliability — moves the other way.
For anyone deploying AI in a business, that is the genuinely interesting story. Not “new best model,” but: a real performance jump at the same price, accompanied by a pattern you should know before you rely on it. Let’s look at it soberly — at what’s new, what it costs, and what the numbers actually say.
At a glance
- What: Claude Opus 5, released July 24, 2026 (model ID
claude-opus-5) — the fourth model in the Claude 5 family in under two months.- New: thinking is on by default; a full effort ladder from
lowtomax; a 1M-token context as standard.- Price: $5/$25 per million tokens (in/out) — exactly the price of its predecessor Opus 4.8. Optional Fast Mode $10/$50.
- Benchmarks: leading to level with the more expensive Fable 5 — and this time broadly confirmed independently, not just claimed.
- The catch: independent analyses report more overconfidence and hallucination; the most expensive reasoning tiers do not automatically deliver the best results.
What’s new: a flagship that thinks by default
Opus 5 is already the fourth model in the Claude 5 family within under two months — the cadence at which the big labs now ship is itself a data point. Technically much stays with the proven architecture; the most visible behavioral change: extended thinking is now on by default. Across the full effort ladder (low, medium, high, xhigh, max) you can set how much compute the model puts into an answer — from a quick reflex to a long chain. Add to that developer-facing refinements: tool changes mid-conversation (beta), server-side fallback answers on refusal, and a lower minimum for prompt caching.
| Key facts | Claude Opus 5 |
|---|---|
| Context window | 1,000,000 tokens (default and maximum) |
| Max output | 128,000 tokens |
| Price (API) | $5 input / $25 output per million tokens |
| Fast Mode (research preview) | $10 / $50 per million tokens |
| Availability | Claude API, Claude.ai (Pro/Max/Team/Enterprise), Amazon Bedrock, Google Vertex AI, Microsoft Foundry |
| Safety level | ASL-3 (triggered by chemical-biological capabilities) |
Availability and price: same price, more model
For most companies, this is where the real news lies. Over the API, Claude Opus 5 costs exactly as much as its predecessor — $5 per million input tokens, $25 per million output tokens. The performance jump is therefore not bought with a price jump. For especially latency-critical cases there is additionally a faster Fast Mode as a research preview, at double the token price.
| Mode | Input | Output |
|---|---|---|
| Standard | $5.00 | $25.00 |
| Fast Mode (preview) | $10.00 | $50.00 |
The second, strategically more important point: Anthropic explicitly positions Opus 5 as the cheaper workhorse below its most expensive model, Fable 5 — “close to Fable 5 intelligence at half the price.” Artificial Analysis’s independent cost accounting supports this: across a task course, Opus 5 comes to roughly $2.03 per task versus $2.75 for Fable 5 — about 26 percent cheaper at a practically equal intelligence index. For many production workloads, that shifts the economically right choice considerably.
Benchmarks: strong — and this time broadly confirmed
The most remarkable difference from some other releases of recent weeks is not the height of the numbers but their robustness. Where individual competitors showed mainly their own, conveniently chosen indexes at launch, Anthropic’s figures here largely match what independent evaluators measure.
Where Opus 5 convinces (independently measured, as of late July 2026):
| Benchmark | Value | Source/condition |
|---|---|---|
| Intelligence Index (aggregate) | 61 (max) · 56 for Opus 4.8 | Artificial Analysis |
| SWE-bench Verified (coding) | ~97% | vals.ai (independent) |
| Terminal-Bench v2.1 | ~89% | independent |
| GPQA Diamond (reasoning) | ~93% | aggregators |
| OSWorld 2.0 (computer use) | 70.6% | Anthropic, independently supported |
| Humanity’s Last Exam | 56.3% / 64.7% with tools | independent |
Anthropic additionally highlights large gains in specialist domains over Opus 4.8 — roughly +10 percentage points in organic chemistry, +9 in financial modeling, +7.7 in protein prediction. And on its own new agentic test Frontier-Bench, Opus 5 doubles its predecessor (43.3 versus 21.1 percent) — though that value comes from a benchmark Anthropic introduced itself and should be read as a vendor figure. On the young abstraction test ARC-AGI-3, early independent analyses report 30.2 percent — far above GPT-5.6 Sol (7.8) and the predecessor (1.5); that number rests on a single source so far and should be treated with caution.

Two measurements of the same blade — and, right at the tip, a tiny offset. The distance between the advertised number and the verified reality is small for Opus 5. Small is not zero.
Where the numbers demand a second look
As solid as the top scores are, three findings belong in the same story — and the first is the most important for practical use.
First: reliability points the wrong way. An independent analysis (the-decoder, drawing on tests by Vals.ai) reports that Opus 5 answers more often when it is actually unsure — on the corresponding measure, the hallucination rate rose to around 50 percent, and the model reportedly stated answers as certain when it was not, only to retract them shortly after. A more capable model that is at the same time more confidently wrong is not pure progress for automated workflows.
Second: more reasoning is not automatically better. The same analysis found that the highest effort tiers (xhigh, max) produce more complex solutions that more often contain errors, while the high tier delivered the better results at lower cost. The marginal return of the most expensive tier declines — more expensive is not the same as better here either.
Third: “best-aligned” is a claim, not a measurement. Anthropic calls Opus 5 its best-aligned model and bases that on safety benchmarks. Independent commentator Zvi Mowshowitz sharply disagrees: one must not confuse a good test score with actual alignment — “Optimization against metrics is practically unavoidable, but ontological confusion between benchmark scores and the North Star of alignment is not.” The point is the same as in the first finding: a score measures the test, not reality.
Fairness requires the other side, and it is independently confirmed: on the plus side there is markedly improved resistance to prompt injection (attack success rate cut from 5.5 to 2.0 percent) and drastically fewer false alarms from the safety classifiers (from 42 to 5 percent) — both genuine gains for use in agent systems. And Anthropic released the model under its ASL-3 safety level, triggered by chemical-biological capabilities — a signal of how seriously the capability gains are taken.
A word on where it sits in the field: per Epoch AI’s index, Opus 5 actually lands just behind Fable 5 overall (159 to 161 points) — a different picture depending on the evaluator. And because the three major vendors no longer share a common benchmark suite, direct comparison tables between Opus 5, GPT-5.6 and Gemini 3 are methodologically fragile. That is exactly why we run a live benchmark monitor on the AI consulting page, drawing its rankings from independent sources (LMArena, Epoch AI). Opus 5 has appeared there for only a few days — credible rankings need one to two weeks of votes to settle.
What this means for your business
Three sober lessons follow from this release — regardless of whether Opus 5 ends up in your stack.
First: benchmarks are a start, not a substitute. That Anthropic’s numbers hold up independently this time is welcome — but it does not replace a test on your own tasks, with your own data. No published score tells you how the model performs on exactly your documents, processes and edge cases. How to measure quality robustly, we described in Evaluating AI outputs.
Second: choose the tier, not the headline. The most interesting practical finding is that here the highest reasoning tier was not the best. For most workloads, high is the smarter choice than max — cheaper and, at the same time, more reliable. And for many cases a smaller model is enough anyway. Hitting the fitting tier is the lever, not booking the most expensive one.
Third: more confidence needs more control, not less. A model that answers confidently more often — even when it is unsure — belongs, in production systems, behind a control layer that verifies results independently rather than trusting them. That is exactly the pattern a hallucination control layer is built against — and with a more confident model it becomes more important, not less.
Which model — and which tier — fits which of your processes is rarely a pure benchmark question. It is precisely at this intersection of technology, cost and safeguards that I work, as a developer and business lawyer in one person. If you want to test Claude Opus 5 or a competing model against your concrete use cases, let’s talk.
FAQ
When was Claude Opus 5 released and what does it cost?
Anthropic released Claude Opus 5 on July 24, 2026 (model ID claude-opus-5). API prices are 5 US dollars per million input tokens and 25 US dollars per million output tokens — unchanged from its predecessor Opus 4.8. A new, faster Fast Mode (research preview) costs 10 and 50 US dollars respectively. The context window is one million tokens, with a maximum output of 128,000 tokens. The model is available via the Claude API, Claude.ai (Pro, Max, Team, Enterprise) as well as Amazon Bedrock, Google Vertex AI and Microsoft Foundry.
How does Claude Opus 5 perform on benchmarks?
Strongly — and this time broadly confirmed by independent measurement. On Artificial Analysis’s Intelligence Index, Opus 5 reaches 61 points at its highest setting (predecessor Opus 4.8: 56). On SWE-bench Verified, independent evaluators measure around 97 percent, on Terminal-Bench v2.1 around 89 percent, and on GPQA Diamond around 93 percent. Anthropic itself highlights large gains in specialist domains (organic chemistry, financial modeling, protein prediction). Unlike some competing releases, the vendor’s own figures largely match the independent measurements here.
Is Claude Opus 5 better than Claude Fable 5 or GPT-5.6?
It depends on the evaluator. Per Artificial Analysis, Opus 5 is essentially level with the more expensive Fable 5 (index 61 vs 60) — at roughly a quarter lower cost per task. Per Epoch AI’s index, however, it sits just behind Fable 5. Against GPT-5.6 Sol, Opus 5 leads on several agentic tests while GPT leads on others. No single model pulls clearly ahead — and because the three major vendors no longer share a common benchmark suite, direct comparison tables should be read with caution.
What should companies keep in mind with Claude Opus 5?
Two things. First, no published score replaces a test on your own tasks — that remains mandatory. Second, independent analyses report increased overconfidence: Opus 5 answers more often even when it is unsure, and the highest reasoning tiers (xhigh, max) sometimes produce more complex, more error-prone solutions than the high tier. For production systems that means: choose the fitting tier — not the most expensive one — and add a control layer that verifies results independently.
Sources — as of 2026-07-27
- Anthropic (primary source): Introducing Claude Opus 5 (2026-07-24) — https://www.anthropic.com/news/claude-opus-5
- Anthropic (technical docs): What’s new in Claude Opus 5 — https://platform.claude.com/docs/en/about-claude/models/whats-new-opus-5
- Anthropic: Activating ASL-3 protections — https://www.anthropic.com/news/activating-asl3-protections
- Artificial Analysis (Intelligence Index, cost per task, Fable 5 comparison) — https://artificialanalysis.ai/models/claude-opus-5
- vals.ai (SWE-bench Verified, independent) — https://www.vals.ai/benchmarks/swebench
- the-decoder (cost/performance, hallucination/overconfidence, Epoch AI index) — https://the-decoder.com/anthropics-claude-opus-5-costs-well-below-fable-5-while-matching-or-beating-it-across-most-benchmarks/
- Zvi Mowshowitz (independent system-card analysis, alignment critique, safety gains) — https://thezvi.substack.com/p/claude-opus-5-the-system-card
- Context/reporting: Axios — https://www.axios.com/2026/07/24/anthropic-releases-new-model-opus-5 ; TechCrunch — https://techcrunch.com/2026/07/24/anthropic-launches-opus-5/ ; VentureBeat — https://venturebeat.com/orchestration/anthropic-launches-claude-opus-5-a-cheaper-ai-model-for-coding-agents-and-enterprise-workflows
Note: the full primary text of the system-card PDF was not directly reachable at the time of research due to its file size; content cited from it comes from secondary sources (chiefly Zvi Mowshowitz’s detailed analysis). The ARC-AGI-3 value rests on a single source so far. Some circulating figures (e.g., IMO 2026 and MMLU-Pro) were only singly sourced and were deliberately omitted. A direct, current head-to-head against Gemini 3 on identical benchmarks was not reliably available.
This article is general information and not legal or investment advice. It covers a very fast-moving market; as of July 27, 2026, please check the current state before making decisions.