In a move that signals a significant shift in the competitive landscape of large language models, Anthropic has officially released Claude Opus 5, a new flagship model designed to bridge the gap between extreme reasoning capabilities and cost-efficiency. Replacing the previous Opus 4.8 iteration, Opus 5 now serves as the premier offering within the Opus-tier, positioned by Anthropic’s engineering team as a model that approaches the sophisticated intelligence of the ultra-high-end Claude Fable 5 while maintaining a significantly lower price point. Despite the substantial jump in performance across reasoning, coding, and agentic tasks, Anthropic has opted to keep its pricing structure stable, maintaining a rate of $5 per million input tokens and $25 per million output tokens, a decision likely intended to consolidate its market share among enterprise developers and power users.
The release marks a pivotal moment for Anthropic, as Claude Opus 5 becomes the default model on the Claude Max platform and the most powerful available option for Claude Pro subscribers. The model’s arrival comes at a time when the industry is increasingly focused on "test-time compute"—the ability of a model to use tools and iterative reasoning to solve complex problems—rather than relying solely on the raw scale of pre-training. By delivering Fable-level intelligence at half the cost, Anthropic is directly challenging the price-to-performance ratios of competitors like OpenAI’s GPT series and Google’s Gemini.
Architectural Refinements and API Enhancements
At the technical and API level, Claude Opus 5 introduces several critical updates that differentiate it from its predecessor. Identified by the model ID claude-opus-5, the new flagship features a standard context window of 1 million tokens, which now serves as both the default and the maximum capacity, removing the need for smaller context variants. This massive window allows users to ingest entire codebases, multi-volume legal documents, or extensive financial histories in a single prompt.
Output capabilities have also seen a dramatic expansion. While the synchronous Messages API supports a maximum output of 128,000 tokens, the Message Batches API—utilizing the output-300k-2026-03-24 beta header—can generate up to 300,000 tokens in a single request. This capability is particularly relevant for long-form content generation and complex data synthesis. Furthermore, Anthropic has optimized the prompt caching mechanism; the minimum cacheable prompt threshold has been reduced from 1,024 tokens to 512 tokens. This change is expected to significantly lower operational costs for developers who frequently reuse system prompts or reference documents in high-frequency API calls.
A Chronology of the Claude Ecosystem
The trajectory of the Claude series has been one of rapid, iterative improvement. Starting with the initial launch of Claude 1, Anthropic focused on "Constitutional AI," a framework designed to make models safer and more honest without heavy manual moderation. By the time Claude 3 and 3.5 were released, the focus shifted toward multimodal capabilities and coding proficiency. The transition from Opus 4.8 to Opus 5 represents the culmination of these efforts, moving from a paradigm of "helpful assistant" to "autonomous agent."
In early 2026, Anthropic’s roadmap emphasized the development of the "Fable" and "Mythos" tiers, intended for specialized research and cybersecurity, respectively. Opus 5 acts as the bridge between the general-purpose utility of the Sonnet tier and the extreme, resource-heavy capabilities of the Fable tier. This release timeline suggests that Anthropic is prioritizing the democratization of high-level reasoning, making it accessible to a broader range of developers before the eventual wider rollout of its next-generation research models.
Breakthroughs in Reasoning and the ARC-AGI-3 Result
One of the most striking achievements of Claude Opus 5 lies in the realm of mathematical reasoning and general intelligence. In a rigorous evaluation, Anthropic prompted the model on all six problems from the International Mathematical Olympiad (IMO) 2026. Notably, the model was not provided with external tools or an agentic harness, relying solely on its internal reasoning. A panel of three independent AI judges scored all 24 generated solutions as correct. This was further verified by human experts who graded one pre-specified solution per problem at a perfect 7/7. The resulting score of 42/42 places Opus 5 at a gold-medal level, far exceeding the 29/42 cutoff typically required for such an honor.
Equally significant is the model’s performance on the ARC-AGI-3 (Abstraction and Reasoning Corpus). The ARC Prize Foundation reported a verified score of 30.16% for Opus 5 at high effort. To put this in perspective, this score is nearly four times higher than the previous leaderboard record. For comparison, GPT-5.6 Sol reached 7.78%, while the previous Opus 4.8 struggled at 1.52%. The ARC-AGI-3 is widely considered one of the most difficult benchmarks in AI because it requires the model to solve novel visual logic puzzles it has never seen before, serving as a proxy for fluid intelligence.
On "Humanity’s Last Exam," a benchmark designed to test the limits of expert-level knowledge across diverse fields, Opus 5 scored 56.3% in a zero-shot setting. When equipped with tools, this score rose to 64.7%, illustrating the model’s ability to effectively leverage external information to solve highly specialized academic problems.
Dominance in Coding and Agentic Workflows
The development of agentic AI—models that can navigate operating systems and use software tools like humans—is a primary focus for Anthropic. On OSWorld 2.0, a benchmark that tests a model’s ability to perform tasks in a real computer environment, Opus 5 achieved a success rate of 70.57%. This is a substantial leap from the 55.7% recorded by Opus 4.8. Similarly, on the Zapier AutomationBench, Opus 5 scored 26.0%, outperforming both Opus 4.8 (17.0%) and the more expensive Fable 5 (17.4%).

In the software engineering domain, Opus 5 has set new records on the SWE-bench. It scored 96.0% on SWE-bench Verified and 79.2% on SWE-bench Pro. While Fable 5 still maintains a slight edge on the Pro version at 80.0%, the gap has narrowed to a negligible margin. The most dramatic improvement occurred in SWE-bench Multimodal, where Opus 5 jumped to 59.4% from the 38.4% achieved by its predecessor.
On FrontierBench v0.1, a successor to Terminal-Bench 2.1 consisting of 74 complex tasks, Opus 5 scored 43.3% at maximum effort. This more than doubles the 18.7% score of Opus 4.8 and notably surpasses the 33.7% of Fable 5. These results suggest that Opus 5 is currently the most capable model on the market for autonomous technical tasks and complex workflow automation.
Multimodal Performance and the Power of Tools
Anthropic’s latest research emphasizes that agentic tool use can scale "test-time compute" more effectively than simply increasing the model’s internal thinking time. This is evidenced by the model’s multimodal performance. On the "Chartography" benchmark, Opus 5’s score rose from 29.6% (without tools) to a staggering 83.0% when given access to a containerized environment and an image-cropping tool.
A similar trend was observed in BenchCAD Vision2Code, where the model’s voxel Intersection over Union (IoU) moved from 0.366 to 0.821. With tools, Opus 5 significantly outperformed Claude Mythos 5, which scored 0.678. This data reinforces the industry’s shift toward "compound AI systems," where the model acts as the central controller for a suite of specialized tools.
Cybersecurity Capabilities and the Evolution of Safeguards
The release of Opus 5 brings a nuanced change to Anthropic’s approach to cybersecurity. While the model was not specifically trained on cyber-offensive tasks, its general intelligence gains naturally translated into increased capability in this area. On ExploitBench, Opus 5 captured 10.14 mean capability flags and produced 99 full arbitrary-code-execution exploits. This is close to the 132 exploits produced by the cybersecurity-specialized Mythos 5.
However, Anthropic has observed a "capability gap" that informed a strategic shift in their safety protocols. Opus 5 is nearly as proficient as Mythos 5 at identifying vulnerabilities in source code but remains significantly less capable of actually executing complex exploits. Consequently, Anthropic has unblocked vulnerability-finding capabilities in source code for all users, while maintaining strict blocks on binary-based scanning, penetration testing, and exploit generation.
This relaxation of safeguards is expected to result in 85% fewer "false positive" refusals compared to Fable 5, making the model more useful for defensive researchers and developers. To support this, Anthropic has introduced the Cyber Verification Program, allowing legitimate defenders to access more robust capabilities. In independent testing by the UK AI Safety Institute (AISI), Opus 5 solved the "The Last Ones" cyber range end-to-end in 80% of attempts and reached the penultimate step of the "Doing Life" range, further than any previously tested model.
Safety Metrics and Prompt Injection Resistance
Despite the increase in raw power, Opus 5 shows improved resistance to adversarial attacks. On the Gray Swan indirect prompt injection benchmark, the success rate for attackers fell from 5.5% in Opus 4.8 to just 2.0%. This is notably superior to GPT-5.6 Sol, which recorded a 20.0% success rate on the same benchmark.
In real-world browser environments simulated through Claude Cowork, the success of prompt injection attacks dropped from 31.5% to 3.70% without any additional safeguards. When "auto mode" safeguards were enabled, the attack success rate fell to 0%. Under Anthropic’s Responsible Scaling Policy (RSP), Opus 5 is classified as having CB-1 (Cyber-Biological) capabilities but not CB-2, meaning it remains within the ASL-3 (AI Safety Level 3) protection tier. This ensures that while the model is highly capable, it does not cross the threshold into providing high-level assistance for creating biological threats or autonomous cyber-warfare.
Broader Implications for the AI Industry
The launch of Claude Opus 5 represents a "sweet spot" in the evolution of generative AI. By delivering frontier-level reasoning and agentic performance at a mid-range price point, Anthropic is forcing a recalibration of value across the industry. The model’s success on the ARC-AGI-3 and IMO benchmarks suggests that the ceiling for LLM-based reasoning is much higher than previously thought, especially when models are optimized for accuracy over mere conversational fluency.
For enterprises, the reduction in prompt caching costs and the massive 1M token context window make Opus 5 a formidable tool for "Big Data" analysis and automated software development. For the broader AI community, the model’s performance serves as a validation of Anthropic’s focus on safety-led development, proving that robust safeguards and world-leading performance are not mutually exclusive. As the industry moves toward the latter half of 2026, the success of Opus 5 will likely serve as the benchmark against which all forthcoming models from OpenAI, Google, and Meta will be measured.
