Claude Opus 5.5

Hacker News by 25 min read 27x views
Claude Opus 5.5

Share Post

We’re introducing Claude Opus 5.5, the archetypal example in our new Claude 5.5 family. It performs at the flat of Claude Fable 5.1 on most activity and expenses 40% small to run than Opus 5.

Claude Opus 5.5 is our archetypal publish since we called for pacing the frontier. It was tested before publish by external evaluators, including Frontier Design and METR. On our automated behavioral audit, the most thorough alignment test we run, Opus 5.5 is the strongest-performing example we’ve tested to date. It additionally comes alongside the safeguards we’ve developed for our most capable models.

Here are several of the improvements you can anticipate from Opus 5.5:

Performance. Opus 5.5 is a important stage up from Opus 5. It’s the new foremost model, and first testers saw ample jumps in achievement on their most complex work. One tester completed a 680,000-line code immigration in small than a day—work that would have taken an engineering squad weeks. It’s fine at finding and fixing inefficiencies in software: whenever we asked it to cut burden times throughout all leaf of a web app, Opus 5.5 succeeded 39 of 40 times, during Opus 5 made smaller improvements that additionally altered the app’s behavior. A distinct tester had multiple Claude models build a equivalent from a sole prompt; Opus 5.5 scored higher than any another example on the power of its graphics and polish.

Safety. Opus 5.5 achieves the finest scores of any example to date on our automated behavioral audit, our alignment suite that tests Claude throughout thousands of simulated scenarios. It is much small apt than latest models to obtain hard-to-reverse actions or act exterior the boundaries it’s been given, and it’s additional resistant than Opus 5 to immediate injection. We’ve additionally broadened our alignment evaluation to shield longer tasks, unattainable tasks, and scenarios modeled on genuine incidents, although it motionless has limits. Full particulars of our evaluation are accessible in the Opus 5.5 System Card.

Because Opus 5.5 is comparable to Claude Mythos 5.1 in existence discipline and cybersecurity, we’re deploying it alongside safeguards akin to those on Claude Fable 5.1. Vetted organizations can use today to our Life Sciences Verification Program to use Opus 5.5 for existence discipline research. In the coming weeks we volition additionally be expanding admission to our Cyber Verification Program, and verified cybersecurity practitioners volition be capable to use Opus 5.5 for their work.

Cost and speed. Opus 5.5 requires small compute to assist than Opus 5, and its pricing reflects that. Our tests display that at default settings it volition disbursal 40% small than Opus 5 on representative workloads. Input and output tokens are $4 and $20 per million, 20% small than Opus 5. Cache says (which create up the bulk of agentic and coding activity costs) are $0.20 per myriad tokens, 60% small than Opus 5. Opus 5.5 additionally generates output additional than 30% faster than Opus 5.

In supplement to the cost drop, we’re expanding five-hour use limits on Pro, Max, Team, and seat-based Enterprise plans. We’re additionally providing subscription users a charge bounds reset, which you can now preserve and use whenever you choose.

Communication. Opus 5.5 communicates additional naturally than previous models. Early testers established its penning clearer and easier to follow, which addresses several of the average feedback we heard concerning Opus 5. It puts the most crucial data up front, and its manner makes it a improved activity partner complete lengthy sessions. As one first tester put it, “it writes the way I do.” In our own use, this has made Opus 5.5’s activity easier to prosecute and check—which is a safety advantage as fine as a applicable one.

Claude Sonnet 5.5 and Claude Haiku 5.5 volition prosecute in the coming weeks, alongside many of the identical improvements to performance, efficiency, and safety.

Performance and cost-effectiveness

On our benchmarks, Claude Opus 5.5 leads in agentic coding, device use, and cognition work. That said, at these levels of capability we’ve established that benchmark margins have rotate into a small dependable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.

Opus 5.5Fable 5.1Opus 5GPT-6 AstraGPT-5.6 Sol
Agentic codingTerminal-Bench 4.0¹66.4%55.8%52.3%57.9%37.3%
Agentic codingFrontierCode v1.1 (Main)54.4%50.3%48.0%53.3%47.5%
Agentic codingCursorBench 4.057.8%51.8%46.6%41.7%
Knowledge workGDPval-AA v2.118461735170815421588
Business workflowsAutomationBench²40.0%31.4%26.9%41.4%28.8%
Multidisciplinary reasoningHumanity's Last Exam67.7%with tools65.6%with tools63.6%with tools57.2%with tools
Agentic specialized researchTerminal-Bench-Science 0.1³58.7%52.6%29.0%64.6%22.4%
Computer useOSWorld 2.081.8%partial80.7%partial74.0%partial
Visual diagram recognitionChartography89.0%with tools88.4%with tools83.4%with tools

Unless alternatively noted, all Claude Opus 5.5 results use adaptive thinking at max effort. Terminal-Bench 4.0 results are reported for Claude Opus 5.5 at xhigh attempt and GPT-6 Astra at elevated effort, as reported by OpenAI; these portray all model’s highest score. Claude Opus 5.5 was evaluated alongside its manufacturing safeguards enabled. When they intervened, cybersecurity tasks were completed by Claude Opus 4.8, and existence discipline and frontier LLM betterment tasks were completed by Claude Opus 5. This apt reduces Claude Opus 5.5’s achievement on these benchmarks.

1 Terminal-Bench 4.0: The norm error is ±2.6 pts for Claude Opus 5.5 and ±1.6–2 pts for the another Claude models. The community leaderboard (5 trials/task, Claude Code harness) reports Claude Opus 5 at 51.8%; our setup reproduces it at 52.3%, inside noise. GPT-6 Astra and GPT-5.6 Sol figures are as reported by OpenAI.

2 AutomationBench: AutomationBench results were run and reported by Zapier. These runs were performed without fallback models, so safeguard interventions were considered failures—this caused a lesser mark than Claude Opus 5.5 would accomplish in practice. Claude Opus 5.5 results arrive from Zapier’s own evaluation during first access. Results for Opus 5, GPT-5.6 Sol, and GPT-6 Astra arrive from Zapier’s community leaderboard.

3 Terminal-Bench-Science 0.1: The norm error is ±3.5–5 pts per model. The community leaderboard (3 trials/task, Claude Code harness) reports Claude Opus 5 at 30.0%; our setup reproduces it at 29.0%, inside noise. The GPT-6 Astra fig is as reported by OpenAI.

Where Opus 5.5’s advantage is extremely apparent is efficiency. It expenses small per token than Opus 5 and uses small tokens per task, which nets out to a 40% autumn in costs.

Pricing

Prices per 1M tokensClaude Opus 5.5Claude Opus 5
Cache reads$0.20$0.50
Input tokens$4$5
Output tokens$20$25
Cache writes$5$6.25

Fast manner for Opus 5.5 is additionally accessible in Claude Code and the Claude Platform alongside up to 2.5x speed. It expenses $8 per myriad input tokens and $40 per myriad output tokens.

Coding

Opus 5.5 is particularly fine at lengthy and sprawling jobs akin codebase-wide migrations and audits. An first tester used it to audit and fix a 200,000-line codebase in under three hours, anywhere Opus 5 took complete 20 hours and used 2.5x as many tokens. In an inner test, we asked Opus 5.5 and Fable 5.1 to translate HAProxy, extensively used application that balances web traffic loads throughout servers, from C into Rust. Both rewrites passed nearly all of HAProxy’s own regression tests, but Opus 5.5 completed in 9.5 hours compared to 12 for Fable 5.1, and disbursal 51% less.

Opus 5.5 delivers frontier results on agentic coding at a fraction of the cost. At its default attempt flat on FrontierCode, it strikes GPT-6 Astra at approximately 20% of the disbursal per task. On Terminal Bench 4.0, it matches Astra for concerning 40% of the cost, during on CursorBench it strikes GPT-5.6 Sol by 11 points for concerning a third of the cost.

Agentic terminal codingAgentic coding: FrontierCodeAgentic coding: CursorBench

Terminal-Bench 4.0Accuracy vs Cost

010203040506070Score (%)251020Cost per attempt (USD, log scale)lowmedhighxhighmax

Our first testers reported akin effectiveness and intellect gains:

GitHubClioLovableQuantiumSpotifyOptiverColumnKiro

Quote

“Developers desire agents that can obtain on genuine application activity and complete it. In our evaluation throughout GitHub Copilot CLI and VS Code, Claude Opus 5.5 used among the fewest tokens and steps we measured. In VS Code, it solved additional terminal tasks than Opus 5 in small than fractional the steps. More than making idiosyncratic tasks additional efficient, it’s making developers’ bigger projects additional achievable.”

CompanyGitHub

AuthorMario Rodriguez, Chief Product Officer

The most safe coding agent

Enterprises that use agents inside their systems need to cognize that those agents are functioning as intended, particularly whenever they run autonomously for many hours. Opus 5.5 has a classifier that screens all act before it runs, an open-source sandbox that safety teams can audit, and code assessment that catches vulnerabilities before they merge.

The example itself additionally has stronger defenses. On immediate injection attacks, it matches or strikes Opus 5 in all environment we tested, including coding, tool use, device use, and web browsing. On a benchmark run by the AI safety resolute Gray Swan, Opus 5.5 ties Fable 5.1 for the lowest immediate injection achievement charge of any example tested.

Knowledge work

Opus 5.5 is a dependable and adept researcher. In one inner test, we asked Opus 5.5, Fable 5.1, and Opus 5 to compose a study on a company’s quarterly achievement using lone the data it could discover on a copy of the web anywhere the earnings publish was difficult to locate. An automated grader checked all fig and citation against sources. Across distinct attempt settings, 16 out of 18 of Opus 5.5’s reports cleared our norm bar, anywhere any invented fig or citation would have failed. Neither Fable 5.1 nor Opus 5 cleared that bar in any attempt.

It’s additionally powerful in financial inspection and endeavor work. Walleye Capital, an funding resolute and first tester, reported that Opus 5.5 mostly solved their evaluation suite on its lowest setting; on higher settings, it performed equal better, noticing an error in their evaluation instructions and correcting for it. No another example had caught this error before.

In another test, we tasked the two Opus 5.5 and Opus 5 alongside analyzing a projected merger between two fictional HR application companies. Each built a financial example in Excel, afterward turned it into an administrator display on whether the agreement made awareness at its price. Both models reached the identical conclusions concerning the deal, but Opus 5.5’s example was additional thorough and its display easier to read, during Opus 5’s had insignificant errors. Opus 5.5 completed in 63 minutes compared to 93 for Opus 5, and disbursal 50% small to produce.

On cognition activity evaluations, Opus 5.5 outperforms another models during additionally using small tokens. On GDPval-AA v2.1, a test of real-world activity throughout 44 occupations, Opus 5.5 scores 1846 Elo, onward of Fable 5.1 and Opus 5. At default attempt (medium), Opus 5.5 strikes GPT-6 Astra at max attempt for concerning a fifth of the disbursal per task. It likewise outperformed another models on benchmarks measuring endeavor workflows and large-scale data collection.

GDPval-AA v2.1AutomationBenchWANDR

GDPval-AA v2.1Elo vs Cost

12001300140015001600170018000Elo0.200.5012510Estimated disbursal per project (USD, log scale)lowmedhighxhighmax

Our customers have reported akin results. Here’s what they told us concerning operating alongside the model:

Deloitte Consulting LLPRogoLexisNexis Legal & ProfessionalWalleye CapitalHexThomson Reuters LabsHebbiaViktor

Quote

“Even at its lowest attempt setting, Claude Opus 5.5 caught 72% of known bugs in our code reviews to Opus 5’s 56% at elevated effort, alongside small false alarms and a fraction of the output. On US consulting analysis, low thinking attempt matched its higher thinking settings on fractional the output and passed our norm checks. When additional lesser thinking efforts are deployed in production, that’s client-ready activity delivered efficiently.”

CompanyDeloitte Consulting LLP

AuthorCarl Bennett, CIO

Communication

We’ve made important improvements to the way Opus 5.5 writes and communicates, among the most average areas of feedback we heard concerning Opus 5. Its messages are much easier to comprehend at a glance, which testers stated helped during lengthy operating sessions. It puts the most crucial data up front, is small apt to use terminology or idiosyncratic phrases, and follows the penning rules you provision it. We discover that this makes Opus 5.5 a noticeably improved collaborator. Here’s a side-by-side difference of the two models:

Claude Opus 5

Please explain the matter to me.

What I found

The extra drop isn’t the free tier it’s a regression in 0552feb (“extract billing period boundaries into periods.py”, labelled “No behaviour change”).

The bug

aggregate.monthly_usage used to do a half-open interval:

start, end = month_start(year, month), next_month_start(year, month) if start <= e.ts < end:

It now does:

if period.contains(e.ts)

and in periods.py:

@property def last_day(self) -> datetime: return next_month_start(...) - timedelta(days=1) # 2026-08-31 00:00:00 UTC def contains(self, ts): return self.first_day <= ts <= self.last_day

last_day is a datetime at midnight, not a whole day. Subtracting one day from Sep 1 00:00 gives Aug 31 00:00, and the comparison is inclusive, so the period effectively ends at the first instant of the 31st. Every event timestamped after midnight on the last day of the month is silently dropped roughly one day of usage per account, per month.

The new test_periods.py doesn’t catch it because it only checks last_day.date() (which is correctly 2026-08-31) and tests contains at Aug 1 00:00, Aug 15 12:00, Jul 31 23:59 and Sep 1 00:00 it never probes anywhere inside the 31st.

Claude Opus 5.5

Please explain the matter to me.

The extra drop is a bug in the billing refactor

The free-tier change accounts for only $1.50 of acme’s August drop. The other $9.92 comes from a bug in commit 0552feb. That commit was labelled “No behaviour change,” but it stops counting usage from the last day of the month.

What changed

Before the merge, aggregate.py used a half-open interval:

start <= e.ts < next_month_start(year, month) # [Aug 1 00:00, Sep 1 00:00)

The refactor moved this into Period.contains and made the interval closed on the wrong boundary:

last_day = next_month_start(...) - timedelta(days=1) # = 2026-08-31T00:00:00Z return self.first_day <= ts <= self.last_day

last_day is a timestamp at midnight, not a whole day. So every event after 2026-08-31T00:00:00Z is excluded. The September period starts Sep 1, so it doesn’t pick these events up either. The usage isn’t moved to another month; it’s never billed at all.

Our customers’ feedback supports these findings:

RampStripeBoxChicago Trading CompanyFactory

Quote

“Verbose, hard-to-follow output has been my biggest frustration alongside frontier models, and Claude Opus 5.5 fixes it. It writes akin a fine colleague, and follows our penning rules. A scheme spec came out usable alongside extremely minimal edits, and whenever it rewrote one of our prompts I preferred its type to my own. When it optimized our test suite, I could prosecute its reasoning effortlessly and shipped the alter alongside confidence.”

CompanyRamp

AuthorJohn Ruelas, Staff Software Engineer

Safety

Pacing the frontier

Last week, our CEO, Dario Amodei, asserted that AI advancement have to be paced so that safety practices remain onward of example capabilities. Pacing is an method to keeping AI safe, remaining rivalrous alongside China, and realizing AI’s benefits, particularly in areas akin existence discipline and medicine.

We mostly comprehend the risks today’s models current and are fine armed to oversee them. However, additional grave risks could appear quickly as capabilities improve, and we need to prepared for them now. For that reason, our safety activity takes location on two period horizons at once:

Safety practices for current models. The current generation of models relies on an established set of practices: extended alignment testing, pre-release evaluation by exterior organizations specified as METR and Frontier Design, and safeguards matched to all model’s capabilities in high-risk areas akin cybersecurity and biology. We refine these practices alongside all release. We accept they are suitable to the worst risks today’s models present, and that they provision us a broad, although not perfect, image of the range of grave risks.

Additionally, we track our capability to train and measure aligned models, and we study on the two our community and inner models in the hazard reports we publish under our Responsible Scaling Policy, our voluntary example for managing catastrophic risks from advanced AI systems.

Preparing for forthcoming models. We’re preparing our training and evaluation processes in anticipation of additional advanced models. We’re tightening how we display the environments used in reinforcement learning, since flawed environments are a important origin of misaligned behavior. Additionally, we’re improving our alignment rewards and evolving automated processes for producing new, varied scenarios for safety training. And we are strengthening our security and monitoring, including a focused attempt to enhance interpretability-based monitoring and evaluation. We anticipation specified techniques volition assistance decrease our reliance on auditing a model’s chain-of-thought, or the reasoning it writes out during it works.

Models alongside greater capabilities—such as those that can completely automate the activity of AI investigation itself—require a higher safety norm still. Our calls for pacing were based in ample part on our anticipation that specified models could be trained soon. For these models, we do not assume the measures described complete volition encounter that safety norm on their own. As AI becomes additional capable, community guideline should perform a larger function in making certain the systems group depend on are safe. That capability takes period to build, and we’ve started to put the infrastructure in location to assistance it, as described in “We Must Pace the Frontier” and our latest notice alongside Accenture; we anticipate to portion additional particulars on these efforts soon. We volition additionally continue to contribute to guideline discussions alongside authorities and industry, including on approaches to regulation and global coordination.

Alignment

On our chief evaluation suite, an automated behavioral audit that assesses Claude throughout nearly 2,000 scenarios, Opus 5.5 scored improved than any latest Claude example on nearly all measure of misaligned behavior. It’s additionally our strongest example on most measures of honesty.

In particular, Opus 5.5 improves complete former models on multiple of the behaviors that contributed to latest cybersecurity incidents, including biased or motivated reasoning, attempting to escape a sandbox, and taking harmful actions following concluding it was in a simulated environment. In a new evaluation designed to test a model’s propensity to cross containment boundaries, Opus 5.5 attempted to circumvent boundaries about 85% small frequently than Opus 5 or Claude Mythos 5.1, and all attempt it made was low severity and self-reported. For teams operating Claude unattended throughout their codebases and systems, this is fair as crucial as raw capability.

However, as we described in our latest alignment assessment, construction evaluations that reliably capture all nonaccomplishment previous to deployment remains an unsolved problem. We see signs that Opus 5.5 frequently suspects it is being evaluated, which challenges our capability to measure how it volition act in the huge assortment of real-world settings it is deployed in. As these settings develop and example capabilities increase, we anticipate this difficulty to grow, unless we create advancement on interpretability. Although we are assured that Opus 5.5 shows broad improvements in the areas we are capable to measure, we brace our own alignment activity alongside the safeguards described below.

Safeguards

As our models develop additional powerful, stricter safeguards are one way we forestall new capabilities from becoming tools for misuse. Opus 5.5 is the archetypal Opus example to initiate alongside a akin category of safeguards to Fable 5.1 on cybersecurity, biology, and distillation, all of which autumn rear to another example transparently.

Cybersecurity. Because Opus 5.5 has extremely powerful cyber capabilities, we’re applying cybersecurity safeguards to Opus 5.5 that are akin to Fable 5.1’s. Users volition be capable to acknowledge and fix bugs in their code as part of the regular application betterment lifecycle, but most cybersecurity tasks volition be re-routed to Opus 4.8.

For cyberdefenders, we’ll shortly be expanding our Cyber Verification Program to contain Opus 5.5. The new program volition contain three tiers for increasingly permissive trusted access, including admission to Claude Mythos models. Claude Security is already accessible alongside admission to Claude Mythos 5.1.

Biology. Opus 5.5 is extremely capable in biology, exceeding Opus 5 and matching or beating Claude Mythos 5.1 throughout many areas of work. For example, Opus 5.5 achieved improvements on a long-horizon molecular prediction and scheme evaluation conducted in collaboration alongside Dyno Therapeutics, and expert red-teamers rated its specialized novelty as comparable to the finest example they had tested.

For this reason, Opus 5.5 uses the identical existence discipline safeguards as Fable 5.1. To use Opus 5.5 for investigation and betterment activity impeded by these safeguards, users can use to our new Life Sciences Verification Program, which gives vetted organizations akin scholarly labs, startups, and pharmaceutical companies admission to safeguards designed for the complete breadth of biology-related work. Interested organizations can use here.

Distillation

Distillation attacks, in which attackers use thousands of counterfeit accounts to extract a model’s capabilities at manufacturing scale, create safety and national safety risks. Distillation allows bad actors to create extremely capable models without the safeguards we build into Claude. Our September 2026 danger intellect report particulars the illicit distillation action we’ve detected and disrupted so far.

Opus 5.5 is launching alongside preserved thinking, the anti-distillation safeguard we introduced alongside Fable 5.1. It stops API users from editing Claude’s previous environment in an attempt to extract Claude’s reasoning. It applies to Fable 5.1 and Opus 5.5 for API accounts created on or following August 31, 2026. Our Help Center article explains the change, and our preserved thinking docs display how to test and update your integrations.

Data preservation and compliance

Like former Opus models, Opus 5.5 is accessible alongside zero data retention.

As alongside Fable 5.1, Opus 5.5 comes alongside our watermarking measures to comply alongside the EU AI Act, discussed here. It is additionally no longer accessible alongside “thinking” manner switched off, as we depict here.

Availability

Claude Opus 5.5 is now accessible on all platforms, including Amazon Web Services, Google Cloud, and Microsoft Azure. On the Claude Platform, developers can get started alongside claude-opus-5-5.

Other Article Hacker News
Close Right Ads
Close Left Ads