Introducing Claude Sonnet 5.5, the second example in the Claude 5.5 family. It’s a apparent upgrade complete Claude Sonnet 5, runs 30%+ faster, and expenses up to 30% small for most work.
Sonnet 5.5 is a faster, lower-cost complement to Claude Opus 5.5. Where Opus 5.5 is built for complex activity requiring careful judgment, Sonnet 5.5 is strongest at well-scoped mundane tasks, fixing bugs, and creating polished documents, slides, and spreadsheets. It’s additionally got a keen eye for design. Claude Haiku 5.5, built for high-volume and cost-sensitive applications, volition associate the Claude 5.5 family in the coming weeks.
Sonnet 5.5 improves complete Sonnet 5 on:
Performance. Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0, an agentic coding evaluation, compared to Sonnet 5’s 10.3%. It scores two points below Opus 5.5 on GDPval-AA, a test of real-world activity throughout a assortment of occupations. And it’s powerful on long-horizon activity and depiction understanding—it’s the archetypal Sonnet example to attack Pokémon Red operating lone from screenshots.
Collaboration. Like Opus 5.5, Sonnet 5.5 writes additional plainly than our former generation of models; first testers described it as a improved partner for collaboration than Sonnet 5. Its speed additionally makes it fine suited to accelerated repeat on small complex tasks.
Cost. Sonnet 5.5 is priced the identical as Sonnet 5 at $2 per myriad input tokens, $10 per myriad output tokens, and $0.20 per myriad tokens for cache reads, but it typically needs far small tokens to do the identical work. In our testing, it expenses up to 30% small per project than its predecessor.
Speed. Sonnet 5.5 generates outputs 30%+ faster than Sonnet 5, making it our fastest Sonnet example to date.
Alignment and safety. On our automated behavioral audit, Sonnet 5.5 improves on or matches Sonnet 5 on most measures of alignment. Because its cybersecurity capabilities are comparable to Opus 5’s, it’s the archetypal Sonnet example to initiate alongside cyber safeguards and fallbacks akin those we’ve developed for our most capable models. Its existence discipline safeguards are the identical as Sonnet 5’s. Both safeguards mark a narrow set of high-risk requests; regular application betterment and most existence sciences activity are unaffected.
Performance
Sonnet 5.5 improves on Sonnet 5 throughout domains—in several cases dramatically. On multiple evaluations, Sonnet 5.5 at Max attempt equal performs comparably to Opus 5.5. However, benchmark scores grasp lone one aspect of a model’s capabilities; in our own testing, and in that of external testers, Opus 5.5 remains plainly stronger at complex, open-ended activity requiring sustained judgment.
| Sonnet 5.5 | Sonnet 5 | Opus 5.5 | GPT-6 Sol | |
|---|---|---|---|---|
| Agentic codingTerminal-Bench 4.0 | 70.6% | 10.3% | 66.4%¹ | — |
| Agentic codingFrontierCode 1.1 (Main) | 46.2%Max² | 42.4% | 54.4% | 49.3% |
| 52.1%Xhigh | ||||
| Agentic codingCursorBench 4.0 | 55.5% | 34.1% | 57.8% | — |
| Knowledge workGDPval-AA v2.1³ | 1844 | 1449 | 1846 | 1487⁴ |
| Knowledge workAA-Briefcase v1.1³ | 1811 | 1359 | 1822 | 1483⁴ |
| Multidisciplinary reasoningHumanity’s Last Exam | 64.5%with tools | 54.9%with tools | 67.7%with tools | — |
| Computer useOSWorld 2.1 | 80.1%partial | 57.0%partial | 81.8%partial | — |
| Visual diagram recognitionChartography | 61.6%no tools | 15.6%no tools | 64.4%no tools | 53.6%⁴no tools |
For particulars on how we run our evaluations, see the Sonnet 5.5 System Card.
The charts below storyline all model’s mark against its disbursal per project at all attempt level. As attempt goes up, models typically activity for longer, foremost to a higher disbursal per project but mostly additionally a higher score. The nearer a item is to the top remaining of the chart, the additional capability it delivers per dollar.
On multiple benchmarks, Sonnet 5.5 at Low or Medium attempt strikes Sonnet 5’s finest mark for concerning a tenth of the disbursal per task. It complements Opus 5.5 finest whenever operating at lesser attempt settings, anywhere it expenses small per task. At higher settings, it can execute comparably at a akin cost.
Agentic terminal codingAgentic coding: FrontierCodeAgentic coding: CursorBenchKnowledge work: AA-Briefcase
Coding
Sonnet 5.5’s jump in achievement is particularly noticeable in coding. At High attempt on FrontierCode, it scores 10 points higher than Sonnet 5 at the identical setting, at concerning one fifteenth of the disbursal per task. On CursorBench, which tests models on tasks from genuine Cursor coding sessions, its finest mark is inside concerning two points of Opus 5.5.
Early testers appreciated how quickly Sonnet 5.5 can comprehend a codebase. They were additionally struck by its efficiency: in head-to-head runs, it batched tool calls together additional than Sonnet 5, foremost to small steps and lesser costs.
Epic GamesEveryCodeRabbitSpaceXAIBase44UnityCreator
Quote
“In Epic’s first testing, Claude Sonnet 5.5 cleared the identical norm bar you’d anticipate from a higher-tier model, holding up on a scheme scheme audit and a data stream review. The new example managed tens of thousands of lines of code for gameplay scheme architecture, kept responses snappy, handled multi-hour tasks, and delivered alongside small prescriptive prompting.”
CompanyEpic Games
AuthorDaniel Vogel, Chief Operating Officer
Knowledge work
Sonnet 5.5 shows gains in multiple areas of cognition work. On GDPval-AA, which tests models on real-world tasks throughout 44 occupations and nine important industries, Sonnet 5.5 scores nearly flat alongside Opus 5.5 and concerning 400 points complete Sonnet 5. It’s near to Opus 5.5 in device use and diagram recognition, and plainly outperforms Sonnet 5 and GPT-6 Sol on long-horizon cognition work.
Early testers highlighted small quantifiable improvements. They established it to be a additional natural conversational partner and remarked on its knack for design, noting that it adds polish to person interfaces and can prosecute glide templates to create decks that necessitate minimal editing. In one inner test, we gave it a community company’s quarterly earnings materials and call transcripts, alongside alongside a glide template, and asked for a 10-slide functioning review. Two experts judged its archetypal outline to be prepared to dispatch as is.
SlackZendeskBalyasny Asset ManagementBoxLovableAtlassian
Quote
“Without changing any of our prompts, Claude Sonnet 5.5 did improved than Sonnet 5 on nearly all of our offline Slackbot evals, in small steps and alongside concerning 14% small output tokens. When person gives Slackbot a task, norm and speed are what matter most, and Sonnet 5.5 allows Slackbot to provision improved outcomes for users, faster.”
CompanySlack
AuthorCurtis Allen, Principal Engineer
Cost and speed
Pricing
| Price per 1M tokens | Claude Sonnet 5.5 | Claude Opus 5.5 |
|---|---|---|
| Cache reads | $0.20 | $0.20 |
| Cache writes | $2.50 | $5 |
| Input tokens | $2 | $4 |
| Output tokens | $10 | $20 |
Sonnet 5.5 requires small tokens per project than Sonnet 5, so it’s small costly to run. It additionally generates output 30%+ faster, and its effectiveness is immediately noticeable:
Prompt:
A murmuration of 400 starlings in one HTML file
Claude Sonnet 5
Claude Sonnet 5.5
Adjusting the effort level lets you balance disbursal and speed against general quality. In Claude Code and our apps, the default attempt is set to Medium, during the Claude Platform defaults to High. At lesser settings, Claude answers faster and uses small tokens, which suits regular work. At higher settings, Claude reasons for longer and checks its activity additional thoroughly.
Safety
Alignment
Sonnet 5.5 doesn’t advancement the frontier of our models’ capabilities, so our alignment appraisal focused on a targeted set of risks that use to models of any capability level, including acting against users’ interests, misleading users, and cooperating alongside high-stakes misuse.
On our automated behavioral audit, which tests Claude throughout approximately 1,850 scenarios, Sonnet 5.5 improves on or matches Sonnet 5 on most measures of alignment, opposition to misuse, and honesty. On our newer containment evaluations, Sonnet 5.5 comes near to Opus 5.5, the finest example we tested, in how rarely it tries to escape its sandbox, and it’s the smallest apt of any of our models to investigation the limits of its containers. Across the complete audit, Opus 5.5 motionless performs slightly improved overall, but we established no evidence that Sonnet 5.5 pursues goals that battle alongside the user’s intention.
As we described in our recent alignment assessment, no set of evaluations reliably catches all failure, and Sonnet 5.5 may have tendencies we haven’t found, which is why we brace our own alignment activity alongside the safeguards described below.
Safeguards
Cybersecurity. Sonnet 5.5’s cyber capabilities are a ample betterment complete Sonnet 5’s, so we’re deploying it alongside safeguards akin to those on Opus 5.5. Users can motionless discover and fix bugs in their code as part of regular application development, but higher-risk cybersecurity tasks volition visibly autumn rear to Sonnet 5. Soon, cyberdefenders volition be capable to use to our expanded Cyber Verification Program for tiered admission to additional advanced capabilities on Sonnet 5.5, Opus 5.5, and Claude Mythos models.
Biology. Sonnet 5.5 uses the identical set of existence discipline safeguards as Sonnet 5. These mark harmful requests; most research, education, and medicinal activity is unaffected, although several microbiology and virology requests may be flagged in error. Organizations can use to our Life Sciences Verification Program for admission to safeguards designed for the complete breadth of biology-related work.
Distillation. Distillation attacks, in which attackers use thousands of counterfeit accounts to extract a model’s capabilities at manufacturing scale, authorize bad actors to create extremely capable models without the safeguards we build into Claude. Because Sonnet 5.5 is far additional capable than its predecessor, it’s the archetypal Sonnet example to initiate alongside safety classifiers that forestall reasoning extraction. Sonnet 5.5 additionally expands preserved thinking, so Claude’s thinking cannot be decoupled from the document that created it. Most developers won’t notice a change. If you move conversations between accounts, including switching accounts mid-session in Claude Code, our docs article explains the change.
Getting started
As alongside Opus 5.5 and Sonnet 5, Claude Sonnet 5.5 is accessible alongside zero data retention.
Claude Sonnet 5.5 is now accessible on all platforms, including Amazon Web Services, Google Cloud, and Microsoft Azure. Developers can get started on the Claude Platform alongside claude-sonnet-5-5. If you run Sonnet alongside thinking off, you’ll need to toggle to the new between_tools setting, which keeps up-front thinking off, before moving to Sonnet 5.5. See our migration guide for details.