Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases

Sep 13, 2026 03:25 AM - 8 hours ago 5

September 2026

Benchmarking frontier AI models connected private, real-world, endeavor codebases.

ResultsAnalysisEffortSetup

01Introduction

Today we are releasing Real-SWE, a benchmark that evaluates frontier AI models connected private, real-world, endeavor codebases. Each task comes from a backstage accumulation codebase that we licensed from a real-world company. These are problems their engineers activity on, pinch each the discourse and complexity that comes pinch an existing product.

  • Private codebases. Agents must navigate proprietary systems whose codification and solutions aren’t disposable connected the nationalist internet.
  • Work pinch business consequences. Getting billing right, calculating taxes, migrating customers. Changes that impact really a business runs, often crossed aggregate services.
  • Company-specific complexity. Every institution has its ain rules and ways of penning code. Agents person to understand those conventions and make changes that activity pinch what’s already there.

Can a coding supplier really do the activity of a package technologist successful the existent world?

  1. 1

    Fable 5.1

    Claude Code

    Resolution rate: 38.8%

  2. 2

    GPT-6 Astra

    Codex CLI

    Resolution rate: 33.8%

  3. 3

    Gemini 3.8 Flash

    Gemini CLI

    Resolution rate: 31.2%

  4. 4

    GLM 5.3

    Claude Code

    Resolution rate: 28.8%

  5. =5

    Grok 4.6

    Grok Build

    Resolution rate: 23.8%

  6. =5

    Muse Spark 1.3

    Muse Code

    Resolution rate: 23.8%

  7. 7

    Kimi K3

    Kimi Code

    Resolution rate: 18.8%

  8. 8

    GPT-5.6 Sol

    Codex CLI

    Resolution rate: 16.2%

Resolution complaint is balanced to pass@1, averaged complete 8 independent runs per task. 95% assurance intervals are shown.

Expert-generated aliases synthetic tasks tin beryllium good designed, but they aren’t the verbatim, existent tasks that engineers successful existent companies request to do. Our tasks disagree connected 2 axes: the underlying coding artifact and specificity of the instruction. Both adhd complexities that situation today’s frontier models.

We usage autochthonal harnesses to bespeak really endeavor engineers activity successful practice, evaluating model-and-harness combinations alternatively than models successful isolation.

Real institution tasks require company-specific context

Correct billing depends connected business rules and outer services

Fix invoice billing truthful each business charges the correct taxation and exempt customers aren't taxed.

View afloat instructionHide afloat instruction

Billing reopens connected Monday and each invoice this work issues is coming retired untaxed. Each business connected the level settles its taxation a different way: immoderate support a complaint themselves, immoderate want each invoice priced against the buyer's destination by our taxation authority provider, and immoderate cod thing astatine all, while a customer we clasp an exemption for is charged thing whichever measurement its business is configured. Pricing a destination intends going to the authority pinch some addresses, the priced lines and the merchandise class that business sells under, connected the sandbox aliases the accumulation authority according to the relationship the business is on; an reside the authority refuses must beryllium reported without stopping the invoice. The rate, the taxation and the gross beryllium connected the issued invoice, and erstwhile an invoice is settled the waste is revenge backmost to the authority nether that invoice's number truthful the returns reconcile. Invoices betwixt European parties show some sides' VAT registrations. The authority and ledger are disposable astatine TAX_JAR_URL, PROD_TAX_JAR_URL and INFLUX_URL.

Services successful the sandbox

  • TaxJar sandbox
  • TaxJar production
  • InfluxDB ledger
  • NestJS service
  • TypeScript

Agents activity crossed code, infrastructure, and business tools

Tools and services crossed Real-SWE task environments. Each task exposes only the services its workflow needs.

  • AWS emulator
  • Docker
  • Kubernetes
  • GitHub
  • Linear MCP
  • PostgreSQL
  • MySQL
  • MongoDB
  • Gel
  • Redis
  • Go
  • Python
  • Node.js
  • Vitest
  • Slack
  • Intercom
  • Google Drive
  • Email
  • ClickUp

Codebase Selection

We selected codebases done a rigorous screening process, focusing connected existent companies pinch important usage, beardown engineering teams, and demanding accumulation workloads. The sample tasks analyzed beneath travel from these codebases, including:

  • A Luma/Partiful competitor pinch 200K+ users and a top 100 App Store ranking
  • A user fintech level processing 100K+ slope statements
  • Enterprise AI income platforms supporting analyzable business workflows

We prioritize codification written to meet an existent personification aliases business request complete codification written solely to create a benchmark task. Production engineering requires knowing existing architecture, preserving behaviour that users trust on, and making changes wrong existent operational constraints.

Brief instructions tin require changes crossed galore files

Our tasks picture the alteration needed, leaving agents to observe implementation specifications successful the codebase and surrounding tools. Any behaviour required by the verifier must beryllium stated aliases reasonably discoverable. This leads to our prompts being somewhat underspecified, astir par pinch DeepSWE and Terminal Bench, but circumstantial capable to not omit instructions.

The activity is cross-functional and complex: a azygous alteration tin span aggregate parts of the application. Agents must understand existing business logic and institution coding patterns while keeping the surrounding strategy working.

Prompt magnitude · median

A emblematic Real-SWE instruction is 1,742 characters.

  • FrontierCode2,056 chars

  • DeepSWE1,975 chars

  • Terminal-Bench 31,584 chars

  • FrontierSWE v2992 chars

  • Real-SWE1,742 chars

Files edited by the reference solution · median

11 files successful Real-SWE, compared pinch 6 successful FrontierCode and DeepSWE.

  • FrontierCode6

  • DeepSWE6

  • Real-SWE11

All figures are medians. FrontierCode and DeepSWE usage Cognition's published comparison; FrontierCode includes task descriptions and codebase guidelines. We measured instruction files from Terminal-Bench 3's 74 tasks, FrontierSWE v2's 34 tasks, and Real-SWE's 8 repository-backed sample tasks. Character counts are rounded to the nearest full character. No comparable files-edited fig is included for Terminal-Bench 3 aliases FrontierSWE v2.

Models neglect moreover successful short rollouts.

71.4% of rollouts nether 10 minutes failed, compared pinch 73.4% of longer rollouts.

Triaging aggregate systems and knowing requirements successful codebases riddled pinch existing business logic and coding patterns is difficult.

Under 10 minUnder 10 min: 70 grounded (71.4%) and 28 passed (28.6%), retired of 98 rollouts.71.4%28.6%

70/98 failed

10 min aliases longer10 min aliases longer: 398 grounded (73.4%) and 144 passed (26.6%), retired of 542 rollouts.73.4%26.6%

398/542 failed

  • Failed
  • Passed

Every task is inspired aliases lifted verbatim from a private, real-world codebase. We find these types of tasks ace absorbing for 3 reasons:

  1. Tasks connected backstage codebases are natively retired of distribution. These types of coding tasks are not disposable anyplace connected the net and are improbable to person ever been trained connected by immoderate different ai model. 99% of tokens successful real-world enterprises are hidden distant from the frontier models.
  2. These tasks are economically viable work. Each task present has a nonstop narration to walk and was assigned to an technologist earning a salary. Most benchmarks trial interesting, experimental capabilities that are often improbable to beryllium wide successful the real-world.
  3. Company-specific engineering patterns matter. Does AI codification lucifer the barroom of a real-world enterprise? Our results show america that we're acold from that reality. Many enterprises attraction astir codification standards and patterns. We've recovered that today's models are weaker astatine knowing institution coding patterns and often miss requirements aliases don't verify their assumptions.

02Analysis

Here's an study of a mini sample of tasks from our benchmark. If you're willing successful the sample, request entree here.

6 of 10 tasks person solution rates beneath 15%

Each task had 8 rollouts per model.

Missed requirements are the astir communal failure

Failures are grouped by observed submission behaviour utilizing the aforesaid taxonomy crossed models, pursuing DeepSWE.

Fable 5.1

GPT-6 Astra

Gemini 3.8 Flash

GLM 5.3

Grok 4.6

Muse Spark 1.3

Kimi K3

GPT-5.6 Sol

Unverified assumptionMissed requirementIntegration errorRegressionWrong file

No exemplary solves each task

One quadrate per rollout: each statement is simply a task, each file a trial, 8 tests per task for each model.

Fable 5.1

01

02

03

04

05

06

07

08

09

10

GPT-6 Astra

01

02

03

04

05

06

07

08

09

10

Gemini 3.8 Flash

01

02

03

04

05

06

07

08

09

10

GLM 5.3

01

02

03

04

05

06

07

08

09

10

Grok 4.6

01

02

03

04

05

06

07

08

09

10

Muse Spark 1.3

01

02

03

04

05

06

07

08

09

10

Kimi K3

01

02

03

04

05

06

07

08

09

10

GPT-5.6 Sol

01

02

03

04

05

06

07

08

09

10

PassUnverified assumptionMissed requirementIntegration errorRegressionWrong file

Different models neglect successful different ways

Percentages are retired of each model's grounded runs, not each runs.

03Effort & the Frontier

Higher costs does not guarantee a higher solution rate

Estimated frontier

Resolution complaint (%)1015202530354045$2$3$5$10Cost per rollout (USD, log scale)Gemini 3.8 Flash: 31.2% · $2.50; Gemini CLIGemini 3.8 Flash31.2% · $2.50GPT-5.6 Sol: 16.2% · $2.65; Codex CLIGPT-5.6 Sol16.2% · $2.65Muse Spark 1.3: 23.8% · $2.74; Muse CodeMuse Spark 1.323.8% · $2.74Grok 4.6: 23.8% · $3.44; Grok Build; incomplete usage, existent costs whitethorn beryllium higherGrok 4.623.8% · $3.44Kimi K3: 18.8% · $3.90; Kimi Code; incomplete usage, existent costs whitethorn beryllium higherKimi K318.8% · $3.90GPT-6 Astra: 33.8% · $4.67; Codex CLIGPT-6 Astra33.8% · $4.67GLM 5.3: 28.8% · $5.12; Claude CodeGLM 5.328.8% · $5.12Fable 5.1: 38.8% · $6.96; Claude CodeFable 5.138.8% · $6.96

Estimated rollout costs scope from $2.50 to $6.96

RankModelEstimated costs (USD)Gemini 3.8 FlashGPT-5.6 SolMuse Spark 1.3Grok 4.6Kimi K3GPT-6 AstraGLM 5.3Fable 5.1
1$2.50
2$2.65
3$2.74
4$3.44
5$3.90
6$4.67
7$5.12
8$6.96

mean per rollout, by task

Swipe the floor plan to spot each tasks.

0100k200k300k400k01020304050607080910taskEntitlement overage lines · Fable 5.1: 34kMulti-region expanse · Fable 5.1: 30kTax jurisdiction · Fable 5.1: 78kAPI token metering · Fable 5.1: 95kAPI keys & environments · Fable 5.1: 71kS3 datastore measurement · Fable 5.1: 62kCustomer personality migration · Fable 5.1: 67kBilling schedule migration · Fable 5.1: 26kLinearizable scan · Fable 5.1: 86kAnalytics watercourse reducer · Fable 5.1: 88kEntitlement overage lines · GPT-6 Astra: 13kMulti-region expanse · GPT-6 Astra: 13kTax jurisdiction · GPT-6 Astra: 24kAPI token metering · GPT-6 Astra: 31kAPI keys & environments · GPT-6 Astra: 32kS3 datastore measurement · GPT-6 Astra: 22kCustomer personality migration · GPT-6 Astra: 25kBilling schedule migration · GPT-6 Astra: 15kLinearizable scan · GPT-6 Astra: 33kAnalytics watercourse reducer · GPT-6 Astra: 29kEntitlement overage lines · Gemini 3.8 Flash: 78kMulti-region expanse · Gemini 3.8 Flash: 67kTax jurisdiction · Gemini 3.8 Flash: 95kAPI token metering · Gemini 3.8 Flash: 134kAPI keys & environments · Gemini 3.8 Flash: 102kS3 datastore measurement · Gemini 3.8 Flash: 106kCustomer personality migration · Gemini 3.8 Flash: 97kBilling schedule migration · Gemini 3.8 Flash: 70kLinearizable scan · Gemini 3.8 Flash: 106kAnalytics watercourse reducer · Gemini 3.8 Flash: 88kEntitlement overage lines · GLM 5.3: 68kMulti-region expanse · GLM 5.3: 53kTax jurisdiction · GLM 5.3: 141kAPI token metering · GLM 5.3: 177kAPI keys & environments · GLM 5.3: 125kS3 datastore measurement · GLM 5.3: 121kCustomer personality migration · GLM 5.3: 90kBilling schedule migration · GLM 5.3: 58kLinearizable scan · GLM 5.3: 172kAnalytics watercourse reducer · GLM 5.3: 169kEntitlement overage lines · Grok 4.6: 7kMulti-region expanse · Grok 4.6: 3kTax jurisdiction · Grok 4.6: 12kAPI token metering · Grok 4.6: 15kAPI keys & environments · Grok 4.6: 16kS3 datastore measurement · Grok 4.6: 13kCustomer personality migration · Grok 4.6: 20kBilling schedule migration · Grok 4.6: 6kLinearizable scan · Grok 4.6: 261kAnalytics watercourse reducer · Grok 4.6: 315kEntitlement overage lines · Muse Spark 1.3: 36kMulti-region expanse · Muse Spark 1.3: 43kTax jurisdiction · Muse Spark 1.3: 67kAPI token metering · Muse Spark 1.3: 152kAPI keys & environments · Muse Spark 1.3: 104kS3 datastore measurement · Muse Spark 1.3: 71kCustomer personality migration · Muse Spark 1.3: 76kBilling schedule migration · Muse Spark 1.3: 38kLinearizable scan · Muse Spark 1.3: 141kAnalytics watercourse reducer · Muse Spark 1.3: 137kEntitlement overage lines · Kimi K3: 30kMulti-region expanse · Kimi K3: 9kTax jurisdiction · Kimi K3: 39kAPI token metering · Kimi K3: 69kAPI keys & environments · Kimi K3: 44kS3 datastore measurement · Kimi K3: 32kCustomer personality migration · Kimi K3: 66kBilling schedule migration · Kimi K3: 19kLinearizable scan · Kimi K3: 71kAnalytics watercourse reducer · Kimi K3: 55kEntitlement overage lines · GPT-5.6 Sol: 12kMulti-region expanse · GPT-5.6 Sol: 8kTax jurisdiction · GPT-5.6 Sol: 22kAPI token metering · GPT-5.6 Sol: 31kAPI keys & environments · GPT-5.6 Sol: 25kS3 datastore measurement · GPT-5.6 Sol: 25kCustomer personality migration · GPT-5.6 Sol: 24kBilling schedule migration · GPT-5.6 Sol: 13kLinearizable scan · GPT-5.6 Sol: 37kAnalytics watercourse reducer · GPT-5.6 Sol: 30k

Fable 5.1 · 64k overallGPT-6 Astra · 24k overallGemini 3.8 Flash · 94k overallGLM 5.3 · 117k overallGrok 4.6 · 67k overallMuse Spark 1.3 · 87k overallKimi K3 · 43k overallGPT-5.6 Sol · 23k overall

View task valuesTask
  • Fable 5.134k
  • GPT-6 Astra13k
  • Gemini 3.8 Flash78k
  • GLM 5.368k
  • Grok 4.67k
  • Muse Spark 1.336k
  • Kimi K330k
  • GPT-5.6 Sol12k

04Evaluation Setup

Each supplier was tally successful an isolated sandbox. All tasks are successful Harbor format, and verifiers are injected astatine grading time. The verifiers are inspired by existing trial suites successful the codebase aliases usage those tests verbatim.

More