September 2026
Benchmarking frontier AI models connected private, real-world, endeavor codebases.
01Introduction
Today we are releasing Real-SWE, a benchmark that evaluates frontier AI models connected private, real-world, endeavor codebases. Each task comes from a backstage accumulation codebase that we licensed from a real-world company. These are problems their engineers activity on, pinch each the discourse and complexity that comes pinch an existing product.
- Private codebases. Agents must navigate proprietary systems whose codification and solutions aren’t disposable connected the nationalist internet.
- Work pinch business consequences. Getting billing right, calculating taxes, migrating customers. Changes that impact really a business runs, often crossed aggregate services.
- Company-specific complexity. Every institution has its ain rules and ways of penning code. Agents person to understand those conventions and make changes that activity pinch what’s already there.
Can a coding supplier really do the activity of a package technologist successful the existent world?
1
Fable 5.1
Claude Code
Resolution rate: 38.8%
2

GPT-6 Astra
Codex CLI
Resolution rate: 33.8%
3

Gemini 3.8 Flash
Gemini CLI
Resolution rate: 31.2%
4
GLM 5.3
Claude Code
Resolution rate: 28.8%
=5
Grok 4.6
Grok Build
Resolution rate: 23.8%
=5
Muse Spark 1.3
Muse Code
Resolution rate: 23.8%
7
Kimi K3
Kimi Code
Resolution rate: 18.8%
8

GPT-5.6 Sol
Codex CLI
Resolution rate: 16.2%
Expert-generated aliases synthetic tasks tin beryllium good designed, but they aren’t the verbatim, existent tasks that engineers successful existent companies request to do. Our tasks disagree connected 2 axes: the underlying coding artifact and specificity of the instruction. Both adhd complexities that situation today’s frontier models.
We usage autochthonal harnesses to bespeak really endeavor engineers activity successful practice, evaluating model-and-harness combinations alternatively than models successful isolation.
Real institution tasks require company-specific context
Correct billing depends connected business rules and outer services
Fix invoice billing truthful each business charges the correct taxation and exempt customers aren't taxed.
Billing reopens connected Monday and each invoice this work issues is coming retired untaxed. Each business connected the level settles its taxation a different way: immoderate support a complaint themselves, immoderate want each invoice priced against the buyer's destination by our taxation authority provider, and immoderate cod thing astatine all, while a customer we clasp an exemption for is charged thing whichever measurement its business is configured. Pricing a destination intends going to the authority pinch some addresses, the priced lines and the merchandise class that business sells under, connected the sandbox aliases the accumulation authority according to the relationship the business is on; an reside the authority refuses must beryllium reported without stopping the invoice. The rate, the taxation and the gross beryllium connected the issued invoice, and erstwhile an invoice is settled the waste is revenge backmost to the authority nether that invoice's number truthful the returns reconcile. Invoices betwixt European parties show some sides' VAT registrations. The authority and ledger are disposable astatine TAX_JAR_URL, PROD_TAX_JAR_URL and INFLUX_URL.
Services successful the sandbox
- TaxJar sandbox
- TaxJar production
- InfluxDB ledger
- NestJS service
- TypeScript
Agents activity crossed code, infrastructure, and business tools
Tools and services crossed Real-SWE task environments. Each task exposes only the services its workflow needs.
- AWS emulator
- Docker
- Kubernetes
- GitHub
- Linear MCP
- PostgreSQL
- MySQL
- MongoDB
- Gel
- Redis
- Go
- Python
- Node.js
- Vitest
- Slack
- Intercom
- Google Drive
- ClickUp
Codebase Selection
We selected codebases done a rigorous screening process, focusing connected existent companies pinch important usage, beardown engineering teams, and demanding accumulation workloads. The sample tasks analyzed beneath travel from these codebases, including:
- A Luma/Partiful competitor pinch 200K+ users and a top 100 App Store ranking
- A user fintech level processing 100K+ slope statements
- Enterprise AI income platforms supporting analyzable business workflows
We prioritize codification written to meet an existent personification aliases business request complete codification written solely to create a benchmark task. Production engineering requires knowing existing architecture, preserving behaviour that users trust on, and making changes wrong existent operational constraints.
Brief instructions tin require changes crossed galore files
Our tasks picture the alteration needed, leaving agents to observe implementation specifications successful the codebase and surrounding tools. Any behaviour required by the verifier must beryllium stated aliases reasonably discoverable. This leads to our prompts being somewhat underspecified, astir par pinch DeepSWE and Terminal Bench, but circumstantial capable to not omit instructions.
The activity is cross-functional and complex: a azygous alteration tin span aggregate parts of the application. Agents must understand existing business logic and institution coding patterns while keeping the surrounding strategy working.
Prompt magnitude · median
A emblematic Real-SWE instruction is 1,742 characters.
FrontierCode2,056 chars
DeepSWE1,975 chars
Terminal-Bench 31,584 chars
FrontierSWE v2992 chars
Real-SWE1,742 chars
Files edited by the reference solution · median
11 files successful Real-SWE, compared pinch 6 successful FrontierCode and DeepSWE.
FrontierCode6
DeepSWE6
Real-SWE11
Models neglect moreover successful short rollouts.
71.4% of rollouts nether 10 minutes failed, compared pinch 73.4% of longer rollouts.
Triaging aggregate systems and knowing requirements successful codebases riddled pinch existing business logic and coding patterns is difficult.
70/98 failed
398/542 failed
- Failed
- Passed
Every task is inspired aliases lifted verbatim from a private, real-world codebase. We find these types of tasks ace absorbing for 3 reasons:
- Tasks connected backstage codebases are natively retired of distribution. These types of coding tasks are not disposable anyplace connected the net and are improbable to person ever been trained connected by immoderate different ai model. 99% of tokens successful real-world enterprises are hidden distant from the frontier models.
- These tasks are economically viable work. Each task present has a nonstop narration to walk and was assigned to an technologist earning a salary. Most benchmarks trial interesting, experimental capabilities that are often improbable to beryllium wide successful the real-world.
- Company-specific engineering patterns matter. Does AI codification lucifer the barroom of a real-world enterprise? Our results show america that we're acold from that reality. Many enterprises attraction astir codification standards and patterns. We've recovered that today's models are weaker astatine knowing institution coding patterns and often miss requirements aliases don't verify their assumptions.
02Analysis
Here's an study of a mini sample of tasks from our benchmark. If you're willing successful the sample, request entree here.
6 of 10 tasks person solution rates beneath 15%
Each task had 8 rollouts per model.Missed requirements are the astir communal failure
Failures are grouped by observed submission behaviour utilizing the aforesaid taxonomy crossed models, pursuing DeepSWE.
Fable 5.1
GPT-6 Astra
Gemini 3.8 Flash
GLM 5.3
Grok 4.6
Muse Spark 1.3
Kimi K3
GPT-5.6 Sol
Unverified assumptionMissed requirementIntegration errorRegressionWrong file
No exemplary solves each task
One quadrate per rollout: each statement is simply a task, each file a trial, 8 tests per task for each model.
Fable 5.1
01
02
03
04
05
06
07
08
09
10
GPT-6 Astra
01
02
03
04
05
06
07
08
09
10
Gemini 3.8 Flash
01
02
03
04
05
06
07
08
09
10
GLM 5.3
01
02
03
04
05
06
07
08
09
10
Grok 4.6
01
02
03
04
05
06
07
08
09
10
Muse Spark 1.3
01
02
03
04
05
06
07
08
09
10
Kimi K3
01
02
03
04
05
06
07
08
09
10
GPT-5.6 Sol
01
02
03
04
05
06
07
08
09
10
PassUnverified assumptionMissed requirementIntegration errorRegressionWrong file
Different models neglect successful different ways
Percentages are retired of each model's grounded runs, not each runs.
03Effort & the Frontier
Higher costs does not guarantee a higher solution rate
Estimated frontier
Resolution complaint (%)1015202530354045$2$3$5$10Cost per rollout (USD, log scale)Gemini 3.8 Flash: 31.2% · $2.50; Gemini CLIGemini 3.8 Flash31.2% · $2.50GPT-5.6 Sol: 16.2% · $2.65; Codex CLIGPT-5.6 Sol16.2% · $2.65Muse Spark 1.3: 23.8% · $2.74; Muse CodeMuse Spark 1.323.8% · $2.74Grok 4.6: 23.8% · $3.44; Grok Build; incomplete usage, existent costs whitethorn beryllium higherGrok 4.623.8% · $3.44Kimi K3: 18.8% · $3.90; Kimi Code; incomplete usage, existent costs whitethorn beryllium higherKimi K318.8% · $3.90GPT-6 Astra: 33.8% · $4.67; Codex CLIGPT-6 Astra33.8% · $4.67GLM 5.3: 28.8% · $5.12; Claude CodeGLM 5.328.8% · $5.12Fable 5.1: 38.8% · $6.96; Claude CodeFable 5.138.8% · $6.96
Estimated rollout costs scope from $2.50 to $6.96
| 1 | Gemini 3.8 Flash$2.50 |
| 2 | GPT-5.6 Sol$2.65 |
| 3 | Muse Spark 1.3$2.74 |
| 4 | Grok 4.6$3.44 |
| 5 | Kimi K3$3.90 |
| 6 | GPT-6 Astra$4.67 |
| 7 | GLM 5.3$5.12 |
| 8 | Fable 5.1$6.96 |
mean per rollout, by task
Swipe the floor plan to spot each tasks.
0100k200k300k400k01020304050607080910taskEntitlement overage lines · Fable 5.1: 34kMulti-region expanse · Fable 5.1: 30kTax jurisdiction · Fable 5.1: 78kAPI token metering · Fable 5.1: 95kAPI keys & environments · Fable 5.1: 71kS3 datastore measurement · Fable 5.1: 62kCustomer personality migration · Fable 5.1: 67kBilling schedule migration · Fable 5.1: 26kLinearizable scan · Fable 5.1: 86kAnalytics watercourse reducer · Fable 5.1: 88kEntitlement overage lines · GPT-6 Astra: 13kMulti-region expanse · GPT-6 Astra: 13kTax jurisdiction · GPT-6 Astra: 24kAPI token metering · GPT-6 Astra: 31kAPI keys & environments · GPT-6 Astra: 32kS3 datastore measurement · GPT-6 Astra: 22kCustomer personality migration · GPT-6 Astra: 25kBilling schedule migration · GPT-6 Astra: 15kLinearizable scan · GPT-6 Astra: 33kAnalytics watercourse reducer · GPT-6 Astra: 29kEntitlement overage lines · Gemini 3.8 Flash: 78kMulti-region expanse · Gemini 3.8 Flash: 67kTax jurisdiction · Gemini 3.8 Flash: 95kAPI token metering · Gemini 3.8 Flash: 134kAPI keys & environments · Gemini 3.8 Flash: 102kS3 datastore measurement · Gemini 3.8 Flash: 106kCustomer personality migration · Gemini 3.8 Flash: 97kBilling schedule migration · Gemini 3.8 Flash: 70kLinearizable scan · Gemini 3.8 Flash: 106kAnalytics watercourse reducer · Gemini 3.8 Flash: 88kEntitlement overage lines · GLM 5.3: 68kMulti-region expanse · GLM 5.3: 53kTax jurisdiction · GLM 5.3: 141kAPI token metering · GLM 5.3: 177kAPI keys & environments · GLM 5.3: 125kS3 datastore measurement · GLM 5.3: 121kCustomer personality migration · GLM 5.3: 90kBilling schedule migration · GLM 5.3: 58kLinearizable scan · GLM 5.3: 172kAnalytics watercourse reducer · GLM 5.3: 169kEntitlement overage lines · Grok 4.6: 7kMulti-region expanse · Grok 4.6: 3kTax jurisdiction · Grok 4.6: 12kAPI token metering · Grok 4.6: 15kAPI keys & environments · Grok 4.6: 16kS3 datastore measurement · Grok 4.6: 13kCustomer personality migration · Grok 4.6: 20kBilling schedule migration · Grok 4.6: 6kLinearizable scan · Grok 4.6: 261kAnalytics watercourse reducer · Grok 4.6: 315kEntitlement overage lines · Muse Spark 1.3: 36kMulti-region expanse · Muse Spark 1.3: 43kTax jurisdiction · Muse Spark 1.3: 67kAPI token metering · Muse Spark 1.3: 152kAPI keys & environments · Muse Spark 1.3: 104kS3 datastore measurement · Muse Spark 1.3: 71kCustomer personality migration · Muse Spark 1.3: 76kBilling schedule migration · Muse Spark 1.3: 38kLinearizable scan · Muse Spark 1.3: 141kAnalytics watercourse reducer · Muse Spark 1.3: 137kEntitlement overage lines · Kimi K3: 30kMulti-region expanse · Kimi K3: 9kTax jurisdiction · Kimi K3: 39kAPI token metering · Kimi K3: 69kAPI keys & environments · Kimi K3: 44kS3 datastore measurement · Kimi K3: 32kCustomer personality migration · Kimi K3: 66kBilling schedule migration · Kimi K3: 19kLinearizable scan · Kimi K3: 71kAnalytics watercourse reducer · Kimi K3: 55kEntitlement overage lines · GPT-5.6 Sol: 12kMulti-region expanse · GPT-5.6 Sol: 8kTax jurisdiction · GPT-5.6 Sol: 22kAPI token metering · GPT-5.6 Sol: 31kAPI keys & environments · GPT-5.6 Sol: 25kS3 datastore measurement · GPT-5.6 Sol: 25kCustomer personality migration · GPT-5.6 Sol: 24kBilling schedule migration · GPT-5.6 Sol: 13kLinearizable scan · GPT-5.6 Sol: 37kAnalytics watercourse reducer · GPT-5.6 Sol: 30k
Fable 5.1 · 64k overallGPT-6 Astra · 24k overallGemini 3.8 Flash · 94k overallGLM 5.3 · 117k overallGrok 4.6 · 67k overallMuse Spark 1.3 · 87k overallKimi K3 · 43k overallGPT-5.6 Sol · 23k overall
View task valuesTask- Fable 5.134k
- GPT-6 Astra13k
- Gemini 3.8 Flash78k
- GLM 5.368k
- Grok 4.67k
- Muse Spark 1.336k
- Kimi K330k
- GPT-5.6 Sol12k
04Evaluation Setup
Each supplier was tally successful an isolated sandbox. All tasks are successful Harbor format, and verifiers are injected astatine grading time. The verifiers are inspired by existing trial suites successful the codebase aliases usage those tests verbatim.
English (US) ·
Indonesian (ID) ·