September 2026
Benchmarking frontier AI models connected private, real-world, endeavor codebases.
01Introduction
Today we are releasing Real-SWE, a benchmark that evaluates frontier AI models connected private, real-world, endeavor codebases. Each task comes from a backstage accumulation codebase that we licensed from a real-world company. These are problems their engineers activity on, pinch each the discourse and complexity that comes pinch an existing product.
- Private codebases. Agents must navigate proprietary systems whose codification and solutions aren’t disposable connected the nationalist internet.
- Work pinch business consequences. Getting billing right, calculating taxes, migrating customers. Changes that impact really a business runs, often crossed aggregate services.
- Company-specific complexity. Every institution has its ain rules and ways of penning code. Agents person to understand those conventions and make changes that activity pinch what’s already there.
Can a coding supplier really do the activity of a package technologist successful the existent world?
1
Fable 5.1
Claude Code
Resolution rate: 38.8%
2

GPT-6 Astra
Codex CLI
Resolution rate: 33.8%
3

Gemini 3.8 Flash
Gemini CLI
Resolution rate: 31.2%
4
GLM 5.3
Claude Code
Resolution rate: 28.8%
=5
Grok 4.6
Grok Build
Resolution rate: 23.8%
=5
Muse Spark 1.3
Muse Code
Resolution rate: 23.8%
7
Kimi K3
Kimi Code
Resolution rate: 18.8%
8

GPT-5.6 Sol
Codex CLI
Resolution rate: 16.2%
Expert-generated aliases synthetic tasks tin beryllium good designed, but they aren’t the verbatim, existent tasks that engineers successful existent companies request to do. Our tasks disagree connected 2 axes: the underlying coding artifact and specificity of the instruction. Both adhd complexities that situation today’s frontier models.
We usage autochthonal harnesses to bespeak really endeavor engineers activity successful practice, evaluating model-and-harness combinations alternatively than models successful isolation.
Real institution tasks require company-specific context
Correct billing depends connected business rules and outer services
Fix invoice billing truthful each business charges the correct taxation and exempt customers aren't taxed.
Billing reopens connected Monday and each invoice this work issues is coming retired untaxed. Each business connected the level settles its taxation a different way: immoderate support a complaint themselves, immoderate want each invoice priced against the buyer's destination by our taxation authority provider, and immoderate cod thing astatine all, while a customer we clasp an exemption for is charged thing whichever measurement its business is configured. Pricing a destination intends going to the authority pinch some addresses, the priced lines and the merchandise class that business sells under, connected the sandbox aliases the accumulation authority according to the relationship the business is on; an reside the authority refuses must beryllium reported without stopping the invoice. The rate, the taxation and the gross beryllium connected the issued invoice, and erstwhile an invoice is settled the waste is revenge backmost to the authority nether that invoice's number truthful the returns reconcile. Invoices betwixt European parties show some sides' VAT registrations. The authority and ledger are disposable astatine TAX_JAR_URL, PROD_TAX_JAR_URL and INFLUX_URL.
Services successful the sandbox
- TaxJar sandbox
- TaxJar production
- InfluxDB ledger
- NestJS service
- TypeScript
Agents activity crossed code, infrastructure, and business tools
Tools and services crossed Real-SWE task environments. Each task exposes only the services its workflow needs.
- AWS emulator
- Docker
- Kubernetes
- GitHub
- Linear MCP
- PostgreSQL
- MySQL
- MongoDB
- Gel
- Redis
- Go
- Python
- Node.js
- Vitest
- Slack
- Intercom
- Google Drive
- ClickUp
Codebase Selection
We selected codebases done a rigorous screening process, focusing connected existent companies pinch important usage, beardown engineering teams, and demanding accumulation workloads. The sample tasks analyzed beneath travel from these codebases, including:
- A Luma/Partiful competitor pinch 200K+ users and a top 100 App Store ranking
- A user fintech level processing 100K+ slope statements
- Enterprise AI income platforms supporting analyzable business workflows
We prioritize codification written to meet an existent personification aliases business request complete codification written solely to create a benchmark task. Production engineering requires knowing existing architecture, preserving behaviour that users trust on, and making changes wrong existent operational constraints.
Brief instructions tin require changes crossed galore files
Our tasks picture the alteration needed, leaving agents to observe implementation specifications successful the codebase and surrounding tools. Any behaviour required by the verifier must beryllium stated aliases reasonably discoverable. This leads to our prompts being somewhat underspecified, astir par pinch DeepSWE and Terminal Bench, but circumstantial capable to not omit instructions.
The activity is cross-functional and complex: a azygous alteration tin span aggregate parts of the application. Agents must understand existing business logic and institution coding patterns while keeping the surrounding strategy working.
Prompt magnitude · median
A emblematic Real-SWE instruction is 1,742 characters.
FrontierCode2,056 chars
DeepSWE1,975 chars
Terminal-Bench 31,584 chars
FrontierSWE v2992 chars
Real-SWE1,742 chars
Files edited by the reference solution · median
11 files successful Real-SWE, compared pinch 6 successful FrontierCode and DeepSWE.
FrontierCode6
DeepSWE6
Real-SWE11
Models neglect moreover successful short rollouts.
71.4% of rollouts nether 10 minutes failed, compared pinch 73.4% of longer rollouts.
Triaging aggregate systems and knowing requirements successful codebases riddled pinch existing business logic and coding patterns is difficult.
70/98 failed
398/542 failed
- Failed
- Passed
Every task is inspired aliases lifted verbatim from a private, real-world codebase. We find these types of tasks ace absorbing for 3 reasons:
- Tasks connected backstage codebases are natively retired of distribution. These types of coding tasks are not disposable anyplace connected the net and are improbable to person ever been trained connected by immoderate different ai model. 99% of tokens successful real-world enterprises are hidden distant from the frontier models.
- These tasks are economically viable work. Each task present has a nonstop narration to walk and was assigned to an technologist earning a salary. Most benchmarks trial interesting, experimental capabilities that are often improbable to beryllium wide successful the real-world.
- Company-specific engineering patterns matter. Does AI codification lucifer the barroom of a real-world enterprise? Our results show america that we're acold from that reality. Many enterprises attraction astir codification standards and patterns. We've recovered that today's models are weaker astatine knowing institution coding patterns and often miss requirements aliases don't verify their assumptions.
02Analysis
Here's an study of a mini sample of tasks from our benchmark. If you're willing successful the sample, request entree here.
6 of 10 tasks person solution rates beneath 15%
Each task had 8 rollouts per model.Missed requirements are the astir communal failure
Failures are grouped by observed submission behaviour utilizing the aforesaid taxonomy crossed models, pursuing DeepSWE.
Fable 5.1
GPT-6 Astra
Gemini 3.8 Flash
GLM 5.3
Grok 4.6
Muse Spark 1.3
Kimi K3
GPT-5.6 Sol
Unverified assumptionMissed requirementIntegration errorRegressionWrong file
No exemplary solves each task
One quadrate per rollout: each statement is simply a task, each file a trial, 8 tests per task for each model.
Fable 5.1
01
02
03
04
05
06
07
08
09
10
GPT-6 Astra
01
02
03
04
05
06
07
08
09
10
Gemini 3.8 Flash
01
02
03
04
05
06
07
08
09
10
GLM 5.3
01
02
03
04
05
06
07
08
09
10
Grok 4.6
01
02
03
04
05
06
07
08
09
10
Muse Spark 1.3
01
02
03
04
05
06
07
08
09
10
Kimi K3
01
02
03
04
05
06
07
08
09
10
GPT-5.6 Sol
01
02
03
04
05
06
07
08
09
10
PassUnverified assumptionMissed requirementIntegration errorRegressionWrong file
Different models neglect successful different ways
Percentages are retired of each model's grounded runs, not each runs.
03Effort & the Frontier
Higher costs does not guarantee a higher solution rate
Estimated frontier
Estimated rollout costs scope from $2.50 to $6.96
| 1 | Gemini 3.8 Flash$2.50 |
| 2 | GPT-5.6 Sol$2.65 |
| 3 | Muse Spark 1.3$2.74 |
| 4 | Grok 4.6$3.44 |
| 5 | Kimi K3$3.90 |
| 6 | GPT-6 Astra$4.67 |
| 7 | GLM 5.3$5.12 |
| 8 | Fable 5.1$6.96 |
mean per rollout, by task
Swipe the floor plan to spot each tasks.
Fable 5.1 · 64k overallGPT-6 Astra · 24k overallGemini 3.8 Flash · 94k overallGLM 5.3 · 117k overallGrok 4.6 · 67k overallMuse Spark 1.3 · 87k overallKimi K3 · 43k overallGPT-5.6 Sol · 23k overall
View task valuesTask- Fable 5.134k
- GPT-6 Astra13k
- Gemini 3.8 Flash78k
- GLM 5.368k
- Grok 4.67k
- Muse Spark 1.336k
- Kimi K330k
- GPT-5.6 Sol12k
04Evaluation Setup
Each supplier was tally successful an isolated sandbox. All tasks are successful Harbor format, and verifiers are injected astatine grading time. The verifiers are inspired by existing trial suites successful the codebase aliases usage those tests verbatim.