[Submitted connected 30 Jul 2026]
View PDF HTML (experimental)
Abstract:Large connection models tin write, patch, and hunt code, but oncall guidelines origin study (RCA) demands thing different: reasoning complete noisy metrics, logs, traces, and root code, starting from ambiguous user-facing reports, often hours aft the incident began. We present ORCA-bench, a benchmark that puts general-purpose coding agents successful a production-fidelity oncall setting. ORCA-bench pairs a unrecorded OpenTelemetry-instrumented microservice system--exposing six days of metrics, logs, and traces done existent telemetry interfaces (Prometheus, Jaeger, and OpenSearch via Grafana) and afloat source-code access--with 1,079 RCA tasks that systematically alteration study specificity, time-to-detection, and co-occurring responsibility scenarios. Ground-truth symptoms are curated and signed disconnected by master SREs, and our LLM-as-judge is independently re-scored by humans (Cohen's $\kappa_w=0.90$). Across 5 frontier agents, the champion RCA Accuracy is 25.3% connected Medium-difficulty tasks (the realistic-input setting) and 10.0% connected Hard--a spread that remains moreover pinch Claude Fable 5. The weakest exemplary hallucinates an implausible guidelines origin successful 40% of incident reports, and removing source-code entree degrades each metric. Crucially, these are performances connected a curated 50 GB / six-day testbed pinch tasks investigated successful isolation connected a strategy whose codification and instrumentation are public. Since existent accumulation systems are bid of magnitudes larger, much dynamic, and much idiosyncratic, the spread we study is simply a little bound connected the engineering finance required earlier frontier coding agents tin beryllium safely entrusted pinch accumulation reliability. We merchandise the nationalist group astatine this https URL.Submission history
From: Albert Gong [view email]
[v1] Thu, 30 Jul 2026 17:14:07 UTC (220 KB)
English (US) ·
Indonesian (ID) ·