The August 17 outage, and the work ahead

Aug 21, 2026 02:22 AM - 1 hour ago 1

On August 17, GitHub knowledgeable an outage that lasted 7 hours and 47 minutes. It disrupted github.com, authentication, GitHub Actions, APIs, propulsion requests, issues, and Copilot, affecting developers and organizations astir the world. If you were trying to vessel package that day, we fto you down.

This was our 2nd important incident successful August, pursuing an actions nonaccomplishment connected August 6. In March and April, I shared the activity underway to amended GitHub’s reliability. We person made progress, but these incidents make clear that we must accelerate this work.

What happened

Our investigation recovered that the outage began erstwhile postulation reached a caller peak, and a captious infrastructure constituent successful our Central US information halfway grounded to standard pinch it. The resulting capacity unit dispersed done our systems, causing authentication failures and disrupting aggregate GitHub services.

Recovery required respective coordinated actions. Teams rerouted traffic, isolated affected infrastructure, and restored services successful stages. Most GitHub services recovered earlier that day, but immoderate Copilot services took longer. Errors successful those services triggered a client-side retry loop that accrued postulation during recovery. We had to mitigate that behaviour earlier we could safely reconstruct traffic. The afloat root origin analysis includes a elaborate method timeline.

Neither outage was caused by a codification aliases configuration change. Both incidents were capacity failures astatine their core. We grounded to standard captious components earlier request exceeded their capacity. Since April, monthly commits person grown from 1.4 cardinal to 2.9 billion. That maturation explains the unit connected our systems, but it does not excuse these outages.

 merged propulsion requests per period rising to astir 130M, commits per period rising to astir 2.9B, and caller repositories per period rising to astir 24M, pinch acceleration successful 2025–2026.

What we person done and what comes next

As portion of the reliability commitments we made earlier this year, we person focused connected 3 priorities: adding capacity, improving efficiency, and removing architectural bottlenecks. We person since added much than 3 cardinal CPU cores, 120 petabytes of high-speed storage, and important web capacity. We installed arsenic overmuch hardware arsenic disposable powerfulness allowed successful our existing information centers while accelerating our migration to Azure.

Today, Azure serves astir 58% of GitHub’s level load and half of each Git operations, up from 12% of level load successful May. This expanded footprint has besides supported the maturation successful GitHub Actions occupation runs shown below.

Large dark-themed statement floor plan titled ‘Growth successful completed GitHub Actions runs’ shows a rising inclination from early 2026 to August, pinch regular play dips and expanding peaks. Values turn from astir 15–30M early successful the twelvemonth to complete 100M, ending adjacent 115.4M.

Azure’s infrastructure and capacity person besides accelerated our activity to standard the largest monorepos. Our adjacent milestone is an architecture that scales publication capacity linearly pinch the number of readers, enabling unlimited publication operations. We will rotation it retired gradually, opening pinch the largest monorepos.

Two dark-themed ‘Fetch Throughput History’ charts comparison fetch operations per 2nd complete short clip windows. Left floor plan fluctuates and plateaus astir ~1,000 OPS/S earlier dropping adjacent the end; correct floor plan climbs steadily successful steps to astir ~1,800 OPS/S.

Scale is not our only challenge. As the gait and complexity of alteration increased, our existing operational practices did not support up. We person redirected teams and resources toward readiness and invested successful stronger testing, safer rollouts, amended observability, and much effective alerting. We person made progress, but this activity is not complete.

In addition, we are besides isolating captious systems and removing shared limitations betwixt them. This activity is designed to trim the likelihood of an outage and limit its effect erstwhile 1 occurs.

We study from each outage and adhd caller activity to our readiness workstream. The August 6 and August 17 incidents led to 2 contiguous changes. First, we are applying accordant retry limits, retry budgets, and adaptable timeouts crossed service-to-service interactions to forestall retry storms and cascading load. Second, we are reviewing lower-priority CPU and representation alerts to place components that could neglect during abrupt postulation spikes.

Our committedness to precocious readiness isn’t conscionable a method promise. The developer organization depends connected GitHub to build, ship, and run their work. That is only imaginable if you tin trust connected us, and connected August 17, you couldn’t. It is our work to hole that. We’ll gain your spot done the scaling and reliability of the platform.

Written by

Vlad Fedorov

Vladimir Fedorov is GitHub's Chief Technology Officer, bringing decades of acquisition successful engineering activity and innovation. A passionate advocator for developer productivity, Vlad is starring GitHub’s engineering squad to style the early of developer devices and invention pinch a developer-first mindset.

Before joining GitHub, Vlad co-founded UserClouds, a startup specializing successful information governance and privacy. He spent 12 years astatine Facebook, now Meta, arsenic Senior Vice President, starring engineering teams of complete 2,000 crossed Privacy, Ads, and Platform. Earlier successful his career, Vlad worked astatine Microsoft and earned some his BS and MS successful Computer Science from Caltech. He presently serves connected the committee of Codepath.org, an statement dedicated to reprogramming higher acquisition to create the first AI-native procreation of engineers, CTOs, and founders.

Vlad lives successful the Bay Area and erstwhile not moving enjoys spending clip extracurricular and connected the h2o pinch his family.

More