The August 17 outage, and the work ahead
GitHub experienced a 7-hour and 47-minute outage on August 17, affecting developers and organizations worldwide. The incident was caused by a critical infrastructure component failure in the Central US data center, leading to authentication failures and disrupting multipl…
Intelligence analysis by Llama

GitHub's second significant incident in August highlights the need to accelerate work on improving the platform's reliability. The company has made progress but acknowledges that more needs to be done to prevent such outages in the future.
Imagine you're trying to build a big Lego castle, but the Lego pieces are all stuck together and can't be added quickly. That's what happened to GitHub on August 17. The company's systems got too busy and couldn't handle the traffic, causing a big outage. Now, GitHub is working hard to fix this problem and make sure it doesn't happen again.
Analysis
What happened
The August 17 outage was caused by a critical infrastructure component failure in GitHub's Central US data center. The resulting capacity pressure spread through the company's systems, causing authentication failures and disrupting multiple GitHub services. The incident was not caused by a code or configuration change, but rather a capacity failure that was exacerbated by the rapid growth in monthly commits.
What we have done and what comes next
As part of its reliability commitments, GitHub has focused on three priorities: adding capacity, improving efficiency, and removing architectural bottlenecks. The company has added more than 3 million CPU cores, 120 petabytes of high-speed storage, and significant network capacity. It has also accelerated its migration to Azure, which now serves roughly 58% of GitHub's platform load and half of all Git operations.
However, the company acknowledges that it has not yet achieved its goal of scaling read capacity linearly with the number of readers, enabling unlimited read operations. It plans to roll out this architecture gradually, beginning with the largest monorepos.
The work ahead
GitHub's commitment to high availability is not just a technical promise, but a responsibility to its developer community. The company recognizes that it must earn the trust of its users through the scaling and reliability of the platform. To achieve this, it will continue to prioritize its availability workstream, investing in stronger testing, safer rollouts, better observability, and more effective alerting.
The company has already made progress in this area, but acknowledges that more needs to be done to prevent such outages in the future. It will continue to learn from every outage and add new work to its availability workstream, with the goal of achieving a more reliable and scalable platform.
Key points
- GitHub experienced a 7-hour and 47-minute outage on August 17, affecting developers and organizations worldwide.
- The outage was caused by a critical infrastructure component failure in the Central US data center.
- GitHub has made progress in improving its reliability but acknowledges that more needs to be done to prevent such outages in the future.
- The company is prioritizing its availability workstream, investing in stronger testing, safer rollouts, better observability, and more effective alerting.
- GitHub plans to roll out a new architecture that scales read capacity linearly with the number of readers, enabling unlimited read operations.
GitHub's efforts to improve its reliability and scalability will likely lead to a more stable and efficient platform, reducing the likelihood of future outages and improving the overall user experience.
If GitHub fails to address its capacity and scalability issues, it may continue to experience outages and disruptions, potentially leading to a loss of user trust and revenue.