GitHub Outage August 2026 Explained: The Autoscaling Failure and Retry Storm Behind 7 Hours of Downtime

GitHub Outage Explained: What Happened and Why Developers Couldn’t Access Their Code GitHub Outage Explained: What Happened and Why Developers Couldn’t Access Their Code

GitHub Outage August 2026: What Happened and What Developers Need to Know

GitHub’s major outage on August 17, 2026, was more than a temporary website failure. The incident exposed how a relatively contained infrastructure capacity problem can turn into a much larger service disruption when autoscaling, load balancing and retry behavior interact in unexpected ways.

The incident lasted 7 hours and 47 minutes, beginning at 13:28 UTC and ending at 21:15 UTC. During the disruption, developers experienced elevated errors and degraded performance across GitHub’s web experience, APIs, Issues, Pull Requests, GitHub Actions, Copilot and other services.

GitHub’s status updates initially showed approximately 20% error rates for web and API traffic, while archive downloads and raw repository content downloads reached approximately 50% error rates. Authentication services including SAML, OIDC, SCIM and Team Sync were also affected.

What was initially an unresolved infrastructure incident can now be understood in considerably more detail. GitHub’s incident analysis points to a chain involving an Istio sidecar reaching its concurrency limit, an autoscaling policy that failed to account for that constraint, saturated load balancers and retry behavior that amplified the disruption.

What Happened During the GitHub Outage?

The incident began in GitHub’s Central US infrastructure during a new peak in traffic. An Istio sidecar reached its configured concurrency limit, creating a capacity constraint inside the service infrastructure.

Normally, an autoscaling system should respond when additional capacity is needed. However, GitHub’s scaling policy was monitoring the capacity of the host service rather than the relevant concurrency constraint of the sidecar.

As a result, the system did not scale in the way engineers expected when the sidecar reached its limit.

This distinction is important because the outage was not simply a case of GitHub receiving “too much traffic.” The deeper problem involved how infrastructure capacity was being measured and how the autoscaling system responded to that measurement.

How the Autoscaling Failure Led to Load-Balancer Saturation

The capacity problem eventually contributed to network saturation on load balancers in GitHub’s Central US environment.

Several HAProxy nodes reached their flow capacity, affecting GitHub’s gateway authentication path. Once this infrastructure became constrained, requests that depended on the affected services began failing or taking longer to complete.

This helped turn a localized capacity problem into a broader platform incident.

The sequence can be simplified as:

  • Traffic reached a new peak.
  • An Istio sidecar reached its concurrency limit.
  • The autoscaling policy did not properly account for that sidecar constraint.
  • Infrastructure capacity became saturated.
  • Load balancers reached their limits.
  • Authentication and other GitHub services began experiencing failures.
  • Retries generated additional traffic and made recovery more difficult.

The Retry Storm Made the Outage Worse

One of the most important lessons from the incident is that the initial infrastructure failure was only part of the problem.

As requests began failing or taking longer to complete, retry mechanisms generated additional requests. Instead of allowing the infrastructure to recover cleanly, some of that retry traffic added more pressure to already-constrained services.

This created a feedback loop:

Failure → retries → additional traffic → more saturation → more failures → more retries.

This type of behavior is commonly known as a retry storm. It is particularly dangerous in distributed systems because a relatively small initial failure can be amplified across multiple dependent services.

The GitHub incident is therefore a useful real-world example of why retry logic needs carefully designed limits, backoff strategies and failure handling.

Why GitHub Copilot Was Still Having Problems After Other Services Recovered

GitHub’s recovery did not happen all at once.

Many core services began recovering earlier in the incident, but Copilot continued to experience authentication-related problems. A separate retry-related behavior in Visual Studio Code became an important factor during this stage of recovery.

According to reporting on GitHub’s incident analysis, a latent retry bug in Visual Studio Code caused a major increase in requests to the Copilot Token Service when responses from an internal endpoint were delayed.

The Copilot Token Service normally handled approximately 7,000–9,000 requests per second. During the incident, traffic reportedly surged to roughly 70,000–100,000 requests per second.

That represents an enormous amplification of demand.

It is important, however, not to misinterpret this finding.

GitHub did not identify AI or Copilot usage as the original cause of the outage. Instead, Copilot-related traffic became a significant factor during recovery because of retry behavior triggered by the earlier service disruption.

Did AI Cause the GitHub Outage?

No. The available incident analysis does not support the claim that AI coding agents or increasing AI usage caused the initial GitHub outage.

There had been considerable speculation among developers that the rapid growth of AI-assisted coding might be putting unusual pressure on GitHub infrastructure. The new information provides a much more specific explanation.

The initial failure involved infrastructure capacity, an Istio sidecar concurrency limit and an autoscaling policy that did not properly account for that limit.

AI-related services did become involved later, particularly through the Copilot Token Service, but that was part of the recovery and amplification phase rather than the original trigger.

What About the Reported Scraping Activity?

GitHub also identified scraping activity against codeload endpoints as something that complicated recovery.

That detail should not be interpreted as evidence that the outage was caused by a cyberattack or DDoS attack.

The distinction matters:

  • Initial infrastructure failure: associated with the sidecar concurrency and autoscaling problem.
  • Retry amplification: increased pressure after services began failing.
  • Copilot recovery issue: amplified by retry behavior in Visual Studio Code.
  • Scraping activity: complicated recovery.

There is therefore no basis for describing the entire August 2026 GitHub outage as a cyberattack.

How Long Did the GitHub Outage Last?

GitHub recorded the incident from 13:28 UTC until 21:15 UTC on August 17, 2026, making the total incident duration 7 hours and 47 minutes.

The recovery happened in stages rather than through one single switch being flipped.

StageWhat Happened
13:28 UTCIncident began.
13:45 UTCGitHub reported approximately 20% error rates across numerous experiences.
14:04 UTCWeb and API errors remained around 20%, while archive and raw-content downloads were around 50% errors.
14:31 UTCGitHub Copilot was reported as degraded.
16:36 UTCGitHub reported strong signs of recovery after corrective actions.
18:03 UTCGitHub Actions had recovered.
21:02 UTCCopilot Token Service recovery was nearing completion.
21:15 UTCThe overall incident was resolved.

What Services Were Affected?

The outage spread across a wide range of GitHub functionality.

  • GitHub.com web experiences
  • API Requests
  • Pull Requests
  • Issues
  • GitHub Actions
  • Webhooks
  • Git Operations
  • GitHub Copilot
  • Archive downloads
  • Raw repository content downloads
  • SAML authentication
  • OIDC authentication
  • SCIM
  • Team Sync

For development teams, the impact was particularly significant because GitHub is no longer simply a place to store Git repositories. Modern software workflows often depend on GitHub for code review, authentication, CI/CD, automation, package distribution and AI-assisted development.

Why the Incident Was More Serious for Modern Development Teams

The outage demonstrates how deeply integrated GitHub has become in modern software delivery.

A developer might use Git locally, but their complete workflow can still depend on GitHub services for pull requests, automated tests, deployments, authentication and AI coding assistance.

That creates a larger dependency surface.

When several interconnected services experience problems simultaneously, teams can lose the ability to perform routine development and deployment tasks even if their source code remains available locally.

GitHub’s Outage Also Highlights the Importance of Retry Design

The most useful engineering lesson from the incident may not be about GitHub specifically. It is about distributed systems.

Retries are necessary in unreliable networks. A temporary failure should not always become a permanent application error. But poorly designed retries can make an outage substantially worse.

A robust retry strategy normally needs mechanisms such as:

  • Exponential backoff
  • Jitter
  • Maximum retry counts
  • Circuit breakers
  • Request deadlines
  • Rate limiting
  • Load shedding
  • Clear failure states

Without those safeguards, millions of clients can effectively become a traffic generator when a service begins responding slowly.

Why Autoscaling Metrics Matter

The GitHub incident also highlights a subtle but important cloud-infrastructure problem: scaling based on the wrong metric.

A system can appear to have sufficient capacity at the host or service level while an individual component inside that system is already saturated.

That is effectively what made the sidecar problem difficult to handle automatically.

For engineers designing cloud-native systems, the lesson is straightforward: autoscaling should monitor the actual bottlenecks that limit application throughput, not simply the most convenient infrastructure metric.

GitHub Outage vs. Git Itself: An Important Distinction

It is also worth separating GitHub from Git.

Git is a distributed version-control system. A developer with a local repository can still create commits, inspect history and work on branches even when GitHub is unavailable.

GitHub provides the hosted collaboration and development platform around Git.

That distinction is why an outage can disrupt pull requests, Actions, authentication and repository access while developers may still be able to continue working with their local Git repositories.

What Developers Can Learn From the GitHub Outage

The incident offers several practical lessons for engineering teams.

1. Don’t Depend Entirely on One Hosted Platform

Teams should understand what happens if their primary Git hosting or CI/CD provider becomes unavailable.

Maintaining local repositories, documented recovery procedures and appropriate backups can reduce the impact of a major service outage.

2. Design Retries Carefully

Every production service that retries requests should have limits and backoff mechanisms. A retry policy that looks harmless under normal conditions can become dangerous during a widespread failure.

3. Monitor the Real Bottlenecks

Infrastructure monitoring should measure the components that actually determine system capacity. Aggregate host metrics may not reveal that a specific sidecar, proxy or connection pool is already saturated.

4. Plan for Dependency Failures

If deployments, authentication, package retrieval or AI tools depend on an external service, teams should know which parts of their workflow can continue when that dependency becomes unavailable.

5. Don’t Confuse a Symptom With the Root Cause

The GitHub incident is a good example of why outage analysis requires looking at the entire chain. A retry storm, Copilot traffic spike or scraping activity might be highly visible during an incident without being responsible for the initial failure.

What GitHub Is Expected to Improve

The incident analysis points toward several areas where GitHub can strengthen its infrastructure and reliability practices, including autoscaling policies, retry behavior, concurrency management, load-balancer capacity and regional recovery mechanisms.

The larger goal is not simply to prevent the exact same failure from happening again. Mature incident response should identify the underlying class of failure and reduce the probability that a similar chain reaction can happen through another component.

The Bigger Lesson From the August 2026 GitHub Outage

The GitHub outage shows how modern cloud platforms can fail in ways that are more complicated than a single server going offline.

The initial problem was relatively specific: an infrastructure component reached a concurrency limit while the autoscaling system failed to respond appropriately. But the consequences spread through load balancers, authentication services, APIs, developer tools and retry mechanisms.

The later Copilot traffic surge demonstrates another important point: recovery itself can create additional load. When thousands or millions of clients attempt to reconnect or retry simultaneously, restoring a service can become harder than fixing the original failure.

For developers and infrastructure engineers, that may be the most valuable takeaway from GitHub’s August 2026 outage.

Reliable systems are not just systems that work when everything is healthy. They are systems designed to fail gradually, recover safely and prevent small problems from becoming cascading failures.

Frequently Asked Questions

What caused the GitHub outage in August 2026?

The incident was traced to an infrastructure capacity problem involving an Istio sidecar reaching its concurrency limit and an autoscaling policy that did not properly account for that constraint. The resulting saturation was amplified by retry behavior.

How long was GitHub down on August 17, 2026?

The incident lasted approximately 7 hours and 47 minutes, from 13:28 UTC until 21:15 UTC.

Did AI cause the GitHub outage?

No. The available incident analysis does not identify AI usage as the original cause. Copilot-related traffic became an important factor during recovery because retry behavior amplified requests to the Copilot Token Service.

Did a cyberattack cause the GitHub outage?

There is no evidence in the incident analysis that the outage itself was caused by a cyberattack or DDoS attack. GitHub did report scraping activity that complicated recovery, but that is different from identifying it as the original cause.

Why was GitHub Copilot affected for so long?

Copilot continued experiencing authentication-related problems after other services began recovering. A latent retry behavior in Visual Studio Code contributed to a major increase in requests to the Copilot Token Service during the recovery period.

Can developers still use Git if GitHub goes down?

Yes. Git is distributed, so developers can continue working with local repositories, committing changes and creating branches. However, hosted functionality such as Pull Requests, GitHub Actions, authentication and collaboration may remain unavailable.

Final Thoughts

The August 2026 GitHub outage is a useful reminder that reliability problems in modern platforms rarely have a single simple cause.

What started with a capacity constraint and an autoscaling blind spot developed into a cascading failure involving load balancers, authentication, retries and eventually Copilot traffic.

For developers, the lesson is not simply that GitHub can go down. It is that every dependency in a modern software delivery pipeline can become part of an outage chain.

Understanding those dependencies, designing safer retry behavior and maintaining practical fallback procedures can make the difference between a temporary inconvenience and a major development shutdown.

Leave a Reply

Your email address will not be published. Required fields are marked *