Microsoft blames automated maintenance bug for major Microsoft 365 outage

Microsoft attributed a wide-reaching Microsoft 365 outage on July 23 to a bug in its automated maintenance request conversion system that removed IP routes from more devices than intended, disrupting traffic to and from the West US Azure region. The incident affected multiple Microsoft 365 applications and several Azure services, and forced Microsoft to roll back the maintenance change to restore connectivity.

Microsoft confirms maintenance-related routing removals triggered outage

Microsoft tracked the incident as Microsoft 365 outage MO1437424. The outage began at 10:44 AM ET on July 23 and primarily affected customers whose network infrastructure connected to Microsoft’s West US Azure region. Microsoft reported that engineers identified large-scale route churn in its wide-area network and traced the problem to a datacenter in the West US region; the company began a rollback at 1:45 PM ET and completed it at 2:26 PM ET.

Service telemetry and customer reports showed recovery following the rollback, and Microsoft said all affected services had fully recovered by 3:41 PM ET. At 11:11 AM ET, Downdetector recorded 2,403 outage reports versus a normal baseline of 29, with SharePoint comprising 78% of complaints, Excel 11% and the Microsoft 365 Admin Center 6%.

How the maintenance automation removed routes

Microsoft’s preliminary Post Incident Review states the disruption began during routine device maintenance in the West US region. The company’s maintenance process converts operator requests into system-readable instructions and is intended to verify that at least one of two redundant paths remains healthy before maintenance proceeds.

A bug in that conversion system incorrectly marked additional network devices as part of the maintenance event. As a result, IP routes were removed from more devices than intended between the affected datacenter and Microsoft’s wide-area network, disrupting traffic entering or leaving the West US region. Microsoft emphasized that traffic remaining entirely within the West US region was not affected.

Microsoft initially tried rerouting traffic through alternate network paths as a mitigation step; that action provided relief for some customers but did not fully restore service until the maintenance change was reverted. The company said engineers began investigating the issue immediately after detection and linked the route removals to the maintenance activity before initiating the rollback.

Why it matters

The outage interrupted access to core productivity and cloud services used by enterprises and administrators. Microsoft listed multiple affected Microsoft 365 services including OneDrive (intermittent access), SharePoint Online (“Something went wrong” errors), Teams (degraded chat and images), the Microsoft 365 Admin Center (slow or unavailable), Power Automate, Copilot Chat, and Microsoft Loop.

Beyond Microsoft 365, the incident also disrupted a wide set of Azure capabilities: App Service, Application Gateway, Azure AD B2C, AI Search, API Management, Cosmos DB, Databricks, Firewall, Kubernetes Service, Monitor, Virtual Desktop, ExpressRoute, Log Analytics, Microsoft Graph, Sentinel, Power BI Embedded, Virtual WAN, and VPN Gateway, among others. Some Microsoft Defender functionality—investigations, workflows and remediation actions—experienced delays or failures.

For enterprises this meant temporary loss or degradation of collaboration, automation, analytics and security tooling. Microsoft advised customers to review business continuity and disaster recovery plans and take any actions appropriate for their environments while the incident was active.

Context and implications for cloud operations

The incident highlights how automated maintenance tooling and safety checks sit at the intersection of operational efficiency and systemic risk. In Microsoft’s account, an automation defect expanded the maintenance scope beyond the intended devices, which in turn removed routing information and produced WAN-level connectivity failures and latency.

Microsoft said it will perform a full internal review focused on safety checks and the automated maintenance request change process, and intends to publish a final Post Incident Review after completing the investigation—typically within 14 days. The company’s initial mitigation (traffic rerouting) provided partial relief but did not replace the rollback of the faulty change.

For network and cloud operators, the episode underscores the importance of conservative automation safeguards, pre-change validation of routing impacts, and robust rollback procedures. Microsoft’s warning to customers to verify their business continuity plans signals that enterprises should assume they may need to enact local mitigations when provider-side networking incidents occur.

What to watch next

  • Microsoft’s final Post Incident Review, expected within approximately 14 days, for specific root-cause detail and the concrete safety-check changes it will implement.

  • Any follow-up notes on limits or changes to automated maintenance tooling, including new pre-deployment validations, expanded telemetry, or operator approvals that Microsoft will add to prevent recurrence.

  • Customer guidance or recommended configuration changes for enterprises that rely on West US region connectivity, and whether Microsoft recommends specific network or routing adjustments to improve resilience.

Microsoft’s preliminary report gives a clear operational timeline and a mechanistic cause: a conversion bug that widened a maintenance window and removed IP routes. The forthcoming Post Incident Review is the primary outstanding source of detail about the specific checks and code changes Microsoft will use to reduce the risk of a similar outage.

Source: BleepingComputer