Massive Microsoft 365 Outage Caused by Maintenance Bug, Leaving Thousands Disrupted
A devastating outage hit Microsoft’s popular productivity suite, Microsoft 365, leaving thousands of users unable to access their critical services. The disruption was caused by a maintenance bug that mistakenly removed IP routes from more devices than intended, disrupting network traffic and affecting Azure and Microsoft 365 services.
The outage began at 10:44 AM ET on Thursday, July 23, primarily affecting customers accessing Microsoft 365 through network infrastructure connected to the West US Azure region. As the incident unfolded, Downdetector recorded an alarming 2,403 outage reports, with SharePoint accounting for a staggering 78% of complaints. Other affected services included Excel, OneDrive, and Teams, with users experiencing intermittent access issues, errors, and degraded functionality.
Microsoft’s automated network maintenance request system was at fault, where a bug in the request conversion process incorrectly marked additional network devices as part of the maintenance event. This led to IP routes being removed from more devices than intended, disrupting network traffic entering or leaving the West US region. Fortunately, Microsoft quickly identified the cause and initiated a rollback of the maintenance change, which was completed at 2:26 PM ET.
The affected services included not only Microsoft 365 but also various Azure services, such as Power BI, Power Apps, Copilot Studio, Windows 365, and Microsoft Defender. Some Defender customers even experienced delays in receiving responses from Microsoft Defender Experts. The outage’s impact was far-reaching, with users experiencing connectivity failures, increased latency, and problems accessing numerous cloud services.
Microsoft’s engineers began investigating the issues immediately after the outage began and quickly identified the root cause. They then initiated a rollback of the maintenance change, which restored the affected network infrastructure and allowed Microsoft 365 services to recover. Some Azure services continued recovering after the fix was put in place, with Microsoft reporting that all affected services had fully recovered by 3:41 PM ET.
The incident has sparked concerns about the reliability of automated maintenance processes and highlights the importance of robust safety checks and fail-safes. As a result, it’s essential for businesses to review their business continuity and disaster recovery plans and take actions tailored to their specific environments.
In light of this outage, users should be aware that even seemingly minor maintenance tasks can have far-reaching consequences if not properly executed. To mitigate the risk of similar incidents in the future, organizations should prioritize regular security audits, invest in robust network monitoring tools, and stay vigilant about potential issues with automated processes. By doing so, businesses can ensure their critical services remain accessible and available when they need them most.
Source: Bleeping Computer — 2026-07-24