Massive Microsoft 365 Outage Caused by Maintenance Bug, Leaving Users Scrambling
In a dramatic display of the intricate dance between human error and automated systems, Microsoft’s massive outage of its popular Microsoft 365 suite was caused by a simple maintenance bug. The glitch mistakenly removed IP routes from more devices than intended, disrupting Azure and Microsoft 365 services for several hours on Thursday.
The outage primarily affected customers accessing Microsoft 365 through network infrastructure connected to the company’s West US Azure region. This included users of Microsoft OneDrive, SharePoint Online, Teams, and other popular services, who reported intermittent access, slow loading times, or complete failures. The extent of the disruption was staggering: by 11:11 AM ET on Thursday, Downdetector had recorded over 2,400 outage reports – a sharp spike from its normal baseline.
Microsoft’s engineers quickly sprang into action to mitigate the damage, attempting to reroute traffic through alternate network paths. However, many services continued to be affected, and some users experienced delays receiving responses from Microsoft Defender Experts or failing investigations and workflows triggered through Threat Explorer and Advanced Hunting.
After a thorough investigation, Microsoft revealed that the root cause of the outage was a bug in its automated network maintenance request system. During routine device maintenance in the West US Azure region, the system incorrectly marked additional network devices as part of the maintenance event. This led to IP routes being removed from more devices than intended between Microsoft’s West US datacenter and its wide-area network.
As a result, network traffic entering or leaving the West US region was disrupted, leading to connectivity failures, increased latency, and problems accessing numerous cloud services. The Azure incident caused widespread disruptions, affecting not only Microsoft 365 but also other critical services such as Azure App Service, Application Gateway, and Power BI Embedded.
Fortunately, Microsoft’s engineers were able to identify the issue quickly and initiated a rollback of the maintenance change at 1:45 PM ET. This was completed at 2:26 PM ET, restoring the affected network infrastructure and allowing Microsoft 365 services to recover. Some Azure services continued recovering after the fix was put in place, with Microsoft reporting that all affected services had fully recovered by 3:41 PM ET.
This incident serves as a stark reminder of the importance of robust safety checks and automated processes in large-scale IT systems. While human error is often seen as an inevitability, it’s essential for companies like Microsoft to learn from these incidents and implement measures to prevent similar outages in the future.
In practical terms, this outage highlights the need for businesses to have contingency plans in place for potential disruptions. Users of Microsoft 365 should review their business continuity and disaster recovery plans and take steps to mitigate any future impacts on their operations.
Source: Bleeping Computer — 2026-07-24