Microsoft says a bug in its automated network maintenance request system caused Thursday’s massive outage by mistakenly removing IP routes from more devices than intended, disrupting Azure and Microsoft 365 services.
The outage began at 10:44 AM ET on Thursday, July 23, and mostly affected customers accessing Microsoft 365 services through network infrastructure connected to Microsoft’s West US Azure region.
At 11:11 AM ET, Downdetector had recorded 2,403 outage reports, sharply above its normal baseline of 29. SharePoint accounted for 78% of the complaints, followed by Excel at 11% and the Microsoft 365 Admin Center at 6%.
Microsoft tracked the Microsoft 365 outage under incident ID MO1437424 and confirmed that multiple Microsoft 365 services were impacted:
Other affected services included Fabric and Power BI, Power Apps, Copilot Studio, Windows 365, and Microsoft Defender.
Some Defender customers experienced delays receiving responses from Microsoft Defender Experts, while investigations, workflows, and remediation actions triggered through Threat Explorer and Advanced Hunting could fail.
Microsoft initially attempted to mitigate the outage by rerouting traffic through alternate network paths, which helped customers, but many services continued to be affected.
Before determining what caused the outage, Microsoft warned customers that they might need to review their business continuity and disaster recovery plans and take actions appropriate for their environments.
The company later identified a recent networking change as the cause and began reverting it.
Microsoft completed the reversion at 2:26 PM ET and confirmed through service telemetry and customer reports that the Microsoft 365 incident had been resolved.
In a preliminary Post Incident Review for the Azure incident, Microsoft said the outage was triggered during routine device maintenance in its West US Azure region, where specific network paths were being isolated.
Microsoft says its maintenance process converts these types of requests into system-readable instructions and checks that at least one of two redundant paths remains healthy before the work begins.
However, a bug in the request conversion system incorrectly marked additional network devices as part of the maintenance event.
As a result, IP routes were removed from more devices than intended between Microsoft’s West US datacenter and its wide-area network.
The removed routes disrupted network traffic entering or leaving the West US region. However, Microsoft said traffic remaining entirely within the region was not affected.
The Azure incident caused connectivity failures, increased latency, and problems accessing numerous cloud services, including Azure App Service, Application Gateway, Azure AD B2C, Azure AI Search, Azure API Management, Azure Cosmos DB, Azure Databricks, Azure Firewall, Azure Kubernetes Service, Azure Monitor, Azure Virtual Desktop, ExpressRoute, Log Analytics, Microsoft Graph, Microsoft Sentinel, Power BI Embedded, Virtual WAN, and VPN Gateway.
Microsoft said its engineers began investigating the issues immediately after the outage began at 10:44 AM ET.
The problem initially presented itself as large-scale route churn in Microsoft’s WAN. Engineers later traced the route removals to a datacenter in the West US region and linked them with the recent maintenance activity.
Microsoft initiated a rollback of the maintenance change at 1:45 PM ET, which was completed at 2:26 PM ET.
The rollback restored the affected network infrastructure and allowed Microsoft 365 services to recover. Some Azure services continued recovering after the fix was put in place, with Microsoft reporting that all affected services had fully recovered by 3:41 PM ET.
Microsoft is now conducting a full internal review focused on the safety checks and automated processes used to execute maintenance requests.
“We will be preforming a full analysis focusing on safety checks, automated maintenance request change process, and more as we progress through our post mitigation internal retrospective,” explained Microsoft.
The company said it will publish a final Post Incident Review after completing its investigation, which is usually within 14 days.
Security teams log 54% of successful attacks and alert on just 14%. The rest move through your environment unseen.
The Picus whitepaper shows how breach and attack simulation tests your SIEM and EDR rules so threats stop slipping by detection.
Microsoft 365 outage affects Teams, SharePoint and other services
Microsoft fixes outage affecting MFA setup, MySignIn service
Microsoft warns of surge in ACR Stealer attacks on customers
OpenAI confirms ChatGPT is down worldwide
Hackers hijack hotel Wi-Fi DNS to steal Microsoft 365 accounts
Large tech companies, if anyone has noticed, have been experiencing a lot of trouble. I attribute this to mass outsourcing to multiple countries for 2 decades taking it’s toll. In America the H1B hires makes this worse. And those same companies are laying off American workers claiming work reduction due to AI when they really are just rehiring in Indian and the Philippines, because “AI” is a palatable excuse. Language barriers, culture differences and silos will continue to take its toll on the stability of major companies.
Yes, I agree. I know too many Indian workers have doctorates but when tested they just have a high school education. The Indian government is behind this. Also, this site needs to add up/down votes and sub chats. Being a tech site this shouldn’t be impossible. Right?
OpenAI confirms ChatGPT is down worldwide
Hackers hijack hotel Wi-Fi DNS to steal Microsoft 365 accounts
Chick-fil-A data breach affects more than 13,000 customers
Rev5 is ending. See what your FedRAMP 20x transition really requires
Calculate what you’d save by replacing your MDR.
Your Scanners Are Green. Your Pipeline Might Not Be. Here’s How to Close the Gap.
AI agents can speed up ransomware attacks. See how Acronis helps reduce the risk.
Overdue a password health-check? Audit your Active Directory for free
Terms of Use – Privacy Policy – Ethics Statement – Affiliate Disclosure
Read our posting guidelinese to learn what content is prohibited.



