Microsoft Incident Halts Azure West US Services for Over 4 Hours

Maintenance error disrupts critical Azure services, causes widespread outage

Cloud users relying on Microsoft’s Azure services in the West US region faced an extended disruption Friday

The global technology leader apologized for service degradation caused by a

mishap during scheduled maintenance work.

Microsoft traced the

outage to a bug in its network equipment that incorrectly processed

maintenance requests, leading to the removal of critical routing data

between datacenters and the wider network.

“During this maintenance, an error in the device

configuration system improperly identified additional network routes

for deactivation, resulting in the loss of essential network paths

and causing significant traffic disruption,” the

company stated in its post-incident report.

The outage, which began at

14:44 UTC and lasted until 18:26 UTC, impacted 27 Azure

services, including popular offerings like Storage, SQL

Database, and Azure Active Directory. Users reported

intermittent access issues, latency spikes, and complete

service downtime.

Microsoft initiated its

troubleshooting process immediately after detecting

anomalies in network routing behavior and realized packet

loss in its Wide-Area Network (WAN).

Initial diagnostics revealed

excessive route churn — a phenomenon where routing tables

experience rapid changes — stemming from the West US region’s

datacenter. Further analysis confirmed that the maintenance

activity had triggered the unintended removal of routes.

“We observed

significant changes in our WAN traffic patterns, which a

minute later aligned with the scheduling of fiber network

maintenance in the West US region,” the engineering team’s

report states.

“The oversight occurred due to a

fault in the system that converts device-specific

configuration requests into standardized commands. This led

to invalid route deletions, impacting connectivity across

multiple Azure services.”

By 18:26 UTC, Microsoft

successfully completed the rollback of the faulty

configuration. “All affected services have returned

to normal operation by 19:41 UTC,” the company stated

concluding its root cause analysis. The timeline reveals

the outage lasted for nearly four and a half hours,

with recovery efforts taking approximately two hours.

Microsoft has acknowledged the

gap between its standard maintenance protocols and the

failure, which triggered a partial outage. While cloud

services like AWS and Google Cloud had previously managed

outages without public disruption, this incident

highlights the vulnerability of even the most

mature cloud infrastructures to internal system flaws.

Also Read

Source link

Exit mobile version