At 04:12 UTC on 11 February, a drive in cluster 12 at FRA6 stopped responding to NVMe admin commands. At 04:16 its mirror partner did the same. Both were from the same manufacturing lot, both had accumulated almost identical write volumes, and both hit the same firmware assertion within four minutes of each other.
Timeline
| UTC | Event |
|---|---|
| 04:12 | First drive drops. Array degrades, alert fires. |
| 04:16 | Mirror partner drops. Array offline. 218 instances lose their disk. |
| 04:19 | On-call engineer paged and online. |
| 04:31 | Cause narrowed to firmware, not the backplane. Decision to restore from the cluster replica rather than attempt recovery. |
| 04:48 | Restore begins on standby chassis. |
| 05:46 | All 218 instances running. Data current to 04:11. |
| 05:52 | Status page updated with preliminary cause. |
Why both halves failed
RAID-10 protects against independent failures. This was a correlated one: same lot, same firmware, same wear. Our procurement had been buying single lots for a chassis because it simplifies inventory, which quietly converted a redundancy story into a single point of failure.
What changed
- 01Mirror pairs are now built from different manufacturing lots. Inventory got harder; the failure domain got smaller.
- 02The affected firmware revision was replaced fleet-wide within eleven days, on all 214 racks.
- 03Wear divergence between mirror halves is now a monitored metric, with an alert when two drives in a pair track within 2% of each other.
- 04The cluster replica restore path, which had never been used at this scale, is now exercised monthly on a live cluster.
What we owe you
Ninety-four minutes of downtime is 0.22% of a month. Our SLA credits at ten times the outage, so affected accounts were credited two days each, applied automatically without anyone opening a ticket. Fourteen customers asked for a full refund instead and got one.
A redundancy story you have not tested is a hypothesis, not a plan.