İçeriğe atla
EchelonVPS
operations · 2026-02-13

Post-mortem: two NVMe drives, one mirror, four minutes

On 11 February a firmware bug took out both halves of a RAID-10 mirror in Frankfurt. 218 instances were affected. Nobody lost data. Here is exactly what happened.

Güncellendi
2026-02-13
Infrastructure
8 min
Tümü
2 / 10

At 04:12 UTC on 11 February, a drive in cluster 12 at FRA6 stopped responding to NVMe admin commands. At 04:16 its mirror partner did the same. Both were from the same manufacturing lot, both had accumulated almost identical write volumes, and both hit the same firmware assertion within four minutes of each other.

Timeline

UTCEvent
04:12First drive drops. Array degrades, alert fires.
04:16Mirror partner drops. Array offline. 218 instances lose their disk.
04:19On-call engineer paged and online.
04:31Cause narrowed to firmware, not the backplane. Decision to restore from the cluster replica rather than attempt recovery.
04:48Restore begins on standby chassis.
05:46All 218 instances running. Data current to 04:11.
05:52Status page updated with preliminary cause.

Why both halves failed

RAID-10 protects against independent failures. This was a correlated one: same lot, same firmware, same wear. Our procurement had been buying single lots for a chassis because it simplifies inventory, which quietly converted a redundancy story into a single point of failure.

What changed

  • 01Mirror pairs are now built from different manufacturing lots. Inventory got harder; the failure domain got smaller.
  • 02The affected firmware revision was replaced fleet-wide within eleven days, on all 214 racks.
  • 03Wear divergence between mirror halves is now a monitored metric, with an alert when two drives in a pair track within 2% of each other.
  • 04The cluster replica restore path, which had never been used at this scale, is now exercised monthly on a live cluster.

What we owe you

Ninety-four minutes of downtime is 0.22% of a month. Our SLA credits at ten times the outage, so affected accounts were credited two days each, applied automatically without anyone opening a ticket. Fourteen customers asked for a full refund instead and got one.

A redundancy story you have not tested is a hypothesis, not a plan.