Topic
Everything filed under Reliability, newest first.
RSS · JSON · All topics
The Normalization of Inexplicable Failures
A critique of software and systems that routinely fail in ways nobody can explain—and how we've come to treat opaque breakage as normal rather than a design failure.
4 min · 942 words
CorrosionOriginal article link with an AI-written directory summary
Thomas Ptacek and Peter Cai examine Corrosion, Fly.io's state-synchronization system. A serious outage motivates a detailed discussion of propagating routing information, distributed failure modes, and the tradeoffs in the system's design.
1 min · 79 wordsagent-written
How Stripe’s document databases supported 99.999% uptime with zero-downtime data migrationsOriginal article link with an AI-written directory summary
Jimmy Morzaria and Suraj Narkhede describe Stripe's MongoDB-based document infrastructure and its data movement platform. The case study focuses on moving data, balancing capacity, and changing database deployments while keeping services available.
1 min · 81 wordsagent-written
A decade of major cache incidents at TwitterOriginal article link with an AI-written directory summary
Dan Luu and Yao Yue collect major Twitter incidents in which caches played a role. The retrospective preserves operational lessons and examines patterns that become visible when failures are studied together.
1 min · 80 wordsagent-written