UmurInan
Reliability Topic

Reliability

Reliable systems are built from unreliable parts. These posts cover the patterns that absorb failure rather than amplify it: retry budgets, idempotency keys, circuit breakers, webhook delivery, and what it actually takes to ship something that does not fall over at 3am.

Devops

Posted on Jun 8, 2026

Your Disaster Recovery Plan Is Fiction Until You Run It

A DR plan you've never run is a hypothesis, not a plan. What only breaks on the real restore, why RTO is fiction until you measure it, and the game-day fix.

Read more
Backend

Posted on May 28, 2026

Why Your Distributed Lock Doesn't Lock

Distributed locks don't provide mutual exclusion. Fencing tokens, GC pauses, clock drift, and why the lock you wrote is actually a polite hint at best.

Read more
Backend

Posted on Apr 19, 2026

The Thundering Herd Problem

Cache stampedes, retry storms, reconnect floods: three failure modes with the same root cause. Synchronized behavior under load amplifies failures every time.

Read more
Backend

Posted on Apr 14, 2026

Webhook Reliability: The Lost Art

Webhooks break predictably: duplicate events, missed deliveries, retry storms. Here is what it actually takes to build receivers that hold up in production.

Read more
Backend

Posted on Apr 6, 2026

Rate Limiting Is Harder Than It Looks

Token bucket, sliding window, fixed counter: rate limiting algorithms all sound simple until you actually implement them correctly across distributed systems.

Read more
Devops

Posted on Apr 1, 2026

Monitoring Is Not a Dashboard

Real monitoring is not a Grafana dashboard. It is knowing which questions to ask, which signals answer them, and what to do when the answer is unexpected.

Read more
Devops

Posted on Mar 29, 2026

The Deploy That Took Down Friday

Friday deploys have a reputation for a reason. Here's why they go wrong, what guardrails actually help, and when it's okay to ship on a Friday anyway.

Read more