We ran into this around push notifications when the cache went cold. We needed a lock so two workers would not process the same job. One option expired itself. The other lived next to our source of truth. Failover behavior mattered more than the benchmark. We started with the store we could operate, then moved the hot path later. Pick the tool your on-call can reason about.
We ran into this around team onboarding after the replica failover. The work was mostly unblocking other people, not closing my own tickets. Clear noes protected the roadmap more than extra hours. Writing the doc nobody wanted still changed how fast the team moved. Impact was hard to see until a project stalled without that context. I am still learning to describe that work without vanity metrics.
Our provider retries aggressively and out of order under failure. Naive fixtures make CI slow and still miss race conditions. Looking for patterns that keep suites fast and realistic. Do you fake the provider clock, or replay recorded payloads? How do you assert idempotency without flaky sleeps? Share a setup that survived production incident recreations.
We moved session checks to the edge to cut latency on every page load. It worked in staging, then failed on preview deploys when cookies crossed domains. Clock skew between edge and origin made short-lived tokens look expired. We fixed cookie domains per environment and added skew-tolerant expiry. Median auth path dropped about 120ms, with fewer cold-start surprises. Lesson: test cookies across every environment before calling a migration done.