We ran into this around email digests after the replica failover. Release notes used to depend on someone remembering to write them. A small bot grouped merged pull requests and asked for missing summaries. It started as a cron job and later listened to webhooks. Product stopped chasing engineers for copy after each weekly release. The rough edges are labeling discipline and weekend merges.
We ran into this around feature flags during a Friday deploy. Staging was quiet. Production was not. A small config drift showed up only after traffic moved. We compared logs from the last healthy deploy before touching code. The fix was smaller than the theory: one timeout and a clearer metric. We kept the old path behind a flag until the numbers settled.
No spike in CPU. Error budgets looked fine at a glance. Users still reported blank pages in a thin slice of traffic. Logs only showed upstream resets with no clear application exception. The culprit was a stale keep-alive timeout between nginx and the app. Aligning idle timeouts stopped the intermittent 502s within an hour. We also added a dashboard for upstream reset reasons so the next page is faster.
We extracted one high-churn billing endpoint behind a strangler facade. Dual-writes ran for two weeks while we compared totals nightly. A feature flag controlled read traffic so we could roll back instantly. The hardest part was matching edge-case rounding in legacy invoices. Cutover finished with no customer-facing downtime and a smaller blast radius. We kept the facade until three more endpoints followed the same path.
We needed locks so queue workers did not process the same job twice. Redis SET NX was faster under load and easy to expire automatically. Postgres advisory locks were simpler operationally for our small team. Failover behavior mattered more than raw latency in our case. We chose Postgres first, then moved hot paths to Redis later. Pick the lock store you can operate confidently at 3am.