We ran into this around session cookies when the cache went cold. I expected a tooling problem. It was an assumption in the schema. EXPLAIN and a short rollback window saved the afternoon. Writes slowed first, then a handful of reads timed out. Dropping the unused index and batching the backfill unblocked the queue. I now review write paths with the same care as the happy-path query.
We ran into this around feature flags during a traffic spike. Docs helped humans. Generated types caught the breakages in CI. Schemas still drift when several repos move at different speeds. Hand-written wrappers stayed flexible and hid mismatches. The noisy diffs were worth it once a breaking change never shipped. We kept the generator and deleted the duplicate sample clients.
We ran into this around search ranking when the cache went cold. The error rate looked fine in the average and ugly in one slice. Two percent of requests reset upstream with no application exception. Idle timeouts between the proxy and the app had drifted apart. Aligning them stopped the intermittent failures within an hour. We added a panel for reset reasons so the next page is faster to find.
We ran into this around CI pipelines during a traffic spike. Incidents were a scavenger hunt across three dashboards. We needed traces and logs without a platform project. OpenTelemetry was the shape. The vendor was the real decision. A weekly budget cap and one starter dashboard got the team using it. Juniors could follow a request without asking who owned the graphs.
We ran into this around invoice generation when the cache went cold. We did not rewrite the service. We moved one endpoint at a time. Dual writes ran until nightly totals matched for several days. A flag sent a small percentage of reads to the new path. The hard part was matching rounding the old code had hidden. Cutover finished with a small blast radius and an easy rollback.
We ran into this around email digests during a traffic spike. Release notes used to depend on someone remembering to write them. A small bot grouped merged pull requests and asked for missing summaries. It started as a cron job and later listened to webhooks. Product stopped chasing engineers for copy after each weekly release. The rough edges are labeling discipline and weekend merges.
We ran into this around push notifications when the cache went cold. We needed a lock so two workers would not process the same job. One option expired itself. The other lived next to our source of truth. Failover behavior mattered more than the benchmark. We started with the store we could operate, then moved the hot path later. Pick the tool your on-call can reason about.
We ran into this around background jobs during a traffic spike. I expected a tooling problem. It was an assumption in the schema. EXPLAIN and a short rollback window saved the afternoon. Writes slowed first, then a handful of reads timed out. Dropping the unused index and batching the backfill unblocked the queue. I now review write paths with the same care as the happy-path query.
We ran into this around team onboarding during a traffic spike. The work was mostly unblocking other people, not closing my own tickets. Clear noes protected the roadmap more than extra hours. Writing the doc nobody wanted still changed how fast the team moved. Impact was hard to see until a project stalled without that context. I am still learning to describe that work without vanity metrics.
We ran into this around rate limits during a traffic spike. The error rate looked fine in the average and ugly in one slice. Two percent of requests reset upstream with no application exception. Idle timeouts between the proxy and the app had drifted apart. Aligning them stopped the intermittent failures within an hour. We added a panel for reset reasons so the next page is faster to find.
We ran into this around image uploads when the cache went cold. Docs helped humans. Generated types caught the breakages in CI. Schemas still drift when several repos move at different speeds. Hand-written wrappers stayed flexible and hid mismatches. The noisy diffs were worth it once a breaking change never shipped. We kept the generator and deleted the duplicate sample clients.
We ran into this around data exports during a traffic spike. Retries arrived out of order and our fixtures only modeled the happy path. Sleeping in tests hid the races and slowed CI. Replaying a small set of recorded payloads found the idempotency bug. We asserted on stored keys, not on elapsed time. The suite is shorter and finally recreates the incident we cared about.