We ran into this around push notifications on a quiet Sunday incident. We needed a lock so two workers would not process the same job. One option expired itself. The other lived next to our source of truth. Failover behavior mattered more than the benchmark. We started with the store we could operate, then moved the hot path later. Pick the tool your on-call can reason about.
We ran into this around invoice generation during a Friday deploy. We did not rewrite the service. We moved one endpoint at a time. Dual writes ran until nightly totals matched for several days. A flag sent a small percentage of reads to the new path. The hard part was matching rounding the old code had hidden. Cutover finished with a small blast radius and an easy rollback.
Curious how teams balance generated clients and hand-written SDKs. OpenAPI docs help humans, but drift still sneaks into multi-repo setups. Generated clients catch breaking changes in CI before they hit prod. They can also create noisy diffs when schemas change often. Plain fetch wrappers stay flexible but hide contract mismatches. What has actually reduced production bugs on your teams?
We need traces and logs without hiring a full-time SRE. Right now we stitch screenshots from three tools during incidents. OpenTelemetry looks right, but the vendor choice is unclear. We ship weekly and cannot afford a six-month platform project. What has worked for small teams that still sleep at night? Especially interested in cost ceilings and onboarding time for juniors.