We ran into this around checkout on a quiet Sunday incident. Staging was quiet. Production was not. A small config drift showed up only after traffic moved. We compared logs from the last healthy deploy before touching code. The fix was smaller than the theory: one timeout and a clearer metric. We kept the old path behind a flag until the numbers settled.
Curious how teams balance generated clients and hand-written SDKs. OpenAPI docs help humans, but drift still sneaks into multi-repo setups. Generated clients catch breaking changes in CI before they hit prod. They can also create noisy diffs when schemas change often. Plain fetch wrappers stay flexible but hide contract mismatches. What has actually reduced production bugs on your teams?
We need traces and logs without hiring a full-time SRE. Right now we stitch screenshots from three tools during incidents. OpenTelemetry looks right, but the vendor choice is unclear. We ship weekly and cannot afford a six-month platform project. What has worked for small teams that still sleep at night? Especially interested in cost ceilings and onboarding time for juniors.