We ran into this around image uploads during a traffic spike.
Staging was quiet. Production was not.
A small config drift showed up only after traffic moved.
We compared logs from the last healthy deploy before touching code.
The fix was smaller than the theory: one timeout and a clearer metric.
We kept the old path behind a flag until the numbers settled.
Comments (0)