what could possibly go wrong?
One instance was down - let's just ask another one
A common strategy!
Backoff spreads requests out, and jitter makes sure that clients don't all retry simultaneously
The green dashed line marks when the service recovers
When a service keeps failing, stop calling it for a while.
Every service has a capacity. Stay under it, or fail gracefully.
Datadog APM surfaces request rates, error rates, retry storms, and circuit-breaker trips across every service — so resilience issues are obvious before they page you at 3am.
Try Datadog APM →