Act 0: Assume The Best

what could possibly go wrong?

Act 1: Route Around It

One instance was down - let's just ask another one

Act 2: Try Again

A common strategy!

Act 3: Jitterbug

Backoff spreads requests out, and jitter makes sure that clients don't all retry simultaneously

The green dashed line marks when the service recovers

Act 4: Calm Yourself

When a service keeps failing, stop calling it for a while.

// the short version

One principle.
Four patterns.

Every service has a capacity. Stay under it, or fail gracefully.

01
Route around it. Health checks plus failover — when there is a healthy path.
02
Retry with backoff and jitter. Spread retries out so they don’t become the next outage.
03
Trip the circuit. Stop calling a service that’s clearly over capacity. Probe later.
04
Know your capacity. You can’t respect a limit you can’t see. Measure request and error rates per service in production.
// see your own capacity

Watch these patterns in your own apps.

Datadog APM surfaces request rates, error rates, retry storms, and circuit-breaker trips across every service — so resilience issues are obvious before they page you at 3am.

Try Datadog APM →