Better Call Fallback

Designing resilient services

Tomasz Nurkiewicz

nurkiewicz.com

About this talk

Building resilient services by assuming failures will happen and preparing for them. It walks through practical patterns like retries, circuit breakers, bulkheads, idempotency, and outbox, plus how to test and observe them. The goal is to design systems that degrade gracefully instead of breaking completely.

The talk focuses on designing services that can survive and recover from failures instead of pretending they won’t happen. It shows why the “happy path” mindset is dangerous in production, where rare bugs surface quickly at scale. Engineers are guided through a set of proven resilience techniques: retries with backoff and jitter, circuit breakers, timeouts, throttling, and bulkheading to contain failures. For asynchronous systems, it covers queues, DLQs, idempotency, and the outbox pattern to handle retries without duplication. In distributed systems, it discusses leader election, offline-first strategies, and clear separation of liveness vs. readiness probes. Architectural safeguards like dry runs, sagas, compensating transactions, graceful degradation, and killswitches are highlighted as essential for containing damage when things go wrong. Testing resiliency with chaos engineering tools and monitoring through metrics, logs, and traces are stressed as key practices. Finally, the cultural side: post-mortems, pre-mortems, and production readiness—ensures teams take failures seriously and learn from them. The message is simple: failures are inevitable, but with the right patterns, they don’t have to take your system down.

Key takeaways from this talk:

  • Failures Are the Norm, Not the Exception
  • Core Resiliency Patterns: fallbacks, retries, circuit breakers, bulkheading
  • Asynchronous & Distributed System Strategies
  • Architectural & Operational Safeguards
  • Cultural Practices for resilient programming

You keep using that word happy path

If builders built houses the way programmers built programs,
the first woodpecker to come along would destroy civilization
Gerald Weinberg

Fragile

One in a million

...happens every 15 minutes at 1k RPS

P99 latency

facebook.com makes 350 HTTP requests

How to...

  • ...avoid
  • ...deal with
  • ...bypass

...failures

Part I: RPC

try-catch

  • Repay loan
    • Send bank transfer
      • Update daily report
        • Generate PDF
          • Send report over email
            • Wait for SMTP confirmation
              • SocketException

Fallback

Retries

Retry-After HTTP header

Retry-After HTTP header in practice

Exponential backoff

...and jitter

Timeouts

12 minutes!

                
                    64 bytes from 216.58.212.14: icmp_seq=9720 ttl=114
                    time=749147.422 ms
                
            

Bulkheading

en.wikipedia.org/wiki/Bulkhead_(partition)

source

Throttling

en.wikipedia.org/wiki/Throttle

Circuit Breaker

en.wikipedia.org/wiki/Circuit_breaker

To consider

  • What conditions open it? 4xx vs 5xx
  • Does it ever close?
  • Granularity?

Thundering herd problem

www.reddit.com/r/CFB/comments/2q7tj8/...

Cache stampede

How a Cache Stampede Caused One of Facebook’s Biggest Outages

Sources: [1], [2]

Part II: Asynchrony

You won't get a failure

if you don't wait for the result

Thread pools

But also:

  • Coroutines/Goroutine
  • Virtual threads
  • Actors/agents

Message queue

Receiver is broken...

  • retries?
  • timeouts?
  • DLQ
  • head-of-line blocking?

At-least once

Idempotency

E.g.: Idempotency-Key HTTP Header

datatracker.ietf.org/doc/draft-ietf-httpapi-idempotency-key-header

Outbox pattern

Part III: distributed systems

Leader election

Leader ≠ singleton worker

Local database

Think: asynchronous replication, local SQLite

Offline-first

Failure backlog

Liveness vs Readiness

Part IV: Architecture

Dry run

aka. simulation

terraform plan

Compensating transaction

Saga

Throttling side effects

"we sent 700 thousand text messages to one person"

mastodon.social/@nurkiewicz/109259595653594613

Killswitch

Knight Capital took a [...] loss of $440 million in 45 minutes

Knight Capital Group: 2012 stock trading disruption

$1m per 6 seconds. Bill Gates makes < $8k per 6 seconds

Graceful degradation

Tiered services

Testing resiliency

Chaos engineering

Who is this Chaos Monkey and why did he crash my server?

Toxiproxy

                
                    Proxy mysqlProxy = 
                        client.createProxy("mysql", "localhost:13306", "localhost:3306");
                    mysqlProxy
                        .toxics()
                        .latency("latency", DOWNSTREAM, 100)
                        .setJitter(15);
                
            

Observability

  • Metrics
  • Logs
  • Traces

Death by a thousand dashboards

...and by millions of $ in APM bill

Part V: Culture

  • Post-mortem
  • Pre-mortem
  • Production readiness

Further reading

  1. Cloud Design Patterns
  2. Resilience4j (Java), Polly (.NET)
  3. Building Resilient Distributed Systems book by Sam Newman
  4. SE Radio 643: Ganesh Datta on Production Readiness
  5. Pre-mortem project management strategy

Thank you!