← DDIA

Chapter 01 · Foundations

Reliable, Scalable, and Maintainable Applications

Why did Twitter almost break under its own timeline?

Twitter serves ~300k timeline reads/s and ~12k tweets/s. Same product, opposite extremes.

Three acts: 1. fan-out 2. performance 3. framework

Understanding Percentiles

The average sits above p50, dragged up by the tail.

1500ms 250ms 50ms p50 p99 p99.9
0 / 3
p50 (Median) 50% of users
typical request 50ms
p99 (99th) 1 in 100 users
slow / mild queueing 250ms
p99.9 (99.9th) 1 in 1000 users
tail events users feel 1.5s

One page makes N backend calls. Its p99 is the slowest of N. A "1 in 100" event fires per call, not per page.

Takeaway

Average latency lies. p99 is what your users feel.

Three pillars of a data-intensive app.

01

Reliability

the system keeps working correctly, even when things go wrong.

HARDWARE FAULTS disks die, RAM flips, racks lose power.
SOFTWARE ERRORS runaway processes, cascading bugs, bad configs.
HUMAN ERRORS operators are the top cause; design to make mistakes safe.
02

Scalability

the system copes when load grows.

LOAD requests/s, follower fan-out, read/write ratio.
PERFORMANCE use percentiles: p99, p99.9. averages hide the tail.
COPING vertical vs horizontal. stateless vs stateful.
03

Maintainability

people can keep working on the system.

OPERABILITY make it easy for ops to keep the lights on.
SIMPLICITY manage accidental complexity.
EVOLVABILITY bend to new requirements without breaking.

everything in ddia hangs off these three concerns. the rest of the book is one long argument about how they trade off.