← DDIA Chapter 01 · Foundations
Reliable, Scalable, and Maintainable Applications
Why did Twitter almost break under its own timeline?
Twitter serves ~300k timeline reads/s and ~12k tweets/s.
Same product, opposite extremes.
Three acts: 1. fan-out
2. performance
3. framework
Act 1 Scalability
every refresh queries each followed author.
1 refresh → N reads.
every tweet is pre-copied into every follower's cache.
1 tweet → N writes. Reads become 1 lookup.
one celebrity tweet becomes 90 million cache writes.
The write path collapses at celebrity scale.
write-fanout for the many · read-merge for the few.
Twitter runs both paths, one per user class.
Act 2 Performance
0 / 3
p50 (Median) 50% of users
typical request 50ms
p99 (99th) 1 in 100 users
slow / mild queueing 250ms
p99.9 (99.9th) 1 in 1000 users
tail events users feel 1.5s
One page makes N backend calls. Its p99 is the slowest of N. A "1 in 100" event fires per call, not per page.
Takeaway
Average latency lies. p99 is what your users feel.
Act 3 Framework
Three pillars of a data-intensive app.
01 Reliability
the system keeps working correctly, even when things go wrong.
HARDWARE FAULTS disks die, RAM flips, racks lose power.
SOFTWARE ERRORS runaway processes, cascading bugs, bad configs.
HUMAN ERRORS operators are the top cause; design to make mistakes safe.
02 Scalability
the system copes when load grows.
LOAD requests/s, follower fan-out, read/write ratio.
PERFORMANCE use percentiles: p99, p99.9. averages hide the tail.
COPING vertical vs horizontal. stateless vs stateful.
03 Maintainability
people can keep working on the system.
OPERABILITY make it easy for ops to keep the lights on.
SIMPLICITY manage accidental complexity.
EVOLVABILITY bend to new requirements without breaking.