Day 2 A-Foundations 2026-09-29 ← All lessons

Scaling out: load balancers, replication, cache, CDN

One server cannot carry millions of users. This lesson shows the exact order of moves that takes a single box to a resilient, multi-tier system.

24Vertical scaling ceiling
10Single master in production
4Sharding hash example

Key points

Outcomes

01Choose between vertical and horizontal scaling using the failover and limit criteria.
02Explain how a load balancer and master/slave replication each restore availability.
03Describe the read-through cache flow and when a CDN is the right tool.
01

Scale up or scale out


Option A

Vertical scaling (scale up)

  • Add CPU, RAM, DISK to an existing server
  • Simple, no code changes
  • Hard limit on CPU and memory
  • No failover or redundancy
  • Powerful servers cost much more
Option B

Horizontal scaling (scale out)

  • Add more servers to the pool
  • Desirable for large scale applications
  • Survives one server going offline
  • Load balancer distributes traffic
  • Requires stateless web tier
GotchaVertical scaling is a great option when traffic is low. Its main advantage is simplicity. Its main danger is that one server going down takes the whole site with it.
02

Load balancer and database replication


How to read: Follow arrows left to right: users hit the load balancer, which routes to web servers, which write to the master and read from slaves.

flowchart LR U[User] --> DNS[DNS] DNS --> LB[Load balancer public IP 88.88.88.1] LB --> S1[Server 1 private IP 10.0.0.1] LB --> S2[Server 2 private IP 10.0.0.2] S1 --> M[(Master DB)] S2 --> M M --> R1[(Slave DB 1)] M --> R2[(Slave DB 2)] S1 --> R1 S2 --> R2
Web tier behind a load balancer and data tier with master/slave replication.
  1. User gets the load balancer IP from DNS.

  2. User connects to the load balancer with that IP.

  3. HTTP request is routed to Server 1 or Server 2.

  4. Web server reads user data from a slave database.

  5. Web server routes write, update, and delete operations to the master database.

GotchaIf the master goes offline, a slave is promoted to master. In production this is harder because the slave may be behind, so data recovery scripts run to fill the gap.

How to read: Watch the dots: each one is a request. The balancer alternates servers, so consecutive dots take different branches.

YoubrowserLoad balancerround robinServer 1Server 2
  1. Your request arrives at the balancer, the single public entry point.

  2. The balancer picks the next server in rotation: Server 1, then Server 2.

  3. If a server stops responding, it leaves the rotation until healthy.

Request path through the load balancer to the web tier.
03

Cache tier and CDN


How to read: Follow arrows left to right: the web server asks the cache first, and only on a miss does it query the database and write the result back.

flowchart LR W[Web server] --> C{Cache has data?} C -->|yes| R[Return data to web server] C -->|no| D[(Database)] D --> S[Save data to cache] S --> R
Read-through cache: check cache first, fall back to database, then store the result.
Option A

Cache tier

  • Temporary data store layer, faster than database
  • Stores expensive responses or frequently accessed data
  • Read-through strategy: check cache, then DB, then store
  • Use when data is read often but modified rarely
  • Risk: single point of failure, so use multiple cache servers
Option B

CDN

  • Network of geographically dispersed servers
  • Delivers static content: images, videos, CSS, JavaScript
  • Closest CDN server serves the user
  • Origin returns file with optional TTL header
  • Risk: cost, stale content, and CDN outage fallback
Interview tipCache eviction happens when the cache is full. Least-recently-used (LRU) is the most popular policy. Other options are LFU and FIFO.
04

Stateless web tier


Option A

Stateful architecture

  • Server remembers client data from one request to the next
  • User A's session and profile live on Server 1
  • Requests from User A must route to Server 1
  • Sticky sessions add overhead
  • Adding or removing servers is difficult
Option B

Stateless architecture

  • Server keeps no state information
  • HTTP requests can go to any web server
  • State is fetched from a shared data store
  • Simpler, more robust, and scalable
  • Autoscaling adds or removes servers based on traffic
Interview tipMove session data out of the web tier into persistent storage such as a relational database, Memcached/Redis, or NoSQL. NoSQL is chosen here because it is easy to scale.
Q&A

Check yourself


Q1Why is horizontal scaling preferred over vertical scaling for large applications?
  • Vertical scaling has a hard limit and no failover, while horizontal scaling adds servers and survives one going offline.
  • Vertical scaling requires rewriting the application for multiple servers.
  • Horizontal scaling removes the need for a database.
✓ Vertical scaling has a hard limit and no failover, while horizontal scaling adds servers and survives one going offline. — Vertical scaling hits a hardware ceiling and has no redundancy, so a single failure takes the site down.
Q2In a master/slave replication setup, which operations go to the master database?
  • All read operations only.
  • All write, update, and delete operations.
  • Both reads and writes are split evenly across all nodes.
✓ All write, update, and delete operations. — The master handles data-modifying commands, while slaves serve reads.
Q3What is the main purpose of a CDN?
  • To store user session data for the web tier.
  • To cache static content on servers close to the user.
  • To replace the database for write operations.
✓ To cache static content on servers close to the user. — A CDN delivers static assets like images and CSS from a server geographically near the user.
Sources: System Design Interview, Ch. 1, Scale from zero to millions of users (pp. 1–30)