Horizontal Scaling & Load Balancing
Turn a single-box web tier into an interchangeable fleet: choose L4 vs L7 and the right algorithm, keep deploys from dropping in-flight requests, and let services find healthy instances under constant churn.
Horizontal vs Vertical Scaling & Stateless Services
Scale out the web tier by externalizing sessions and files so any node serves any request; scale up (then shard) the stateful data tier.
Load Balancer Fundamentals: L4 vs L7
L4 is fast, protocol-agnostic, and content-blind; L7 routes on content and terminates TLS at a CPU cost; production stacks L4 at the edge in front of an L7 fleet, all made HA.
Load-Balancing Algorithms & Session Affinity
Least-connections for variable durations, power-of-two-choices for large pools, consistent hashing for cache-warm stickiness, and affinity used deliberately, not by default.
Health Checks, Draining & Graceful Rollout
Separate liveness (restart) from readiness (pull from pool), drain in-flight work before terminating, slow-start cold nodes, and keep deep checks from failing the whole fleet.
Service Discovery & Client vs Server-Side Load Balancing
A registry (heartbeats or k8s readiness-driven Endpoints) keeps healthy addresses current within seconds; choose server-side simplicity or client-side/mesh locality deliberately.
Global Traffic & Gateway
Route global users to the nearest healthy region and fail a region out in under a minute, design gateways and BFFs that keep services thin, and manage TLS termination plus long-lived connections at scale.
Global & DNS-Level Load Balancing (GSLB, Anycast)
GeoDNS steers coarsely but is TTL-bound; anycast plus BGP withdrawal gives seconds-scale failover; active-active regions need headroom to absorb a lost region.
API Gateway & Backend-for-Frontend
A thin, horizontally scaled gateway owns north-south cross-cutting concerns, BFFs shape payloads per client type, the mesh owns east-west, and business logic stays in services.
TLS Termination & Connection Management
Terminate TLS at the edge and re-encrypt with mTLS inside, pool and keep-alive connections, and fix H2/gRPC/WebSocket pinning with client-side or per-stream balancing.
Rate Limiting & Overload
Keep a service alive when demand exceeds supply: pick a rate-limiting algorithm with an exact response contract, enforce one global limit across a fleet, and shed load by priority instead of collapsing.
Rate Limiting Algorithms
Token bucket for burst-friendly limits, sliding-window counter for accuracy without the log's memory, never raw fixed window, and always a 429 + Retry-After contract.
Distributed Rate Limiting
Naive per-node limits grant Nx; enforce exactly with atomic Redis ops or hybrid local-cache-plus-async-sync for bounded overshoot, always with a fail-open plan.
Load Shedding, Adaptive Concurrency & Backpressure
Shed early and by priority, adapt concurrency limits via Little's Law, bound every queue, propagate deadlines, and brown out features instead of failing everything.
Autoscaling & Isolation
Match compute to demand with reactive, event-driven, and predictive autoscaling despite scaling lag, size fleets from Little's Law plus redundancy math, and bound blast radius with cells and shuffle sharding.
Autoscaling: Reactive, Event-Driven & Predictive
Scale on leading signals (queue depth, RPS) not lagging CPU, and hide the 2-5 minute reactive lag with warm pools, scheduled pre-scaling, and standing headroom.
Capacity Planning & Back-of-Envelope Sizing
Size with Little's Law, divide by a 50-70% utilization target because queues explode near 100%, add N+1 AZ redundancy, and split capacity across reserved/on-demand/spot.
Cell-Based Architecture & Shuffle Sharding
Cells are self-contained stacks behind a dumb HA router that cap any failure at one cell's share; shuffle sharding makes full overlap between two tenants statistically rare.