Start with a workload, not a diagram
System design is the process of deciding how components cooperate to meet requirements. A diagram with load balancers, queues, caches, and many services is not automatically a good design. Every component solves a problem while adding operational work and new failure modes.
Begin with concrete questions: how many requests arrive, what does each request do, how much data exists, how quickly must responses arrive, and what happens when a dependency fails? Separate current needs from plausible growth so the design does not solve an imaginary scale problem at the expense of today's maintainability.
Define success in measurable terms
Latency is how long an operation takes. Throughput is how much work is completed per unit time. Concurrency is how much work is in progress at once. They are related but not interchangeable. A service can handle many requests per second while giving some users unacceptable delays.
Percentiles reveal tail behavior. A p95 latency of 300 milliseconds means 95 percent of measured requests complete within that time for the chosen window and population. An average can hide a minority of extremely slow requests.
Define a service-level objective around user-visible behavior, such as successful checkout requests within a time budget. Include availability, correctness, and recovery objectives where relevant. Be precise about measurement: a server returning HTTP 200 with an application error is not necessarily a successful user action.
Find the bottleneck before adding capacity
A typical request may spend time in network setup, application code, database queries, and external calls. Use traces, query statistics, and resource metrics to locate the constraint. CPU, memory, database connections, disk I/O, and locks can each become limiting resources.
Browser → Web service → Database
↘ External payment API
If slow queries dominate, adding web servers may increase pressure on the database. If an external payment service has a strict rate limit, more workers may increase errors instead of throughput. Fix inefficient work before multiplying it.
Scale the application tier
Vertical scaling adds resources to one instance. It is often the simplest first step, though every machine has a practical capacity limit and remains a potential failure point. Horizontal scaling adds instances behind a load balancer.
To make horizontal scaling predictable, avoid storing critical state only in an instance's memory or local filesystem. Use an appropriate session store, durable storage, and shared database. Configure graceful shutdown so instances stop accepting new work while existing requests complete within a bounded period.
A load balancer distributes traffic; it does not make the application resilient by itself. Health checks, timeouts, instance capacity, and deployment behavior all need attention.
Cache expensive, repeatable reads
A cache can avoid repeated computation or database access. A CDN caches suitable public responses near users; an application cache may store query results or derived data.
Define a key, lifetime, invalidation policy, and failure behavior. Include permission scope in keys for private data. Decide whether stale results are acceptable. A cache miss storm after expiration or a restart can overload the original dependency; request coalescing, jittered expiry, and controlled warming can help.
Do not rely on cached stock counts as the final authority for purchasing inventory. Use the database or another authoritative transactional mechanism for decisions that must remain correct under concurrent requests.
Move suitable work to queues
Tasks such as sending emails, generating previews, and processing uploaded files can often run asynchronously. A queue separates request handling from slower work and absorbs bursts.
Many systems provide at-least-once delivery, so a worker can receive a job more than once. Make operations idempotent where possible, using identifiers or uniqueness constraints to avoid duplicate effects. Set bounded retries, backoff, dead-letter handling, and visibility into queue age.
A queue changes the user experience: completion becomes eventual. Show a meaningful pending state, and define what happens if processing fails. It cannot hide unlimited work forever; an arrival rate above processing capacity will grow the backlog.
Understand data scaling tradeoffs
Read replicas can distribute some database reads, but replication lag may produce stale results. Route correctness-sensitive reads appropriately. Partitioning or sharding can distribute larger datasets, while making joins, transactions, rebalancing, and operational procedures harder.
Before sharding, examine indexes, query patterns, data retention, hardware, and connection pooling. A well-tuned relational database can support substantial workloads, and premature distribution can introduce more difficulty than it removes.
| Component | Helps with | Adds questions about |
|---|---|---|
| Cache | Repeated reads and latency | Staleness and invalidation |
| Queue | Bursts and background work | Retries and eventual completion |
| Replica | Read load and some resilience | Lag and failover |
| Sharding | Dataset or write distribution | Routing and cross-shard operations |
Design for partial failure
Use deadlines and timeouts for dependency calls. Retry only appropriate operations, with backoff and jitter; uncontrolled retries can amplify an outage. Limit concurrency, shed excess load, and provide graceful degradation when a nonessential dependency fails.
Keep enough simplicity that the team can understand a failure under pressure. Load-test representative workflows, inspect the system near its limits, and document recovery steps. Scale is achieved by removing measured constraints while preserving correctness and operability, one justified change at a time.