- Published on
System Design Building Blocks: A Reference Map
- Authors

- Name
- Mehdi Akiki
I kept running into the same pattern while studying system design: someone explains how to "design WhatsApp" and the diagram ends up full of boxes labeled Kafka, Redis, CDN, WebSocket, and sharding, with not much explanation for what any of those boxes are actually doing there.
So I started collecting notes from the other direction. Instead of starting with big systems, I wanted a clear map of the recurring building blocks. What each one is, what specific problem it solves, when it tends to appear, and what breaks when you misuse it.
This page is that map. It's not an exhaustive reference, just a personal index. Each block will eventually have a dedicated deep-dive page. For now, the summaries here are the starting point.
How to read this
Each entry tries to answer four things: what it is, why it exists, when it shows up, and what tradeoff it brings in.
The quick versions:
- A cache exists because repeated reads are expensive.
- A queue exists because not everything should happen in the request path.
- A CDN exists because users are far from your servers.
- A rate limiter exists because fairness and survival matter.
Once you have a mental model for each block, large system diagrams start feeling less like random chaos and more like familiar combinations of recognizable parts. That's the goal anyway.
The building blocks
Grouped loosely by the kind of problem they address. The groupings aren't rigid (many blocks span categories), but clustering them makes it easier to see which ones tend to appear together and why.
Routing and delivery: how requests reach your system
1. Client
The thing making requests: a web app, mobile app, desktop app, or another service. Rich clients can reduce server load and improve UX, but they add frontend complexity and versioning concerns. Present in every system by definition.
2. DNS
Maps human-readable domain names to IP addresses. Gives you flexibility and indirection. The catch: caching and propagation mean changes can be slow to take effect. Present in any public-facing service.
3. Load Balancer
Distributes incoming traffic across multiple backend instances. You need one as soon as you have more than one server, any failover requirement, or rolling deployments. Improves availability and scalability, but introduces another moving part and sometimes session-routing complexity.
4. API Gateway
A single entry point that handles request routing, authentication, rate limiting, and other cross-cutting concerns. Useful when you have many services and want to centralize shared logic rather than duplicating it everywhere. The risk is it becomes a bottleneck or an overloaded control layer if you put too much in it.
5. CDN
A geographically distributed cache that serves static or cacheable content from locations closer to users. Lower latency and less load on your origin servers. The discipline required is around cache invalidation and making sure freshness rules are explicit.
Compute: where work gets done
6. Stateless Application Server
A backend server that processes requests without holding state between them. The standard choice in modern architectures because stateless servers scale horizontally without friction. Anything that needs to persist lives in external systems: databases, caches, session stores.
7. Background Worker
A process that consumes queued work outside the synchronous request path. Email, image resizing, billing jobs, report generation, retries, anything that would make a user wait. Keeps latency low and isolates failures, but debugging gets harder because the work is no longer linear or immediate.
8. Scheduler / Cron
Triggers tasks on a time-based schedule. Cleanup jobs, billing cycles, digest emails, analytics aggregation. Sounds trivial until you run it across multiple nodes, then duplicate execution, missed runs, and coordination problems become real concerns.
Storage: where data lives
9. Relational Database
Structured, durable storage with tables, rows, schemas, SQL, and transactional guarantees. Payments, orders, users, inventory, any place where a wrong answer costs real money. Horizontal scaling is harder than with NoSQL, and schema changes need to be managed carefully.
10. NoSQL Database
A broad family of databases optimized for simpler access patterns, flexible schemas, or horizontal scale. Key-value lookups, document storage, wide-column workloads, high-ingest pipelines. The fit can be excellent for certain patterns, but joins are often weaker, consistency models differ, and more responsibility moves to the application layer.
11. Read Replica
A copy of a primary database that serves read traffic. Useful when the primary is getting hammered by read queries. Helps with read scalability, but replication lag can break read-after-write expectations. This is the companion to block 24 (Replication): same underlying mechanism, different purpose.
12. Object Storage
Durable blob storage for images, videos, backups, logs, large files. Cheap and nearly infinitely scalable, but the access patterns and retrieval latency are different from databases. Not a good fit for hot, frequently updated records.
13. Cache
A fast storage layer (usually in-memory) that keeps frequently accessed data close to the application. Appears when latency matters or read traffic is heavy. Faster reads and lower backend load, but cache invalidation, stale data, hot keys, and stampedes are very real problems. This is the block people reach for too quickly and also the one they overcomplicate the most.
14. Search Index
A specialized system built for full-text search, ranking, and filtering. SQL LIKE '%query%' stops being acceptable fast. Product search, document search, log search, people search. Good capabilities, but now you have another data store that must stay synchronized with your source of truth.
15. Data Warehouse / Analytics Store
Optimized for large-scale analytical queries rather than transactional workloads. BI dashboards, trend analysis, reporting, historical aggregations. Strong for analytics, but data freshness is often delayed and ETL pipelines must be maintained.
Communication: how components talk to each other
16. Queue
Stores work to be processed asynchronously. Background jobs, email, notifications, video processing, retries, load smoothing. Adds decoupling and resilience, but also eventual consistency, retry logic, and operational overhead. Tightly paired with Background Workers (block 7).
17. Event Stream / Kafka
A durable append-only log for publishing and consuming streams of events across multiple independent consumers. Analytics pipelines, feed fan-out, change data capture, audit logging, large notification systems. It's tempting to treat it as "just a fancier queue," but the differences matter: durability, replayability, multiple consumer groups, ordering within partitions. Partitioning, ordering rules, lag, retention, and operational burden are easy to underestimate.
18. Message Fan-Out Layer
Takes one event and delivers it to many consumers or users. Chat groups, feed generation, push notifications, pub-sub systems. Efficient one-to-many distribution, but scaling hot topics or large recipient lists becomes the core challenge quickly.
19. WebSocket / Realtime Channel
A persistent connection allowing the server to push updates to the client in near real time. Chat apps, collaborative editing, live dashboards, trading apps, presence. Better realtime UX than polling, but connection management, fan-out, and state tracking add complexity that standard request-response doesn't have.
20. Notification System
Delivers messages to users through email, SMS, push, or in-app channels. Present in most consumer apps and SaaS products. Delivery guarantees, retries, deduplication, and user preferences each add their own layer of complexity.
Reliability and coordination: keeping things alive
21. Rate Limiter
Restricts how many requests a client can perform in a given period. Login endpoints, APIs, OTP flows, scraping protection. Improves safety and fairness, but rules that are too strict end up punishing legitimate users.
22. Circuit Breaker / Resilience Layer
Stops repeated calls to failing services to prevent cascading failure. Service-to-service calls, external APIs, payment providers. Better fault tolerance, but fallback behavior must be carefully designed, otherwise users get degraded experiences in confusing ways.
23. Distributed Lock / Coordination Service
Coordinates actions across multiple machines to prevent conflicting work. Leader election, job ownership, failover, deduplication, scheduling. Prevents races, but incorrect locking assumptions can create outages or false confidence. This is the block where "it works on one machine" doesn't transfer to two.
24. Replication
Maintains multiple copies of data across machines or locations. Databases, object storage, distributed logs, caches, anything that needs durability or high availability. Better resilience, but consistency and synchronization become design concerns.
25. Sharding / Partitioning
Splits data or traffic across multiple machines so no single node holds or processes everything. Needed when single-machine limits become real limits. Improves scale, but cross-shard queries, rebalancing, hotspot handling, and operational complexity increase significantly.
Platform and operations: the layer beneath features
26. Service Discovery
A way for services to find each other dynamically. Necessary in microservices, container orchestration, and autoscaling environments where instances come and go. Enables flexibility, but adds a layer that must always be healthy and correct.
27. Authentication and Authorization
Authentication verifies who a user or service is. Authorization determines what they can do. Present in almost every system. Necessary for security, but easy to make inconsistent or hard to reason about across many services.
28. Observability Stack
Logs, metrics, traces, dashboards, and alerts used to understand system behavior. Every serious production system needs one. A distributed system you can't observe is a distributed system you can't operate. Poor instrumentation design produces noise instead of clarity.
29. Feature Flag System
Turns features on or off without redeploying code. Staged rollouts, A/B tests, canary releases, kill switches. Safer releases, but stale flags accumulate over time and make codebases messier than you'd expect.
30. Data Pipeline
Ingests, transforms, moves, and stores data for downstream use. Analytics, ML features, ETL, audit systems. Enables richer downstream capabilities, but pipeline correctness, latency, backfills, and schema evolution become ongoing maintenance concerns.
Five questions that matter more than definitions
Definitions help, but the real skill is knowing how to choose and combine blocks under constraints. Most good answers to system design questions show you've thought through these:
1. What exact pain is this component solving? Don't say "use Redis for speed." Say what is slow, who is overloaded, and why this layer helps.
2. What new problem does this component introduce? Every useful component creates new failure modes. Caches go stale. Queues create eventual consistency. Replication creates lag. Sharding makes cross-entity queries harder.
3. Is this on the critical path or the async path? That one distinction changes latency requirements, reliability expectations, and user experience.
4. What is the source of truth? A cache, search index, or analytics store may be important, but they are rarely the authoritative source. Know which system owns the data.
5. What breaks first at scale? Different systems hit different walls: database writes, read traffic, connection count, queue lag, search indexing delay, or storage growth. Identifying the first bottleneck is often more useful than drawing a perfect diagram.
How this connects to larger system design
Once these blocks feel familiar, systems like WhatsApp, YouTube, Uber, Stripe, Dropbox, and Twitter stop looking mysterious. They're still complex, but they stop being made of magic.
They become combinations of recurring decisions:
- synchronous vs asynchronous work
- stateful vs stateless components
- latency vs consistency
- cost vs simplicity
- throughput vs correctness
- push vs pull
- centralization vs distribution
That's the real reason to study building blocks before studying whole systems.
What comes next
Each block will eventually get its own page covering:
- a deeper definition and mental model
- the exact problem it solves
- examples from real systems
- common misuse and anti-patterns
- alternatives and when to choose them
- failure modes
- small diagrams
The deeper layer is where the understanding becomes durable enough to use under pressure, whether in an interview or an actual production incident.
The goal of system design isn't to know fashionable infrastructure names. It's to understand why certain components keep reappearing across very different products, and to be able to reason about when they help, when they hurt, and what breaks when you get the choice wrong.
That's what this index is trying to build toward.
I build and scale reliable production systems. Open to full-time and freelance work with U.S.-based teams that value ownership and execution.
Got something in mind?
Book a Discovery Call