How to Approach a System Design Interview
A system design interview is a conversation about trade-offs, not a quiz with a right answer. The interviewer gives you a vague product ("design a URL shortener", "design a news feed") and watches how you turn it into a structured, defensible design. Candidates rarely fail because they do not know a technology. They fail because they start drawing boxes before they know what they are building.
This chapter gives you a repeatable order of work and the arithmetic you need to size a system in two minutes.
1. What is actually being assessed
Interviewers usually score four things, whatever the company calls them:
- Problem framing. Do you ask the questions that change the design, or do you assume?
- Technical breadth. Do you know the standard building blocks and when each applies?
- Depth. Can you take one component, such as the database or the queue, and reason about it down to failure modes?
- Communication and judgement. Do you name the trade-off you are making, and can you change course when the interviewer adds a constraint?
Notice that "arrive at the same architecture as the reference solution" is not on the list. Two strong candidates often produce different designs for the same question.
2. A working sequence for forty-five minutes
A reasonable split for a forty-five minute round:
- Requirements and scope (5 minutes). Functional requirements, non-functional requirements, and what is explicitly out of scope.
- Estimation (5 minutes). Traffic, storage and bandwidth, enough to pick the shape of the system.
- API and data model (5 minutes). The two or three core calls and the entities behind them.
- High-level design (10 minutes). The boxes and the request path for the main use cases.
- Deep dive (15 minutes). One or two components the interviewer cares about.
- Bottlenecks, failures and wrap-up (5 minutes). Single points of failure, scaling limits, what you would build next.
Treat the times as a guide. If the interviewer pulls you into the data model early, follow them. The order exists so that you never design a database before you know the read-to-write ratio.
<!--fig:steps-->3. Requirements: the questions that change the design
Ask questions whose answer would change the architecture. Generic questions waste time.
Functional. What are the two or three things users do most? For a URL shortener: create a short link, redirect to the long link. Analytics, custom aliases and expiry are extras you can confirm are in or out.
Non-functional. These shape the design more than features do:
- Scale: daily active users, requests per second, data volume.
- Read to write ratio: a read-heavy system leans on caches and replicas, a write-heavy one on partitioning and queues.
- Latency target: a redirect should take tens of milliseconds, a report may take seconds.
- Availability versus consistency: can a user briefly see stale data, or must every read be current?
- Durability: is losing a write acceptable, as with a view counter, or catastrophic, as with a payment?
State your assumptions aloud and write them on the board. If you are wrong, the interviewer corrects you early, which is the point.
4. Back-of-the-envelope estimation
Estimation is not about precision. It is about knowing whether the system fits on one machine, needs a cluster, or needs a global footprint.
Time conversions worth memorising. A day has 86,400 seconds, which you can round to about 100,000 for fast arithmetic. A month has roughly 2.5 million seconds, and a year about 31.5 million.
Traffic. Suppose 100 million daily active users each make 10 requests per day. That is requests a day, so an average of requests per second. Peak traffic is commonly two to three times the average, so plan for about 25,000 to 35,000 requests per second.
Storage. Suppose 500 million new records a month, each about 500 bytes. That is 250 GB a month, or 3 TB a year before replication. With three replicas it is about 9 TB a year. Say this out loud, then say whether it fits on a single database server. It does not need a cluster for capacity, but you may still want one for throughput and availability.
Bandwidth. 30,000 requests per second at 2 KB per response is 60 MB per second, or about 480 Mbit per second. That is within a single well-provisioned network link, which tells you the bottleneck lies elsewhere.
Latency intuitions (orders of magnitude, hardware changes them):
| Operation | Rough cost |
|---|---|
| Read from main memory | about 100 nanoseconds |
| Read from a local SSD | tens to a hundred microseconds |
| Round trip inside one data centre | about 0.5 milliseconds |
| Read from a spinning disk (seek) | about 10 milliseconds |
| Round trip between continents | 100 to 200 milliseconds |
The lesson in the table is the gap: memory is thousands of times faster than a network hop, which is why caches exist, and a cross-region call is hundreds of times slower than a local one, which is why data is placed near users.
<!--fig:latency-->5. API and data model
Sketch the interface before the internals. For a URL shortener:
POST /v1/urls { "long_url": "...", "alias": "optional" } -> { "short_url": "..." }
GET /{code} -> 301 or 302 redirect
Then list the entities and how they are accessed: a url record keyed by the short code, looked up by code on every read. The access pattern ("always fetch by key") drives the storage choice far more than the entity shape does.
Mention the details that show experience: idempotency for create calls, pagination for lists, authentication and rate limiting at the edge, and versioning in the path.
6. High-level design
Draw the minimum that serves the main use cases: clients, a load balancer, stateless application servers, a cache, a primary datastore, and any asynchronous worker. Trace one write and one read through the diagram, naming each hop. If you cannot trace a request through your picture, the picture is not a design yet.
Add components only when a requirement or a number forces them. A queue appears because writes must be absorbed during spikes. A cache appears because the read ratio is high. Say the reason each time.
7. Deep dive and trade-offs
The interviewer will probe where the design is weakest or most interesting. Prepare to go deep on:
- the data store: choice, indexing, partitioning key, replication;
- consistency: what a user can observe after a write;
- failure: what happens when a node, a zone or a dependency dies;
- hot spots: one celebrity, one viral link, one tenant;
- cost: whether the design is wasteful for the stated load.
Whenever you pick an option, name the cost: "I am using asynchronous replication, so a failover can lose the last few writes; for this feature that is acceptable". That sentence is worth more than a perfect diagram.
8. Handling a changed constraint
Interviewers add constraints on purpose: "now the traffic is a hundred times larger" or "now it must work across regions". Do not restart. Identify which assumption broke, which component it hits, and what you would change there. The ability to adapt a design is exactly the skill the job needs.
A complete walk-through: applying the framework
Here is the framework on a real prompt, so you can hear what each stage sounds like. The prompt: "Design a service that lets people share text snippets by link."
Stage 1: requirements (about five minutes)
You: "Let me confirm scope. Users paste text and get a link. Anyone with the link can read it. Do snippets expire? Can they be edited or deleted? Are there accounts, or is it anonymous? Is there search?"
Interviewer: "Anonymous, no edits, optional expiry, no search."
You: "Then the core requirements are: create a snippet and receive a link, and read a snippet by link. Non-functionally, reads must be fast and highly available, snippets are small, and I will assume 10 million new snippets a month with 20 reads for each. Durability matters: a snippet must not be lost while it is live."
Notice the pattern: questions that remove ambiguity, a one-line restatement, and numbers stated as assumptions the interviewer can correct.
Stage 2: estimation (about five minutes)
- Writes: per second. Trivial.
- Reads: 20 times that, about 80 per second. Also small.
- Storage: assume 10 KB per snippet, so 100 GB a month and about 1.2 TB a year.
You: "Both the traffic and the data are small. One database server could hold this for years. So I will keep the design simple: a stateless service, a database and a cache for hot snippets, and I will explain where it would break first."
This stage decides the design. A candidate who skips it might add sharding and a message queue to a system that needs neither, and the interviewer will notice.
Stage 3: API and data model (about five minutes)
POST /v1/snippets { "text": "...", "expires_in": 86400 } -> 201 { "url": "https://sn.ip/k9Xz2Qa" }
GET /k9Xz2Qa -> 200 text | 404 | 410 expired
One table keyed by the short code, holding the text (or a pointer to object storage if large), creation time and expiry.
Stage 4: high-level design (about ten minutes)
A client, a load balancer, stateless app servers, a cache, and the database. Writes go to the database. Reads check the cache first. For large snippets, store the text in object storage and keep a reference. I trace one create and one read through the picture, naming each hop.
Stage 5: deep dive (about fifteen minutes)
The interviewer asks: "How do you generate the codes?" This is where depth is shown: collisions, a key generation approach, guessability, and the trade-off of each, as in the URL shortener chapter. If they ask about expiry, you discuss lazy deletion on read plus a background sweep.
Stage 6: wrap-up (about five minutes)
"The weak point is the single database primary, which I would put behind replicas with automatic failover. The first thing to break at ten times the load is the read path, which the cache and replicas absorb. Next I would add monitoring on error rate and latency, and abuse controls on the create endpoint."
That is a complete round: shape, numbers, simple design, depth where asked, honest limits.
Estimation toolkit
You will do arithmetic out loud, so keep a few anchors ready.
Powers of two and ten
| Power | Approximately | Name |
|---|---|---|
| Thousand (KB) | ||
| Million (MB) | ||
| Billion (GB) | ||
| Trillion (TB) | ||
| Quadrillion (PB) |
Time
- 1 day is about seconds (precisely 86,400).
- 1 month is about seconds.
- 1 year is about seconds.
Rules of thumb
- Requests per second is daily requests divided by . A billion requests a day is about 10,000 per second.
- Peak is commonly two to three times the average, and a daily-cycle service has a quiet night and a busy evening.
- A character of text is one byte (ASCII) or up to four (Unicode). A typical row of metadata is hundreds of bytes. A photo is hundreds of KB to a few MB. A minute of compressed video is on the order of several to tens of MB depending on quality.
- A single database server comfortably handles thousands of simple queries per second, and tens of thousands on good hardware with a hot cache. A single cache node handles on the order of operations per second. Treat both as starting points to benchmark.
- Replicate data three times when estimating durable storage.
- Availability: 99.9 percent is about 8.8 hours of downtime a year, 99.99 percent about 53 minutes.
The habit. Round aggressively, write the formula, state the assumption, and finish with what the number means for the design. A wrong assumption that you stated is a small cost, because the interviewer corrects it. A silent assumption is a large one.
Potential deep dives into the process itself
Deep dive 1: How much detail should the high-level design have?
Weak: a diagram of twenty boxes, each labelled with a product name, and no explanation of why any is there. It shows memorised architecture and no judgement.
Solid: the minimum components that serve the stated requirements, each added for a stated reason: "a cache, because reads outnumber writes by twenty to one".
Excellent: the solid design, plus an explicit statement of what you left out and the condition that would make you add it: "I am not adding a queue because writes are four per second. If it were four thousand, I would buffer them."
Deep dive 2: How do you respond when the interviewer changes a constraint?
Weak: restart the design from scratch, or defend the original and ignore the change.
Solid: identify which assumption changed and update the affected component. "At a hundred times the traffic, the single database is the bottleneck, so I would add replicas and a larger cache."
Excellent: the same, and use the change to show the design's evolution: what still works unchanged, what breaks first, what you would do next, and what it would cost. Interviewers add constraints precisely to see this adaptability.
Deep dive 3: How do you handle a topic you do not know well?
Weak: bluff, or go silent.
Solid: say what you know and reason from principles. "I have not used that product, but I know the problem it solves is X, and here is how I would approach it."
Excellent: connect it to concepts you do know, state the trade-off in general terms, and ask whether the interviewer wants you to go deeper. Honesty with reasoning scores well. Confident wrongness does not.
What is expected at each level
Mid-level. You produce a working design for a clear problem, with guidance. You ask basic requirement questions, make reasonable estimates and know the standard building blocks. The interviewer may need to prompt you toward failure cases and scaling.
Senior. You lead the conversation without prompting. You clarify requirements that change the design, use numbers to choose it, go deep on one or two components, and discuss failure, consistency and trade-offs unprompted. You can defend decisions and adapt when challenged.
Staff and above. You question the problem as well as solve it: what the product really needs, what is the simplest thing that works, what the operational and organisational costs are. You discuss evolution over time, cost, team boundaries, risks and how you would validate the design. You often identify the most important decision in the problem within the first few minutes.
Interview questions and model answers
Q: You are asked to design something you have never built. What do you do first? Ask what it is used for, who uses it and how many, then restate the problem in my own words. I pick the two or three core operations and design for those, and I park everything else as explicitly out of scope.
Q: How do you decide between a simple design and a distributed one? By the numbers. If the data and traffic fit comfortably on one primary with a replica, I start there and explain the point at which I would split it. Distributed systems add failure modes, so I add distribution only when a requirement forces it.
Q: What if you realise mid-interview that an earlier choice was wrong? Say so, explain what changed, and fix it. Revising a decision openly reads as good judgement. Defending a mistake reads as the opposite.
Common mistakes
- Designing before asking any question.
- Naming technologies without saying why they fit.
- Skipping estimation, then choosing a design the numbers do not justify.
- Drawing a single database and calling it done, with no mention of failure.
- Talking continuously and never checking the interviewer's interest.
- Treating the design as finished instead of listing its weaknesses.
What a strong opening sounds like
"Before I draw anything, let me confirm scope. I assume a read-heavy service, around a hundred million daily users, redirects under fifty milliseconds, and links that rarely change. I will leave analytics out unless you want it. Roughly that is twelve thousand reads per second on average, so I expect the cache to do most of the work. Does that match what you have in mind?"
It is specific, it states assumptions, it commits to numbers, and it invites correction.