0% completed
The Five Questions
On This Page
- Why You Need a Fixed Procedure
- Question 1: How Does Data Move?
- Question 2: Where Does Data Live?
- Question 3: How Is Data Served Fast?
- Question 4: What Happens When Something Fails?
- Question 5: How Does the System Grow?
- Example: The Five Questions on a URL Shortener
- How to Use the Five Questions
- TL;DR
1. Why You Need a Fixed Procedure
The interviewer says, "Design a URL shortener," gives you the marker, and waits for you to begin.
The difficult part is often not missing technical knowledge but lacking a procedure. Organized candidates use the same questions for every design, which gives them a clear next step when the system is unfamiliar.
This course uses five questions in a fixed order. Each question reveals a set of problems, and each problem leads to patterns taught later in the course. The previous lesson explained how to reason with patterns, while these questions tell you where to search for them.
Get the numbers first. Before asking the five questions, clarify users, requests per second, data size, read-to-write ratio, and cost of downtime. A design for 100 users differs from one for 100 million users. The next lesson explains how to estimate these values quickly.
2. Question 1: How Does Data Move?
Ask who communicates with whom and whether the caller must wait for an answer. Then determine whether the work should happen once, whether several services need the same fact, whether the call crosses a company boundary, and whether data must flow continuously.
This question determines the connections in your diagram, and each connection requires a communication decision. Start with request-response, in which a caller sends a request and waits for a response. Use a message queue when the work can wait and should be stored until a consumer processes it. Use pub/sub when several subscribers need the same event. Use webhooks, which are HTTP calls from one company's system to another, when the call crosses a company boundary. Use streams when the server must send data continuously.
The Moving Data module explains each option and its costs.
Example answer: "Order confirmation must be immediate, so it uses a direct call. Fulfillment can start later, so it uses a queue. Five teams need each order event, so we publish that event to all five."
3. Question 2: Where Does Data Live?
Ask which store contains the authoritative version of the data, how large the data will become, and how quickly it will grow. You also need the read-to-write ratio and whether the system must preserve history or only the current state.
This question places the databases in your design. Add replicas, which are copies of a database, when reads exceed the capacity of one machine. Use sharding, which splits rows across several machines, when writes or total data exceed that capacity. Use consistent hashing when the number of shards must change without moving most of the data.
When history matters, consider the log-based patterns: a write-ahead log records a change before applying it, event sourcing stores the sequence of events, and CQRS uses separate models for writes and reads. The Storing Data module explains these patterns.
Example answer: "Order history matters because customers may dispute a charge, so we store the events instead of keeping only the current state."
4. Question 3: How Is Data Served Fast?
Ask which reads repeat and how long each type of data may remain stale, meaning out of date. Then estimate the hit ratio, which is the share of reads answered by the cache, and decide what happens when a popular cache entry expires.
This question determines what the system caches and how those cached copies remain correct. Start with cache-aside, in which the application reads the database after a cache miss and then stores the result in the cache. Use the write-based variants when data changes frequently, and use stampede prevention when many requests may miss the same popular entry at once.
The caching module follows one central decision: determine how long each type of data may remain stale, then choose a mechanism that enforces that limit.
Example answer: "Product descriptions may be one hour old. Prices may be several seconds old if an update invalidates the cache. The checkout price must be exact, so checkout always reads it from the database."
5. Question 4: What Happens When Something Fails?
For every connection from Question 1, ask what happens when the other service does not respond, fails once, or remains unavailable for ten minutes. You must also determine whether a retry can repeat an action and what the user will see when the operation cannot complete.
This question turns a functional design into a reliable one. Use timeouts to end calls that do not return. Use retries with backoff to repeat a failed call after progressively longer waits. Use idempotency so a repeated request cannot repeat an effect such as charging a customer. Use circuit breakers to stop calls during a sustained outage, and use bulkheads to keep one failing feature from consuming all shared resources. Finally, use graceful degradation to provide a smaller result instead of an error.
The Surviving Failure module explains how these patterns work together.
Example answer: "If the payment provider becomes slow, we end the call after twice its normal worst-case latency. Retries use idempotency keys so they cannot charge the customer twice. After sustained failures, the circuit breaker stops new calls, while the product saves the cart and emails the customer instead."
6. Question 5: How Does the System Grow?
Ask what would reach its limit first if traffic increased by ten times: read capacity, write capacity, storage, or one unusually popular key. Then identify what can scale by adding machines and what requires a design change.
This question tests the capacity limits created by your previous decisions. The Growing Under Load module covers load balancing, automatic scaling, and horizontal scaling, which means adding machines to increase capacity. It also explains which limits cannot be removed by adding more machines.
Example answer: "Replicas and caching increase read capacity, but writes remain limited by the single primary. Sharding by customer ID is the next change, and we should plan it before the primary reaches its limit."
The remaining modules add consistency between machines, traffic management at the system's entry point, and production operations. The data and AI modules then apply the same five questions to data pipelines and machine learning systems.
7. Example: The Five Questions on a URL Shortener
The following example applies the complete procedure to the opening question.
- How does data move? There is one write path: submit a long URL and receive a short one. There is also a read-heavy redirect path, with about 100 reads for each write. Both use direct calls. Click analytics is optional work that does not delay the redirect, so it goes to a queue.
- Where does data live? Each row contains two URLs and a counter, so the rows are small, but the system may store billions of them. One primary database with replicas is sufficient until the data becomes much larger. At billions of rows, shard by the short code because every redirect query already contains it. The system does not need a history of changes, so simple tables are sufficient.
- How is it served fast? A redirect returns the same destination because the mapping does not change. This makes cache-aside suitable, with a long TTL, or time-to-live, which controls how long an entry stays in the cache. A small share of links receives most clicks, so the most popular links require stampede prevention and should not expire during normal operation.
- What happens when it fails? The redirect path has a strict availability target, so a cache miss reads from a replica. Creating a link may return a clear "try again shortly" message because a brief failure causes less harm on that path. The two paths therefore use different reliability targets.
- How does it grow? Reads scale by adding replicas and cache nodes, while writes remain light. Key generation is the main growth concern. Random codes require collision checks, while a counter-based scheme requires coordination so two servers cannot create the same code. You can identify this as future work and explain both options.
This explanation takes about three minutes. The resulting design has separate read and write paths, a storage plan with a shard key, a caching plan, distinct reliability targets, and one identified growth risk. If a requirement changes, you can apply the five questions again to determine which decisions must change.
8. How to Use the Five Questions
💡 In an interview, state the procedure as your agenda: "I will work through data movement, storage, caching, failure handling, and growth in that order." This tells the interviewer how you will organize the discussion and confirms that you plan to cover each major concern.
Use the questions as a checklist. If you cannot decide what to discuss next, or you finish much earlier than expected, review the five questions. Question 4 is commonly missed and often produces the most follow-up discussion.
Use the questions during design reviews. Apply them to designs written by other engineers as well as your own. Asking what happens when the other service does not return is one of the most useful reliability checks in a distributed system.
9. TL;DR
| The procedure | Numbers first, then five fixed questions: How does data move? Where does it live? How is it served fast? What happens when it fails? How does it grow? |
| Why it works | Each question reveals problems, each problem leads to a pattern, and the fixed order guarantees you cover everything. |
| The structure | The five questions match the next five modules of this course, in order. Later modules refine the answers. |
| The most-skipped question | Question 4, failure handling. It is where designs become reliable, and where interviewers spend their follow-ups. Ask it about every arrow in your diagram. |
| Use it as | Your spoken agenda in interviews, your checklist when stuck, and your review question at work. |
Reading Progress
0%
On This Page
- Why You Need a Fixed Procedure
- Question 1: How Does Data Move?
- Question 2: Where Does Data Live?
- Question 3: How Is Data Served Fast?
- Question 4: What Happens When Something Fails?
- Question 5: How Does the System Grow?
- Example: The Five Questions on a URL Shortener
- How to Use the Five Questions
- TL;DR