System Design Patterns: From Fundamentals to Real Systems
Vote

0% completed

​

Back-of-Envelope Math

  1. Why Engineers Estimate
  1. Numbers to Memorize
  1. Calculation 1: Rate x Duration = Work in Progress
  1. Calculation 2: Availability Drops Along a Chain
  1. Calculation 3: Cache Hit Ratios
  1. Calculation 4: Sizing
  1. Example: Sizing the URL Shortener
  1. Using Estimation in Interviews
  1. TL;DR

1. Why Engineers Estimate

Back-of-the-envelope calculations provide a useful approximation rather than a precise forecast. They help you find large mistakes early: whether one machine can handle the load, whether 600 or 60,000 requests are active at once, and whether a table fits in memory. A rough answer in thirty seconds can change the design, while a precise answer produced next week cannot guide the current decision.

The five questions need numbers before their answers become useful, and this lesson provides the required calculations. Estimation is also a standard part of system design interviews. When an interviewer asks how much storage a design requires, you should state reasonable assumptions, round the values, and explain each step.

Round aggressively. A day contains 86,400 seconds, but you can use 100,000 for a quick estimate. A month contains about 2.5 million seconds, but 3 million is easier to calculate with. The goal is to estimate orders of magnitude, meaning powers of ten, rather than exact decimal values. When the direction matters, overestimate demand and underestimate capacity so the design keeps a safety margin.

2. Numbers to Memorize

Most latency estimates use the following values. Each row is roughly 10 to 1,000 times slower than the row above it:

OperationTime
Read from memory (RAM)~100 nanoseconds
Read from a fast SSD~100 microseconds
Network call inside a datacenter (for example, a cache read)~0.5-1 ms
A simple database query~5-15 ms
Disk seek on a spinning disk, or a slow query~10-100 ms
Network round trip between regions~100 ms
What users perceive as "instant"up to ~100-200 ms

Three facts from this table influence many designs. First, a cache hit takes about 1 ms, while a database query takes about 10 to 15 ms, which explains why caching can reduce read latency. Second, sequential disk writes are much faster than writes scattered across the disk, which supports the write-ahead log pattern. Third, every sequential round trip between regions adds about 100 ms, so several such calls can exceed the application's response-time limit.

Size estimates require only a few reference values. A character uses about one byte, an identifier uses about 16 to 36 bytes, and a typical database row uses about 1 KB. An image commonly uses 100 KB to 1 MB. Each unit is one thousand times the previous unit: KB, MB, GB, and TB. Therefore, one million rows of 1 KB each require about one gigabyte, while one billion such rows require about one terabyte.

The latency numbers to memorize
The latency numbers to memorize

3. Calculation 1: Rate x Duration = Work in Progress

To estimate how many operations are active at the same time, multiply their arrival rate by the time each operation requires:

<center> things in flight = arrival rate x duration </center>

Suppose 300 requests arrive each second and each request takes 2 seconds to complete. The service has 300 x 2 = 600 requests active at any moment, and each request may hold a thread or connection. If the connection pool, which is a fixed set of open database connections, has 200 slots, the service does not have enough connections for this load.

The same values can estimate how long a queue needs to process its backlog. If a queue contains 552,000 messages and its consumers process 400 messages per second, it needs 552,000 / 400 = about 1,400 seconds, or 23 minutes.

This is the most frequently used calculation in the course. It helps size thread pools, connection pools, bulkhead limits, and queue processing time. Its formal name is Little's law, which is also useful terminology in an interview.

4. Calculation 2: Availability Drops Along a Chain

When a request depends on several parts, multiply their availability values to calculate the availability of the complete request. The result is always lower than the availability of the least reliable part:

<center> chain availability = availability of each part, multiplied together </center>

Five sequential dependencies with 99.9% availability each produce 0.999^5, or about 99.5% availability for the request. For small failure rates, you can use a simpler shortcut by adding those rates: five parts x 0.1% failure gives about 0.5% total failure.

Percentages are easier to understand when converted into time:

AvailabilityDowntime per yearPer month
99%~3.7 days~7 hours
99.9%~9 hours~43 minutes
99.99%~53 minutes~4 minutes
99.999%~5 minutes~26 seconds

Two design conclusions follow from this calculation. Each additional nine in an availability target requires roughly ten times the engineering effort of the previous one, while each required dependency can only reduce total availability. The practical response is not to make every part perfect but to reduce the number of dependencies required for a successful result.

The graceful degradation lesson develops this idea by showing how a system can return a smaller useful result when an optional dependency fails.

5. Calculation 3: Cache Hit Ratios

For any cache, two values follow from the hit ratio h, which is the share of requests answered by the cache:

<center> average latency = h x (fast time) + (1 - h) x (fast time + slow time) </center> <center> load reaching the database = (1 - h) x request rate </center>

Consider 12,000 requests per second, with 1 ms for a cache hit and 15 ms for a database read.

  • At a 99% hit ratio, average latency is about 1.15 ms and the database receives 1% x 12,000 = 120 queries per second.
  • At a 90% hit ratio, average latency is about 2.5 ms, but the database receives 1,200 queries per second, which is ten times more load.

💡 A change from a 99% hit ratio to 90% does not increase database load by nine percent. It increases the miss rate from 1% to 10%, so the database receives ten times as many queries. Evaluate a cache by its misses as well as its hits.

This calculation also explains why a cache stampede can overload a database. A stampede occurs when many requests miss the same cache entry at once. When a popular entry expires, its hit ratio briefly approaches zero, and most requests reach the database.

6. Calculation 4: Sizing

Requirements often describe daily traffic, while capacity plans require per-second traffic. Divide the daily value by 100,000 for a quick conversion:

<center> per second = per day / 100,000 </center>

For example, 10 million orders per day is about 100 orders per second on average. Traffic is not constant, so plan for a peak between three and five times the average, or 300 to 500 orders per second in this example.

Storage is the number of records multiplied by the size of each record and the retention period:

<center> 10 million orders/day x 1 KB x 365 days = ~3.7 TB per year </center>

This amount fits on one ordinary database machine. The estimate therefore prevents unnecessary sharding, which means splitting the rows across several machines.

The four calculations
The four calculations

7. Example: Sizing the URL Shortener

The previous lesson outlined a URL shortener. You can now estimate its capacity while explaining each assumption:

  • "Assume 100 million redirects per day. Dividing by 100,000 gives 1,000 redirects per second on average, so we should support about 4,000 per second at peak."
  • "Each mapping contains two URLs and metadata, which we can estimate as 1 KB. With 500 million links, the database stores about 500 GB. One database has sufficient capacity today, so sharding can wait until the data grows."
  • "Reads outnumber writes by about 100 to 1, and a small share of links receives most of the traffic. Caching the most popular 1% stores 5 million rows, or about 5 GB. With a 99% hit ratio, only about 10 redirect queries per second reach the database."
  • "At peak, 4,000 requests per second x 2 ms gives 8 requests active at once. This load requires one capable database, a replica, a cache, and explicit failure handling."

This explanation requires about sixty seconds and uses five numbers. Every major part of the design now has a capacity reason, and the design does not include parts for demand that has not appeared.

8. Using Estimation in Interviews

  1. State assumptions before calculating. Say, "Assuming 1 KB per record and 10 million records per day," before using those values. The interviewer can correct an assumption you state, but cannot correct one you leave unstated.
  2. Round before multiplying, and check the units at the end. A per-second rate multiplied by seconds produces a count, while bytes multiplied by a record count produces storage. If the units do not match the expected result, the calculation is wrong.
  3. Use the numbers to reject unnecessary parts. A useful interview statement is, "At this scale, we do not need that yet." This is the cost awareness from the first lesson, supported by arithmetic.

9. TL;DR

The ruleRound aggressively and estimate orders of magnitude. A day is 100,000 seconds. A million 1 KB rows is a GB. Peak is 3-5x average.
The tableMemory ~100 ns, SSD ~100 us, in-datacenter call ~1 ms, DB query ~10 ms, cross-region round trip ~100 ms, "instant" ends at ~100-200 ms.
The four calculationsIn flight = rate x duration. Chain availability = multiply availabilities (or add failure rates). Cache: average latency and DB load both follow from the hit ratio. Sizing: per day / 100,000 = per second; rows x bytes x retention = storage.
The traps90% is not close to 99% (it is 10x the misses). Every extra nine costs ~10x. Dependency chains can only lower the total, so shorten them.
In interviewsState assumptions, check units, and let arithmetic rule out over-engineering. "We do not need that yet" is a strong answer.

Reading Progress

0%


Vote for new content

On This Page

  1. Why Engineers Estimate
  1. Numbers to Memorize
  1. Calculation 1: Rate x Duration = Work in Progress
  1. Calculation 2: Availability Drops Along a Chain
  1. Calculation 3: Cache Hit Ratios
  1. Calculation 4: Sizing
  1. Example: Sizing the URL Shortener
  1. Using Estimation in Interviews
  1. TL;DR