0% completed
Capacity Planning
On This Page
- Average Is the Wrong Number
- Headroom Is Not Waste
- The Growth Multiplier
- Whole Machines, in a Few Sizes
- What Autoscaling Does Not Fix
- The Senior Decision
1. Average Is the Wrong Number
Capacity planning turns an estimate of demand into a count of machines. The first mistake is almost always made in the first step, by planning against average demand.
Systems do not experience averages. They experience peaks, and the gap between the two is wider than most people expect.
- Daily shape. Consumer products typically peak at two to four times their daily average, in a window a few hours wide.
- Weekly shape. Business products often see Monday at double Saturday, which means the daily average across a week understates the day that matters.
- Seasonal shape. Retail at its annual peak can run ten times its ordinary day, and that peak is known months ahead, which makes failing to plan for it a choice.
- Event-driven shape. A product launch, a broadcast, a push notification to every user. These are the sharpest, and the only one of the four that can arrive without warning.
So the input to capacity planning is peak demand, not average demand. And there is a second peak hiding inside the first: the busiest minute inside the busiest hour, which can be several times the hourly rate if something synchronizes arrivals. If your traffic is driven by a scheduled push, your real peak is the minute after it is sent.
State which peak you are planning for. "Four thousand requests per second" means nothing without "at the busiest minute of the busiest day we expect this quarter."
2. Headroom Is Not Waste
The second mistake is sizing the fleet so that peak demand exactly equals capacity. That looks efficient and it does not work, for two independent reasons.
Queueing makes the last ten percent unusable. As utilization rises toward one hundred percent, waiting time does not rise gently. For a simple queue, the time a request spends in the system is its service time divided by one minus utilization.
- At 50 percent utilization, a request takes about twice its service time.
- At 90 percent, about ten times.
- At 95 percent, about twenty times.
Nothing broke between 90 and 95 percent. The arithmetic simply stops being forgiving, which is why a system can look healthy on a utilization dashboard right up to the moment its latency becomes unacceptable. Most teams target somewhere around 70 percent at peak, which keeps the multiplier near three and leaves room to absorb a surprise.
Failure takes capacity away at the worst moment. If you run across three availability zones and lose one, the surviving two must carry all the traffic. Each goes from a third of the load to a half, which is 1.5 times what it was handling.
That gives a hard ceiling. To survive losing one of three zones, steady-state utilization must stay at or below about 66 percent. With two zones, losing one doubles the load on the survivor, so the ceiling is 50 percent. This is why the familiar advice to run at 70 percent is not arbitrary: it is roughly what three zones and a single-zone failure allow.
Headroom is therefore not slack. It is the capacity that the failure you have planned for will consume.
3. The Growth Multiplier
Capacity planning is always for a horizon, and the horizon should be set by how long it takes you to add capacity, not by how far ahead you can imagine.
If new machines take a week to provision, plan a few months out. If adding capacity means a hardware order or a quota increase that takes a quarter, plan a year ahead. Planning further than you can act is a forecasting exercise rather than a capacity exercise.
Then apply growth, and respect that it compounds.
- Ten percent per month is about 3.1 times in a year.
- Twenty percent per month is about 8.9 times in a year.
- Five percent per month is about 1.8 times in a year.
Those numbers surprise people repeatedly, because monthly growth rates sound small and annual multipliers do not. What matters is not the rate but which kind of growth you are facing.
- Growth you can absorb is handled by adding instances behind the same design. Plan it, budget it, and move on.
- Growth that forces a redesign is growth that crosses a threshold from the previous lesson, such as writes outgrowing one primary. That needs a date attached, because the work takes time and the date is when the work must already be finished rather than when it must start.
The useful output is a sentence of this shape: "At current growth we cross the single-primary write limit in about seven months, and partitioning takes a quarter to build, so it starts within three months."
4. Whole Machines, in a Few Sizes
Arithmetic gives you a number of machines with a decimal point. Vendors sell whole machines, in a small set of sizes, and the gap between those two facts is where capacity plans lose accuracy.
- You cannot buy 3.7 machines. You buy four, and four is the honest number. Across a small fleet this rounding is a large fraction of the total.
- Instance sizes are coarse. If the next size up is double, then needing 10 percent more capacity than one size can mean paying for 100 percent more.
- Workloads have more than one dimension. A service that exhausts memory while using a fifth of its CPU does not pack with one that does the reverse. Capacity is not a single number; it is whichever dimension runs out first, and that dimension differs per service.
- Stranded capacity is real. A machine with plenty of CPU left and no memory left is full. Fleet utilization figures that average across dimensions hide this and routinely overstate how much room remains.
So plan capacity in the dimension that binds, name it out loud, and check the other dimensions for stranding. A fleet reported as 60 percent utilized may have no usable space at all.
5. What Autoscaling Does Not Fix
Automatic scaling is often treated as a reason not to plan capacity. It changes the problem rather than removing it, and it has four limits worth knowing before relying on it.
- It reacts more slowly than traffic arrives. Detecting load, starting an instance, warming it and passing health checks takes from tens of seconds to several minutes. A spike that arrives in ten seconds is served entirely by the capacity you already had.
- Something must survive the gap. Whatever the scale-up delay is, your existing fleet handles the whole spike for that long. That sets a floor on steady-state capacity that autoscaling cannot lower.
- It scales what you told it to. Adding application instances when the database is the constraint makes the problem worse, because each new instance opens its own connections and adds load to the thing that was already saturated.
- Scaling down has its own risk. Aggressive scale-down leaves you with the wrong floor when traffic returns, and scale-down during a partial failure removes capacity exactly when the survivors need it.
Autoscaling is genuinely good at following a daily curve and at absorbing moderate, gradual growth. It is poor at absorbing sharp spikes and it does nothing for a bottleneck that is not the thing being scaled.
6. The Senior Decision
A capacity answer is not a number. It is four statements, and leaving any of them out makes the number impossible to check.
- Which peak. "The busiest minute of the busiest day we expect this quarter."
- What headroom, and why. "Sized to 65 percent, so that losing one of three zones keeps us under capacity."
- What horizon, and at what growth. "Six months out at 10 percent monthly, which is 1.8 times current load."
- Which dimension binds. "Memory, at roughly 40 gigabytes per instance; CPU is at a third."
Stated that way, every assumption is visible and arguable, which is exactly what you want. A reviewer can disagree with the growth rate without rejecting the whole plan, and the plan can be revisited by changing one input.
What the interviewer is scoring: "how many servers do you need?" is a question about assumptions, not arithmetic. The weak answer produces a number. The strong answer produces a number with its peak, its headroom, its horizon and its binding resource attached, and it says which assumption the answer is most sensitive to. Two follow-ups are near certain. "What happens if you lose an availability zone?" is checking whether your headroom was chosen or inherited. "Why not just autoscale?" is checking whether you know that scaling takes longer than a spike does, and that the existing fleet carries everything until it finishes.
Flashcards Review
Why is average demand the wrong planning input?
Reading Progress
0%
On This Page
- Average Is the Wrong Number
- Headroom Is Not Waste
- The Growth Multiplier
- Whole Machines, in a Few Sizes
- What Autoscaling Does Not Fix
- The Senior Decision