The latency cliff

Your API is fine at 70% utilization and unusable at 90%, and nothing in between changed. Waiting time scales with the reciprocal of the free capacity.

Traffic goes up 15% and response times go up 400%. Nothing was deployed. No query got slower. The dashboards show CPU at 85% instead of 70%, which doesn't look like an emergency, and yet the p99 has gone from 200ms to two seconds.

The cause is that queueing does not scale linearly with load. Waiting time is governed by how much spare capacity remains, and spare capacity runs out much faster than utilization suggests.

Response time against utilization

50%utilization
2.0xslower than idle

Utilization is how much of the pool's capacity is in use. The curve is the same shape for CPUs, connection pools, and checkout queues.

For a single queue, response time is service time divided by the fraction of capacity still free. At 50% utilization requests take twice as long as they would on an idle system. At 90% they take ten times as long. At 95%, twenty.

The dangerous property is that the curve is nearly flat for the entire range you spend most of your time in. Going from 20% to 50% barely moves it. Going from 85% to 92% moves it enormously, and it's the same seven points.

Where the queue is

The formula describes any resource with a fixed number of servers and a line of work waiting for them. In a backend that's usually not the CPU.

It's the connection pool. Twenty connections, each request holding one for the duration of a query. It's the worker pool, the thread pool, the number of concurrent requests your runtime will handle, the rate limit on a downstream API. Anything where work waits for a slot.

A pool under load

pool
waiting: 0
0.80utilization
0msp50
0msp99
0dropped

Raise the arrival rate slowly. Nothing happens, then everything happens at once.

Raise the arrival rate one step at a time. The queue stays empty, then stays empty, then doesn't. The transition takes one adjustment, and it's the same size adjustment as all the ones before it.

Utilization here is arrivals times service time divided by pool size. Twenty requests per second against a 200ms query needs 4 concurrent slots to break even. A pool of 5 puts you at 0.8 and a pool of 4 puts you at 1.0, where the queue grows without limit and latency is bounded only by your timeout.

Little's Law

The relationship that ties this together is L = λW. The average number of items in a system equals the arrival rate times the average time each one spends there.

It's useful because you usually know two of the three. If you're handling 500 requests per second and each takes 200ms, you have 100 requests in flight on average. That number is what your pool has to accommodate, and if the pool is 50 then half of those requests are waiting rather than working.

Run it backwards to size things. A pool of 20 with a 50ms service time supports 400 requests per second at full saturation, and full saturation is exactly where you don't want to be. Target 60 to 70% and the same pool gives you 240 to 280.

Why variance makes it worse

The curve above assumes randomly distributed arrivals and service times. Real traffic is worse than random on both counts.

Arrivals cluster. Users act on the same triggers, cron jobs fire on the minute, a retry storm arrives as one burst. A system provisioned for the mean is running well above the mean for a meaningful fraction of every minute.

Service times have long tails. If 99% of your queries take 10ms and 1% take 500ms, the mean is 15ms but the slow ones occupy a slot fifty times longer, and while they do it the pool is effectively smaller. One slow query type can push a healthy pool over the edge without changing the average much at all.

This is why a single unindexed query in a rarely-hit endpoint can degrade everything else on the same pool. The cost is the connection it occupies, not the CPU it burns.

What to do at the cliff

Adding capacity moves the cliff, it doesn't remove it. Doubling the pool halves utilization once, and traffic growth walks it back.

The things that actually help are the ones that keep the queue short:

  • Bound the queue. An unbounded queue converts a capacity problem into a latency problem, and a request that waits 30 seconds for a slot has usually been abandoned by the client already. Rejecting work fast is better than accepting work you can't finish.
  • Time out at every layer, with budgets that shrink going down. A caller with a 5s timeout calling a service with a 10s timeout means the downstream work continues, holding a slot, for a response nobody will read.
  • Shed load by priority rather than uniformly. Health checks and paying customers should not queue behind bulk exports.
  • Separate the pools. Slow endpoints on their own pool cannot exhaust the one serving fast ones. This is bulkheading, and it's the cheapest structural fix on the list.

The metric worth alerting on is utilization, not latency. Latency tells you that you're already over the edge. Utilization crossing 0.7 tells you roughly how long you have.