Resource Requests Come Before Autoscaling

More from the blog
October 1, 2026

Autoscaling is the most requested Kubernetes feature and the last one worth adding. A HorizontalPodAutoscaler multiplies capacity. It does not tell the scheduler where that capacity belongs, and it cannot repair a cluster that never had a plan for its workloads in the first place.

Before you wire a single replica count to CPU, you need resource requests. Scheduling, eviction, autoscaling and cost control all sit on top of them.

What a request actually does

A request is a promise to the scheduler. When you set requests.cpu and requests.memory on a container, you tell kube-scheduler how much capacity a node must have free before that Pod can land there. A limit is a different thing entirely: a runtime ceiling enforced by the kubelet and cgroups.

Without requests, the scheduler treats every Pod as free. It will stack twenty of them onto one node and leave the rest of the cluster idle. This is the single most common cause of "one node is out of memory while most nodes are empty".

Requests, limits and QoS

  • requests.cpu reserves capacity for scheduling, and is the denominator for HPA CPU utilisation percentages.
  • requests.memory reserves memory and largely determines the Pod's QoS class.
  • limits.cpu throttles the container when exceeded; it does not kill it.
  • limits.memory exceeded means OOMKilled, restart, repeat.
  • QoS class is Guaranteed when requests equal limits, Burstable in between, and BestEffort when nothing is set. BestEffort Pods are evicted first under pressure.

Why autoscaling without requests misbehaves

An HPA targeting CPU utilisation computes a percentage of the request, not of the node and not of the limit. Remove the CPU request and the percentage has no denominator, so the HPA either emits no usable metric or flails between one replica and the maximum with no change in traffic.

A Vertical Pod Autoscaler is not a shortcut either. In recommendation mode it reports the requests it would apply, which makes it a reasonable way to discover numbers nobody ever measured — but applying those numbers automatically on a live cluster without review is how a stable workload gets restarted for no good reason.

The Cluster Autoscaler has the same blind spot from the other direction. It estimates node utilisation by summing the requests of the Pods running there. Pods with no requests look like free tenants, so it never scales down and the bill grows while every dashboard stays green.

Autoscaling is arithmetic performed on requests. If the requests are wrong, the scaling is wrong — only faster.

Getting the numbers right

Do not guess and do not copy from a tutorial. Run the workload, then measure it. kubectl top pods gives you a live snapshot; a metrics stack such as Prometheus gives you the p95 over a week, and that is the number that matters. Set memory requests near the p95 working set rather than the average, and CPU requests around sustained usage with limits left open enough for genuine bursts.

Then enforce it. A ResourceQuota per namespace plus a LimitRange with default requests means no manifest can ship without them. This also removes the noisy-neighbour incidents that are impossible to attribute after the fact.

Where requests stop being enough

Requests solve placement, not elasticity. Once coverage is complete, an HPA can scale replicas on real metrics, and the Cluster Autoscaler or Karpenter can add nodes when Pods sit Pending because no single node satisfies them. Note the ordering: node autoscaling reacts to unschedulable Pods, so it only functions when those Pods declare what they need.

What to check first

Run kubectl describe node and read the allocated resources table. If requests sit near zero while usage is high, you have found the problem. Then list containers with no request at all and count them. That count is your backlog. Sort namespaces by request coverage and start with whichever one carries the most traffic.

How we approach it at Weeltec

Cluster reviews almost always start with a request-coverage audit, because it is cheap to measure and it explains most of the incidents teams bring to us. We profile real workloads over time, set requests and limits from evidence rather than habit, wire in quotas so the settings survive the next deploy, and only then switch on autoscaling.

The rule

Requests first. Quotas second. Metrics third. Autoscaling fourth. Skip a step and every later step inherits the damage.

We run Kubernetes platforms for teams that need them to behave predictably. If your cluster schedules badly or your autoscaling numbers make no sense, get a quote and we will start with the measurements.