{"id":57,"date":"2026-10-01T15:51:39","date_gmt":"2026-10-01T15:51:39","guid":{"rendered":"https:\/\/pomax-v3.weeltec.com\/?p=57"},"modified":"2026-10-01T15:53:06","modified_gmt":"2026-10-01T15:53:06","slug":"resource-requests-before-autoscaling","status":"publish","type":"post","link":"https:\/\/pomax-v3.weeltec.com\/?p=57","title":{"rendered":"Resource Requests Come Before Autoscaling"},"content":{"rendered":"<div class=\"wt-post\">\n<style>.wt-post { --wt-bg: #0A0C0F; --wt-bg-2: #0E1116; --wt-surface: #12161C; --wt-line: #1F252E; --wt-line-strong: #2B333E; --wt-text: #E9ECEF; --wt-muted: #98A2AD; --wt-accent: #B9F24D; --wt-accent-hover: #C8FF5E; --wt-accent-dim: rgba(185, 242, 77, 0.10); --wt-accent-line: rgba(185, 242, 77, 0.25); --wt-danger: #F07A6B; --wt-ok: #7ED99A; --wt-sans: ui-sans-serif, system-ui, -apple-system, \"Segoe UI\", Roboto, \"Helvetica Neue\", Arial, sans-serif; --wt-mono: ui-monospace, \"Cascadia Code\", \"JetBrains Mono\", \"SF Mono\", Menlo, Consolas, monospace; --wt-radius: 4px; --wt-h1: var(--wt-text); --wt-h2: var(--wt-text); --wt-h3: var(--wt-text); box-sizing: border-box; background: var(--wt-bg); color: var(--wt-text); font-family: var(--wt-sans); font-size: 1rem; line-height: 1.7; padding: clamp(1.75rem, 4vw, 3rem); border: 1px solid var(--wt-line); border-radius: 0; -webkit-font-smoothing: antialiased; text-rendering: optimizeLegibility; } .wt-post *, .wt-post *::before, .wt-post *::after { box-sizing: border-box; } .wt-post ::selection { background: var(--wt-accent); color: var(--wt-bg); } .wt-post.wt-post p { color: var(--wt-text); font-family: var(--wt-sans); font-size: 1rem; line-height: 1.7; margin: 0 0 1.15rem; max-width: 68ch; } .wt-post.wt-post p:last-child { margin-bottom: 0; } .wt-post.wt-post h1, .wt-post.wt-post h2, .wt-post.wt-post h3, .wt-post.wt-post h4, .wt-post.wt-post h5, .wt-post.wt-post h6 { font-family: var(--wt-sans); font-weight: 700; letter-spacing: -0.025em; line-height: 1.15; text-wrap: balance; } .wt-post.wt-post h2 { color: var(--wt-h2); font-size: clamp(1.45rem, 3vw, 2rem); margin: 2.4rem 0 0.9rem; display: flex; align-items: baseline; gap: 0.6rem; } .wt-post.wt-post h2::before { content: \"\"; flex: none; width: 8px; height: 8px; background: var(--wt-accent); transform: translateY(-2px); } .wt-post.wt-post h3 { color: var(--wt-h3); font-size: 1.15rem; font-weight: 650; margin: 1.8rem 0 0.7rem; padding-left: 0.85rem; border-left: 2px solid var(--wt-accent-line); } .wt-post.wt-post h2:first-child, .wt-post.wt-post h3:first-child { margin-top: 0; } .wt-post.wt-post strong { color: #FFFFFF; font-weight: 650; } .wt-post.wt-post em { color: var(--wt-muted); font-style: italic; } .wt-post.wt-post a { color: var(--wt-accent); text-decoration: none; border-bottom: 1px solid var(--wt-accent-line); transition: color 0.15s ease, border-color 0.15s ease; } .wt-post.wt-post a:hover { color: var(--wt-accent-hover); border-bottom-color: var(--wt-accent-hover); } .wt-post.wt-post ul, .wt-post.wt-post ol { margin: 0 0 1.3rem; padding: 0; list-style: none; max-width: 68ch; } .wt-post.wt-post li { position: relative; padding-left: 1.6rem; margin-bottom: 0.55rem; color: var(--wt-text); line-height: 1.65; } .wt-post.wt-post ul > li::before { content: \"\\25AE\"; color: var(--wt-accent); position: absolute; left: 0; top: 0; font-size: 0.85em; line-height: 1.65; } .wt-post.wt-post ol { counter-reset: wt-li; } .wt-post.wt-post ol > li { counter-increment: wt-li; } .wt-post.wt-post ol > li::before { content: counter(wt-li) \".\"; font-family: var(--wt-mono); font-size: 0.8em; color: var(--wt-accent); position: absolute; left: 0; top: 0; line-height: 1.9; } .wt-post.wt-post blockquote { margin: 1.8rem 0; padding: 1.1rem 1.4rem; background: var(--wt-accent-dim); border-left: 2px solid var(--wt-accent); border-radius: 0; color: var(--wt-text); font-size: 1.05rem; font-style: normal; line-height: 1.6; } .wt-post.wt-post blockquote p { margin: 0; color: var(--wt-text); font-style: normal; } .wt-post.wt-post blockquote::before { content: none; } .wt-post.wt-post code, .wt-post.wt-post kbd, .wt-post.wt-post pre { font-family: var(--wt-mono); font-size: 0.88em; } .wt-post.wt-post code { background: var(--wt-surface); border: 1px solid var(--wt-line); border-radius: var(--wt-radius); padding: 0.1em 0.4em; color: var(--wt-accent); } .wt-post.wt-post pre { background: #0C0F13; border: 1px solid var(--wt-line-strong); border-radius: var(--wt-radius); padding: 1.1rem 1.25rem; overflow-x: auto; color: var(--wt-text); line-height: 1.7; margin: 0 0 1.3rem; } .wt-post.wt-post pre code { background: none; border: 0; padding: 0; color: inherit; } .wt-post.wt-post hr { border: 0; border-top: 1px solid var(--wt-line); margin: 2.2rem 0; } .wt-post.wt-post img { max-width: 100%; height: auto; border-radius: var(--wt-radius); border: 1px solid var(--wt-line); } @media (max-width: 640px) { .wt-post.wt-post h2 { font-size: 1.35rem; } .wt-post.wt-post blockquote { padding: 0.9rem 1.1rem; } } @media (prefers-reduced-motion: reduce) { .wt-post.wt-post a { transition: none; } }<\/style>\n<p>Autoscaling is the most requested Kubernetes feature and the last one worth adding. A HorizontalPodAutoscaler multiplies capacity. It does not tell the scheduler where that capacity belongs, and it cannot repair a cluster that never had a plan for its workloads in the first place.<\/p>\n<p>Before you wire a single replica count to CPU, you need resource requests. Scheduling, eviction, autoscaling and cost control all sit on top of them.<\/p>\n<h2>What a request actually does<\/h2>\n<p>A request is a promise to the scheduler. When you set <strong>requests.cpu<\/strong> and <strong>requests.memory<\/strong> on a container, you tell kube-scheduler how much capacity a node must have free before that Pod can land there. A limit is a different thing entirely: a runtime ceiling enforced by the kubelet and cgroups.<\/p>\n<p>Without requests, the scheduler treats every Pod as free. It will stack twenty of them onto one node and leave the rest of the cluster idle. This is the single most common cause of \"one node is out of memory while most nodes are empty\".<\/p>\n<h3>Requests, limits and QoS<\/h3>\n<ul>\n<li><strong>requests.cpu<\/strong> reserves capacity for scheduling, and is the denominator for HPA CPU utilisation percentages.<\/li>\n<li><strong>requests.memory<\/strong> reserves memory and largely determines the Pod's QoS class.<\/li>\n<li><strong>limits.cpu<\/strong> throttles the container when exceeded; it does not kill it.<\/li>\n<li><strong>limits.memory<\/strong> exceeded means OOMKilled, restart, repeat.<\/li>\n<li><strong>QoS class<\/strong> is Guaranteed when requests equal limits, Burstable in between, and BestEffort when nothing is set. BestEffort Pods are evicted first under pressure.<\/li>\n<\/ul>\n<h2>Why autoscaling without requests misbehaves<\/h2>\n<p>An HPA targeting CPU utilisation computes a percentage of the request, not of the node and not of the limit. Remove the CPU request and the percentage has no denominator, so the HPA either emits no usable metric or flails between one replica and the maximum with no change in traffic.<\/p>\n<p>A Vertical Pod Autoscaler is not a shortcut either. In recommendation mode it reports the requests it would apply, which makes it a reasonable way to discover numbers nobody ever measured \u2014 but applying those numbers automatically on a live cluster without review is how a stable workload gets restarted for no good reason.<\/p>\n<p>The Cluster Autoscaler has the same blind spot from the other direction. It estimates node utilisation by summing the requests of the Pods running there. Pods with no requests look like free tenants, so it never scales down and the bill grows while every dashboard stays green.<\/p>\n<blockquote>\n<p>Autoscaling is arithmetic performed on requests. If the requests are wrong, the scaling is wrong \u2014 only faster.<\/p>\n<\/blockquote>\n<h3>Getting the numbers right<\/h3>\n<p>Do not guess and do not copy from a tutorial. Run the workload, then measure it. <strong>kubectl top pods<\/strong> gives you a live snapshot; a metrics stack such as Prometheus gives you the p95 over a week, and that is the number that matters. Set memory requests near the p95 working set rather than the average, and CPU requests around sustained usage with limits left open enough for genuine bursts.<\/p>\n<p>Then enforce it. A <strong>ResourceQuota<\/strong> per namespace plus a <strong>LimitRange<\/strong> with default requests means no manifest can ship without them. This also removes the noisy-neighbour incidents that are impossible to attribute after the fact.<\/p>\n<h2>Where requests stop being enough<\/h2>\n<p>Requests solve placement, not elasticity. Once coverage is complete, an HPA can scale replicas on real metrics, and the Cluster Autoscaler or Karpenter can add nodes when Pods sit Pending because no single node satisfies them. Note the ordering: node autoscaling reacts to unschedulable Pods, so it only functions when those Pods declare what they need.<\/p>\n<h3>What to check first<\/h3>\n<p>Run <strong>kubectl describe node<\/strong> and read the allocated resources table. If requests sit near zero while usage is high, you have found the problem. Then list containers with no request at all and count them. That count is your backlog. Sort namespaces by request coverage and start with whichever one carries the most traffic.<\/p>\n<h2>How we approach it at Weeltec<\/h2>\n<p>Cluster reviews almost always start with a request-coverage audit, because it is cheap to measure and it explains most of the incidents teams bring to us. We profile real workloads over time, set requests and limits from evidence rather than habit, wire in quotas so the settings survive the next deploy, and only then switch on autoscaling.<\/p>\n<h2>The rule<\/h2>\n<p>Requests first. Quotas second. Metrics third. Autoscaling fourth. Skip a step and every later step inherits the damage.<\/p>\n<p><em>We run Kubernetes platforms for teams that need them to behave predictably. If your cluster schedules badly or your autoscaling numbers make no sense, <a href=\"https:\/\/weeltec.com\/#contact\">get a quote<\/a> and we will start with the measurements.<\/em><\/p>\n<\/div>\n","protected":false},"excerpt":{"rendered":"<p>Autoscaling only multiplies whatever your requests already said. Fix the requests first, or every later step inherits the same scheduling damage.<\/p>\n","protected":false},"author":1,"featured_media":54,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[7],"tags":[32,13,31,33],"class_list":["post-57","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-kubernetes","tag-autoscaling","tag-kubernetes","tag-resource-requests","tag-scheduling"],"_links":{"self":[{"href":"https:\/\/pomax-v3.weeltec.com\/index.php?rest_route=\/wp\/v2\/posts\/57","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/pomax-v3.weeltec.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/pomax-v3.weeltec.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/pomax-v3.weeltec.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/pomax-v3.weeltec.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=57"}],"version-history":[{"count":3,"href":"https:\/\/pomax-v3.weeltec.com\/index.php?rest_route=\/wp\/v2\/posts\/57\/revisions"}],"predecessor-version":[{"id":84,"href":"https:\/\/pomax-v3.weeltec.com\/index.php?rest_route=\/wp\/v2\/posts\/57\/revisions\/84"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/pomax-v3.weeltec.com\/index.php?rest_route=\/wp\/v2\/media\/54"}],"wp:attachment":[{"href":"https:\/\/pomax-v3.weeltec.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=57"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/pomax-v3.weeltec.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=57"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/pomax-v3.weeltec.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=57"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}