{"id":41,"date":"2026-10-01T16:28:34","date_gmt":"2026-10-01T16:28:34","guid":{"rendered":"https:\/\/pomax-v3.weeltec.com\/?p=41"},"modified":"2026-10-01T16:28:34","modified_gmt":"2026-10-01T16:28:34","slug":"day-2-operations-after-cluster-works","status":"publish","type":"post","link":"https:\/\/pomax-v3.weeltec.com\/?p=41","title":{"rendered":"Day-2 Operations: Life After the Cluster Works"},"content":{"rendered":"<div class=\"wt-post\">\n<style>.wt-post { --wt-bg: #0A0C0F; --wt-bg-2: #0E1116; --wt-surface: #12161C; --wt-line: #1F252E; --wt-line-strong: #2B333E; --wt-text: #E9ECEF; --wt-muted: #98A2AD; --wt-accent: #B9F24D; --wt-accent-hover: #C8FF5E; --wt-accent-dim: rgba(185, 242, 77, 0.10); --wt-accent-line: rgba(185, 242, 77, 0.25); --wt-danger: #F07A6B; --wt-ok: #7ED99A; --wt-sans: ui-sans-serif, system-ui, -apple-system, \"Segoe UI\", Roboto, \"Helvetica Neue\", Arial, sans-serif; --wt-mono: ui-monospace, \"Cascadia Code\", \"JetBrains Mono\", \"SF Mono\", Menlo, Consolas, monospace; --wt-radius: 4px; --wt-h1: var(--wt-text); --wt-h2: var(--wt-text); --wt-h3: var(--wt-text); box-sizing: border-box; background: var(--wt-bg); color: var(--wt-text); font-family: var(--wt-sans); font-size: 1rem; line-height: 1.7; padding: clamp(1.75rem, 4vw, 3rem); border: 1px solid var(--wt-line); border-radius: 0; -webkit-font-smoothing: antialiased; text-rendering: optimizeLegibility; } .wt-post *, .wt-post *::before, .wt-post *::after { box-sizing: border-box; } .wt-post ::selection { background: var(--wt-accent); color: var(--wt-bg); } .wt-post.wt-post p { color: var(--wt-text); font-family: var(--wt-sans); font-size: 1rem; line-height: 1.7; margin: 0 0 1.15rem; max-width: 68ch; } .wt-post.wt-post p:last-child { margin-bottom: 0; } .wt-post.wt-post h1, .wt-post.wt-post h2, .wt-post.wt-post h3, .wt-post.wt-post h4, .wt-post.wt-post h5, .wt-post.wt-post h6 { font-family: var(--wt-sans); font-weight: 700; letter-spacing: -0.025em; line-height: 1.15; text-wrap: balance; } .wt-post.wt-post h2 { color: var(--wt-h2); font-size: clamp(1.45rem, 3vw, 2rem); margin: 2.4rem 0 0.9rem; display: flex; align-items: baseline; gap: 0.6rem; } .wt-post.wt-post h2::before { content: \"\"; flex: none; width: 8px; height: 8px; background: var(--wt-accent); transform: translateY(-2px); } .wt-post.wt-post h3 { color: var(--wt-h3); font-size: 1.15rem; font-weight: 650; margin: 1.8rem 0 0.7rem; padding-left: 0.85rem; border-left: 2px solid var(--wt-accent-line); } .wt-post.wt-post h2:first-child, .wt-post.wt-post h3:first-child { margin-top: 0; } .wt-post.wt-post strong { color: #FFFFFF; font-weight: 650; } .wt-post.wt-post em { color: var(--wt-muted); font-style: italic; } .wt-post.wt-post a { color: var(--wt-accent); text-decoration: none; border-bottom: 1px solid var(--wt-accent-line); transition: color 0.15s ease, border-color 0.15s ease; } .wt-post.wt-post a:hover { color: var(--wt-accent-hover); border-bottom-color: var(--wt-accent-hover); } .wt-post.wt-post ul, .wt-post.wt-post ol { margin: 0 0 1.3rem; padding: 0; list-style: none; max-width: 68ch; } .wt-post.wt-post li { position: relative; padding-left: 1.6rem; margin-bottom: 0.55rem; color: var(--wt-text); line-height: 1.65; } .wt-post.wt-post ul > li::before { content: \"\\25AE\"; color: var(--wt-accent); position: absolute; left: 0; top: 0; font-size: 0.85em; line-height: 1.65; } .wt-post.wt-post ol { counter-reset: wt-li; } .wt-post.wt-post ol > li { counter-increment: wt-li; } .wt-post.wt-post ol > li::before { content: counter(wt-li) \".\"; font-family: var(--wt-mono); font-size: 0.8em; color: var(--wt-accent); position: absolute; left: 0; top: 0; line-height: 1.9; } .wt-post.wt-post blockquote { margin: 1.8rem 0; padding: 1.1rem 1.4rem; background: var(--wt-accent-dim); border-left: 2px solid var(--wt-accent); border-radius: 0; color: var(--wt-text); font-size: 1.05rem; font-style: normal; line-height: 1.6; } .wt-post.wt-post blockquote p { margin: 0; color: var(--wt-text); font-style: normal; } .wt-post.wt-post blockquote::before { content: none; } .wt-post.wt-post code, .wt-post.wt-post kbd, .wt-post.wt-post pre { font-family: var(--wt-mono); font-size: 0.88em; } .wt-post.wt-post code { background: var(--wt-surface); border: 1px solid var(--wt-line); border-radius: var(--wt-radius); padding: 0.1em 0.4em; color: var(--wt-accent); } .wt-post.wt-post pre { background: #0C0F13; border: 1px solid var(--wt-line-strong); border-radius: var(--wt-radius); padding: 1.1rem 1.25rem; overflow-x: auto; color: var(--wt-text); line-height: 1.7; margin: 0 0 1.3rem; } .wt-post.wt-post pre code { background: none; border: 0; padding: 0; color: inherit; } .wt-post.wt-post hr { border: 0; border-top: 1px solid var(--wt-line); margin: 2.2rem 0; } .wt-post.wt-post img { max-width: 100%; height: auto; border-radius: var(--wt-radius); border: 1px solid var(--wt-line); } @media (max-width: 640px) { .wt-post.wt-post h2 { font-size: 1.35rem; } .wt-post.wt-post blockquote { padding: 0.9rem 1.1rem; } } @media (prefers-reduced-motion: reduce) { .wt-post.wt-post a { transition: none; } }<\/style>\n<p>Cluster provisioning is a solved problem. A control plane, a node pool and a handful of operators get you to a working cluster in an afternoon. Day 1 is easy, which is why it gets all the attention.<\/p>\n<p>Day 2 is everything that happens after \"it works\". It is also where platforms fail \u2014 slowly, through accumulation, rather than dramatically.<\/p>\n<h2>What day 2 actually contains<\/h2>\n<p>Upgrades, certificate expiry, node image rotation, capacity drift, dependency patching, cost creep, incident response, and the slow decay of conventions nobody wrote down. None of it is a project with a finish line. All of it is ongoing.<\/p>\n<h3>The things that page you at night<\/h3>\n<ul>\n<li><strong>Expired certificates.<\/strong> Admission webhooks, internal service certs and ingress certificates all carry clocks. Rotation has to be automated, not remembered.<\/li>\n<li><strong>Version skew.<\/strong> The control plane may be supported while a node pool sits three minor versions behind, and the upgrade path requires stepping through each one.<\/li>\n<li><strong>Disk pressure.<\/strong> Container images, logs and etcd on modest control-plane nodes fill up quietly, and eviction storms follow.<\/li>\n<li><strong>Cluster CA expiry.<\/strong> A date in the far future that gets forgotten for years and then ends the cluster.<\/li>\n<li><strong>Deprecated APIs.<\/strong> A manifest that worked on 1.24 fails on 1.29. Removal is announced, ignored, then enforced.<\/li>\n<\/ul>\n<h2>Operating habits that hold up<\/h2>\n<p>Upgrade one minor version at a time and never jump. Read the release notes for removals rather than features. Patch node images on a schedule, and drain nodes with <strong>kubectl drain --ignore-daemonsets<\/strong> only after confirming that your PodDisruptionBudgets permit the disruption, not when the drain hangs and you have to find out live.<\/p>\n<p>Back up etcd, and more importantly, restore a snapshot on a schedule. A snapshot that has never been loaded is a file, not a backup.<\/p>\n<p>Do the upgrades during business hours, deliberately, with the team awake and watching. A controlled ten-minute disruption beats an uncontrolled one at 4 a.m. every time.<\/p>\n<p>Keep a written runbook for the handful of things that genuinely break: a node stuck NotReady, an unreachable control plane, ingress returning 502, and a workload that will not schedule. Everything else is improvisation under pressure, which is another way of saying outage time.<\/p>\n<blockquote>\n<p>Day 1 is a deployment. Day 2 is a practice. One of them ends, and it is not this one.<\/p>\n<\/blockquote>\n<h3>Observability, minimally<\/h3>\n<p>Four signals matter before anything else: are the nodes healthy, are workloads running the version you believe they are, are the certificates valid, and is anything stuck Pending. Dashboards are optional; alerting on those four is not. Most day-2 incidents announce themselves as a Pending Pod or a NotReady node long before a user notices anything.<\/p>\n<h3>Change management on a live cluster<\/h3>\n<p>Everything you merge lands on a running system, so keep changes small and reversible, and exercise the rollback path at least once before you rely on it. A deployment strategy that exists only in documentation is a plan, not a capability. On clusters carrying stateful workloads, confirm that PersistentVolumeClaims and their storage classes survive a rollback before the night you need them to.<\/p>\n<h2>The cost dimension<\/h2>\n<p>Clusters drift toward waste without a single dramatic cause. Load balancers orphaned by deleted Services, PersistentVolumes that outlive their claims, node pools sized for a launch that ended two quarters ago, and requests set once during a migration and never revisited. A monthly review of requested against used capacity typically finds between a fifth and a third of spend recoverable, and it takes an hour. None of these appear on an architecture diagram, and all of them appear on the invoice.<\/p>\n<h2>How Weeltec runs day 2<\/h2>\n<p>Ongoing cluster operations are mostly calendar work: scheduled upgrades with a tested rollback path, certificate and cluster CA rotation that is automated rather than hoped for, monthly capacity and cost reviews, and a runbook set a new engineer can follow at 3 a.m. without calling anyone. The difference between a stable platform and an unstable one is rarely the architecture. It is whether someone does the boring work on time.<\/p>\n<h2>The rule<\/h2>\n<p>Treat the cluster as a product with a roadmap, not a project with an end date. Schedule the upgrades, automate the rotation, measure the spend, and write down what breaks.<\/p>\n<p><em>We operate Kubernetes platforms for teams that would rather ship than babysit. If day-2 work keeps landing on your engineers' evenings, <a href=\"https:\/\/weeltec.com\/#contact\">get a quote<\/a> and we will take it off your hands.<\/em><\/p>\n<\/div>\n","protected":false},"excerpt":{"rendered":"<p>Provisioning a cluster is the easy part. Upgrades, rotation, capacity drift and cost creep decide whether the platform survives its second year.<\/p>\n","protected":false},"author":1,"featured_media":40,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[7],"tags":[14,13,16,15],"class_list":["post-41","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-kubernetes","tag-day-2-operations","tag-kubernetes","tag-platform-engineering","tag-upgrades"],"_links":{"self":[{"href":"https:\/\/pomax-v3.weeltec.com\/index.php?rest_route=\/wp\/v2\/posts\/41","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/pomax-v3.weeltec.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/pomax-v3.weeltec.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/pomax-v3.weeltec.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/pomax-v3.weeltec.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=41"}],"version-history":[{"count":2,"href":"https:\/\/pomax-v3.weeltec.com\/index.php?rest_route=\/wp\/v2\/posts\/41\/revisions"}],"predecessor-version":[{"id":78,"href":"https:\/\/pomax-v3.weeltec.com\/index.php?rest_route=\/wp\/v2\/posts\/41\/revisions\/78"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/pomax-v3.weeltec.com\/index.php?rest_route=\/wp\/v2\/media\/40"}],"wp:attachment":[{"href":"https:\/\/pomax-v3.weeltec.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=41"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/pomax-v3.weeltec.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=41"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/pomax-v3.weeltec.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=41"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}