Cluster provisioning is a solved problem. A control plane, a node pool and a handful of operators get you to a working cluster in an afternoon. Day 1 is easy, which is why it gets all the attention.
Day 2 is everything that happens after "it works". It is also where platforms fail — slowly, through accumulation, rather than dramatically.
What day 2 actually contains
Upgrades, certificate expiry, node image rotation, capacity drift, dependency patching, cost creep, incident response, and the slow decay of conventions nobody wrote down. None of it is a project with a finish line. All of it is ongoing.
The things that page you at night
- Expired certificates. Admission webhooks, internal service certs and ingress certificates all carry clocks. Rotation has to be automated, not remembered.
- Version skew. The control plane may be supported while a node pool sits three minor versions behind, and the upgrade path requires stepping through each one.
- Disk pressure. Container images, logs and etcd on modest control-plane nodes fill up quietly, and eviction storms follow.
- Cluster CA expiry. A date in the far future that gets forgotten for years and then ends the cluster.
- Deprecated APIs. A manifest that worked on 1.24 fails on 1.29. Removal is announced, ignored, then enforced.
Operating habits that hold up
Upgrade one minor version at a time and never jump. Read the release notes for removals rather than features. Patch node images on a schedule, and drain nodes with kubectl drain --ignore-daemonsets only after confirming that your PodDisruptionBudgets permit the disruption, not when the drain hangs and you have to find out live.
Back up etcd, and more importantly, restore a snapshot on a schedule. A snapshot that has never been loaded is a file, not a backup.
Do the upgrades during business hours, deliberately, with the team awake and watching. A controlled ten-minute disruption beats an uncontrolled one at 4 a.m. every time.
Keep a written runbook for the handful of things that genuinely break: a node stuck NotReady, an unreachable control plane, ingress returning 502, and a workload that will not schedule. Everything else is improvisation under pressure, which is another way of saying outage time.
Day 1 is a deployment. Day 2 is a practice. One of them ends, and it is not this one.
Observability, minimally
Four signals matter before anything else: are the nodes healthy, are workloads running the version you believe they are, are the certificates valid, and is anything stuck Pending. Dashboards are optional; alerting on those four is not. Most day-2 incidents announce themselves as a Pending Pod or a NotReady node long before a user notices anything.
Change management on a live cluster
Everything you merge lands on a running system, so keep changes small and reversible, and exercise the rollback path at least once before you rely on it. A deployment strategy that exists only in documentation is a plan, not a capability. On clusters carrying stateful workloads, confirm that PersistentVolumeClaims and their storage classes survive a rollback before the night you need them to.
The cost dimension
Clusters drift toward waste without a single dramatic cause. Load balancers orphaned by deleted Services, PersistentVolumes that outlive their claims, node pools sized for a launch that ended two quarters ago, and requests set once during a migration and never revisited. A monthly review of requested against used capacity typically finds between a fifth and a third of spend recoverable, and it takes an hour. None of these appear on an architecture diagram, and all of them appear on the invoice.
How Weeltec runs day 2
Ongoing cluster operations are mostly calendar work: scheduled upgrades with a tested rollback path, certificate and cluster CA rotation that is automated rather than hoped for, monthly capacity and cost reviews, and a runbook set a new engineer can follow at 3 a.m. without calling anyone. The difference between a stable platform and an unstable one is rarely the architecture. It is whether someone does the boring work on time.
The rule
Treat the cluster as a product with a roadmap, not a project with an end date. Schedule the upgrades, automate the rotation, measure the spend, and write down what breaks.
We operate Kubernetes platforms for teams that would rather ship than babysit. If day-2 work keeps landing on your engineers' evenings, get a quote and we will take it off your hands.