A Kubernetes upgrade strategy is not a maintenance window — it's a sequence, and the order matters more than the timing. Most upgrade incidents are not caused by the new version. They're caused by something that was already broken and only became visible when a node was asked to drain.
This guide covers what to check before you start, how to sequence the components, when blue-green beats in-place, and how to configure node rollouts so they finish.
Why Kubernetes Upgrades Break Production
The release cadence is not optional
Kubernetes ships three minor versions a year and supports each for roughly 14 months — twelve months of active patching plus two months of maintenance. Version 1.37 was released in August 2026 and reaches end of life in October 2027.
That maths is unforgiving. Falling two releases behind means you have under a year of runway, and catching up requires sequential upgrades because you cannot skip minor versions. Teams that defer upgrades for "a quieter quarter" typically discover they now need three upgrades instead of one, with the same amount of quiet.
Managed platforms soften this but do not remove it:
| Platform | Standard support | Extended option |
|---|---|---|
| Amazon EKS | 14 months | Extended support, ~12 extra months at a significantly higher per-cluster hourly rate |
| Google GKE | Varies by release channel | Extended support up to roughly 24 months |
| Azure AKS | 12 months for GA versions | Long-term support, roughly 2 years, opt-in |
Extended support is a way to buy time, not a strategy. On EKS the cluster-hour cost increases several-fold, and you are still running a version that is no longer receiving upstream feature work.
The four things that actually break
In practice, upgrade incidents cluster into four causes:
- Removed APIs. A manifest, Helm chart or operator still uses an API version the new release deleted. Deployments start failing after the control plane moves, not during it.
- Add-on incompatibility. The CNI, CSI driver or ingress controller wasn't upgraded alongside the cluster. Symptoms are indirect — pods schedule but get no IP, or volumes fail to attach.
- PodDisruptionBudgets that forbid disruption. A PDB with zero allowed disruptions makes a node drain impossible. The rollout hangs rather than failing.
- No surge capacity. Draining a node with nowhere for its pods to go means the pods sit Pending and the service degrades for the length of the rollout.
None of these are version-specific. All of them are findable before you start.
Downtime is almost always the node phase
Control plane upgrades on managed platforms are effectively transparent — the API server is replaced behind a load balancer, and running pods are unaffected because the kubelet doesn't need the API server to keep containers running.
Traffic drops during node replacement. That is where the engineering effort belongs.
Pre-Upgrade Audit: Deprecated APIs, Add-ons and Version Skew
Find removed APIs before they find you
Three sources, used together, give complete coverage:
- Live cluster traffic. The API server exposes
apiserver_requested_deprecated_apis, a metric that reports which deprecated APIs are actively being called and which release removes them. This catches controllers and operators you forgot existed. - Stored manifests. Scan your Git repositories and Helm charts with a tool like
plutoorkubent. This catches things not currently deployed but ready to be. - The upstream removal list. Check the Kubernetes deprecated API migration guide for the target version. It lists every removal with its replacement.
An important subtlety: objects stored in etcd under an old API version are served under the new one automatically, so a kubectl get returning results does not prove your manifests are current. The problem appears on the next apply.
Audit add-ons against the target version
Every add-on has its own compatibility matrix, and "it works today" tells you nothing about the next minor version.
| Component | Where it breaks if mismatched |
|---|---|
| CNI plugin | Pods schedule but never get an IP |
| CoreDNS | Intermittent DNS resolution failures across the cluster |
| kube-proxy | Service routing degrades; must not lead the control plane version |
| CSI drivers | Volumes fail to attach; StatefulSets stall |
| Ingress controller | Admission webhook rejects or config translation changes |
| Metrics server | HPA stops scaling; no metrics to scale on |
| Cluster autoscaler / Karpenter | Version-pinned to the cluster; a mismatch stops node provisioning |
| Service mesh | Frequently the tightest constraint — check it first |
Build the matrix once, as a table in your runbook, and update it at each upgrade. Service meshes and CSI drivers are the two that most often force a delay.
Understand the version skew you're allowed
Kubernetes defines precise limits on how far components may drift apart, and these are what make rolling upgrades possible:
- kube-apiserver must be the newest component
- kubelet may be up to three minor versions older than the API server
- kube-proxy must match its node's kubelet and must not be newer than the API server
- kube-controller-manager and kube-scheduler must be within one minor version of the API server
- kubectl may be one minor version either side
The practical consequence is significant: nodes can lag the control plane by up to three minor versions. You are not obliged to finish the node rollout the same night. You can upgrade the control plane on Tuesday and roll nodes across the following two weeks, one node group at a time, in daylight.
This single fact removes most of the pressure from the process. The official version skew policy has the complete rules.
The audit checklist
In-Place vs Blue-Green Cluster Upgrades
In-place
Upgrade the control plane, then roll nodes within the same cluster. This is the default on every managed platform and the right choice most of the time.
Strengths: no data migration, no DNS changes, no duplicate cost, and the process is well-trodden.
Weaknesses: rollback is limited — control plane downgrades are generally not supported, so your recovery path is restore-from-backup rather than switch-back. The cluster is in a mixed-version state during the node rollout.
Blue-green
Stand up a new cluster at the target version, deploy workloads to it, shift traffic, then decommission the old one.
Strengths: rollback is a traffic switch. You can validate the new cluster under real load before committing. It also cleans up accumulated cluster drift — the CRDs nobody remembers installing, the orphaned ConfigMaps.
Weaknesses: double infrastructure cost for the overlap period, and every piece of state has to be dealt with. Persistent volumes do not move between clusters without deliberate work. External dependencies — IAM roles, security groups, database allow-lists, webhook endpoints — all need to point at the new cluster.
Which to choose
| Situation | Approach |
|---|---|
| Routine minor version, stateless workloads | In-place |
| Catching up several versions at once | Blue-green — avoids repeated sequential upgrades |
| Regulated environment needing proven rollback | Blue-green |
| Changing something structural (CNI, node OS, control plane config) | Blue-green |
| Heavy stateful workloads with large PVs | In-place — volume migration usually outweighs the benefit |
| Cost-constrained, single environment | In-place |
The middle path most teams land on: in-place for the control plane, blue-green at the node group level. Create a new node group at the target version, cordon the old one, migrate workloads across, delete the old group. You get switchable rollback on the risky part without duplicating the cluster.
This works particularly well with Karpenter, where node pools can be versioned and drained progressively — an approach we cover in Karpenter and EKS autoscaling.
Node Rollouts: Surge, Drain and PodDisruptionBudgets
This is the phase that drops traffic. Four things determine whether it hurts.
1. Surge before you drain
Set maxSurge to at least one node. The rollout then adds a new node at the target version, waits for it to become Ready, and only then drains an old one.
Without surge, drained pods have nowhere to go and sit Pending until a replacement node joins — which, including provisioning and image pull, can be three to six minutes of reduced capacity per node.
Set maxUnavailable to zero where the platform allows it. On EKS managed node groups, maxUnavailable: 1 with maxSurge: 1 gives a steady one-in-one-out rollout; raising surge to 25% or 33% of the group finishes faster at the cost of temporary over-provisioning.
2. Make PodDisruptionBudgets permit motion
A PDB is a promise about voluntary disruption. If it permits zero, kubectl drain will wait indefinitely rather than break it — which is correct behaviour and a common cause of stalled upgrades.
| PDB setting | Effect during drain |
|---|---|
minAvailable: 1 on a 1-replica deployment | Blocks forever. No eviction is ever permitted. |
maxUnavailable: 0 | Blocks forever. Same problem, different spelling. |
minAvailable: 50% on 2 replicas | Allows one at a time. Workable. |
minAvailable: 1 on 3+ replicas | Allows two at once. Fast, possibly too fast. |
maxUnavailable: 1 on 3+ replicas | Predictable, one at a time. Usually the right answer. |
Audit every PDB before the rollout, not during it. The pattern that causes most incidents is a single-replica Deployment protected by minAvailable: 1 — created with good intentions, and guaranteed to block.
3. Give pods time to leave cleanly, but cap it
Three settings interact here:
terminationGracePeriodSeconds— how long the pod has to shut down after SIGTERM. The default of 30 seconds is too short for anything holding long connections.- A
preStophook with a short sleep — gives the endpoint controller and load balancer time to deregister the pod before the process starts shutting down. Without it, traffic is still arriving at a terminating pod. - Drain timeout — how long the rollout waits before giving up on a node. Unbounded means an indefinite hang; too short means pods get killed mid-request.
The sequence that produces clean drains: readiness probe fails → endpoint removed → preStop sleep → SIGTERM → in-flight requests complete → SIGKILL at grace period. Getting the ordering right here eliminates most of the 502s people attribute to "the upgrade."
4. Canary one node group first
Never roll every node group at once. Upgrade the least critical group, watch for 30 to 60 minutes, then proceed.
What to watch during and after the rollout:
- Pod restart count and CrashLoopBackOff across all namespaces
- Error rate and P99 latency at the ingress
- Pending pods (indicates scheduling pressure or taint mismatch)
- Node Ready status and kubelet errors
- PVC attach failures
- DNS resolution errors — CoreDNS problems often surface as unrelated application errors
If a group behaves, the rest usually will. If it doesn't, you have lost one node group instead of a cluster. Good monitoring and observability makes this phase a decision rather than a guess.
Stateful workloads need explicit handling
StatefulSets roll one pod at a time by ordinal and wait for each to become Ready. That is deliberate and safe, but it means an eight-replica database StatefulSet takes eight sequential restarts.
Before rolling nodes carrying stateful workloads:
- Confirm the volume type supports the move — a zonal EBS volume cannot follow a pod to another AZ
- For quorum systems, verify quorum is maintained at every point in the sequence
- For databases, fail over deliberately rather than letting the drain trigger it
- Consider draining stateful node groups last, after everything else has proven fine
Building an Upgrade Cadence That Holds
The reason upgrades feel dangerous is usually that they happen rarely. A team that upgrades twice a year has a hard, memorable, high-stakes event. A team that upgrades every release has a routine.
What that looks like in practice:
- Upgrade non-production within two weeks of a release reaching your platform. Let it soak. Real workloads find real problems.
- Production one release behind the newest is a reasonable standing target — recent enough to have runway, mature enough to avoid early-release surprises.
- Patch versions monthly, without ceremony. They are low-risk and carry the security fixes.
- Treat the add-on matrix as a living document, updated at every upgrade rather than rebuilt from scratch each time.
- Automate the audit. Deprecated API scanning belongs in CI, not in a checklist someone works through at 10pm.
A note on managed platforms: EKS Auto Mode, GKE Autopilot and AKS auto-upgrade channels will handle node upgrades for you. They are genuinely useful and they do not remove your responsibility for the audit — an automated node rollout still stalls on a PDB that permits zero disruptions. Automation makes a correct process faster; it does not make an incorrect one safe.
If upgrades are the thing that keeps getting deferred, our managed Kubernetes service and 24×7 SRE teams run this cadence continuously. The related Kubernetes production readiness checklist covers the controls that make upgrades boring in the first place.
The teams for whom upgrades are boring are the ones who do them often. Cadence is the strategy.
If upgrades keep slipping, or you are on a version approaching end of life, our managed Kubernetes service handles this as routine work rather than a quarterly event.
Talk to our team → We will start with your current version, your add-on matrix and your PDBs — which is where the surprises are.