A Kubernetes upgrade strategy is not a maintenance window — it's a sequence, and the order matters more than the timing. Most upgrade incidents are not caused by the new version. They're caused by something that was already broken and only became visible when a node was asked to drain.

This guide covers what to check before you start, how to sequence the components, when blue-green beats in-place, and how to configure node rollouts so they finish.

Why Kubernetes Upgrades Break Production

The release cadence is not optional

Kubernetes ships three minor versions a year and supports each for roughly 14 months — twelve months of active patching plus two months of maintenance. Version 1.37 was released in August 2026 and reaches end of life in October 2027.

That maths is unforgiving. Falling two releases behind means you have under a year of runway, and catching up requires sequential upgrades because you cannot skip minor versions. Teams that defer upgrades for "a quieter quarter" typically discover they now need three upgrades instead of one, with the same amount of quiet.

Managed platforms soften this but do not remove it:

PlatformStandard supportExtended option
Amazon EKS14 monthsExtended support, ~12 extra months at a significantly higher per-cluster hourly rate
Google GKEVaries by release channelExtended support up to roughly 24 months
Azure AKS12 months for GA versionsLong-term support, roughly 2 years, opt-in

Extended support is a way to buy time, not a strategy. On EKS the cluster-hour cost increases several-fold, and you are still running a version that is no longer receiving upstream feature work.

The four things that actually break

In practice, upgrade incidents cluster into four causes:

  • Removed APIs. A manifest, Helm chart or operator still uses an API version the new release deleted. Deployments start failing after the control plane moves, not during it.
  • Add-on incompatibility. The CNI, CSI driver or ingress controller wasn't upgraded alongside the cluster. Symptoms are indirect — pods schedule but get no IP, or volumes fail to attach.
  • PodDisruptionBudgets that forbid disruption. A PDB with zero allowed disruptions makes a node drain impossible. The rollout hangs rather than failing.
  • No surge capacity. Draining a node with nowhere for its pods to go means the pods sit Pending and the service degrades for the length of the rollout.

None of these are version-specific. All of them are findable before you start.

Downtime is almost always the node phase

Control plane upgrades on managed platforms are effectively transparent — the API server is replaced behind a load balancer, and running pods are unaffected because the kubelet doesn't need the API server to keep containers running.

Traffic drops during node replacement. That is where the engineering effort belongs.

Pre-Upgrade Audit: Deprecated APIs, Add-ons and Version Skew

Find removed APIs before they find you

Three sources, used together, give complete coverage:

  • Live cluster traffic. The API server exposes apiserver_requested_deprecated_apis, a metric that reports which deprecated APIs are actively being called and which release removes them. This catches controllers and operators you forgot existed.
  • Stored manifests. Scan your Git repositories and Helm charts with a tool like pluto or kubent. This catches things not currently deployed but ready to be.
  • The upstream removal list. Check the Kubernetes deprecated API migration guide for the target version. It lists every removal with its replacement.

An important subtlety: objects stored in etcd under an old API version are served under the new one automatically, so a kubectl get returning results does not prove your manifests are current. The problem appears on the next apply.

Audit add-ons against the target version

Every add-on has its own compatibility matrix, and "it works today" tells you nothing about the next minor version.

ComponentWhere it breaks if mismatched
CNI pluginPods schedule but never get an IP
CoreDNSIntermittent DNS resolution failures across the cluster
kube-proxyService routing degrades; must not lead the control plane version
CSI driversVolumes fail to attach; StatefulSets stall
Ingress controllerAdmission webhook rejects or config translation changes
Metrics serverHPA stops scaling; no metrics to scale on
Cluster autoscaler / KarpenterVersion-pinned to the cluster; a mismatch stops node provisioning
Service meshFrequently the tightest constraint — check it first

Build the matrix once, as a table in your runbook, and update it at each upgrade. Service meshes and CSI drivers are the two that most often force a delay.

Understand the version skew you're allowed

Kubernetes defines precise limits on how far components may drift apart, and these are what make rolling upgrades possible:

  • kube-apiserver must be the newest component
  • kubelet may be up to three minor versions older than the API server
  • kube-proxy must match its node's kubelet and must not be newer than the API server
  • kube-controller-manager and kube-scheduler must be within one minor version of the API server
  • kubectl may be one minor version either side

The practical consequence is significant: nodes can lag the control plane by up to three minor versions. You are not obliged to finish the node rollout the same night. You can upgrade the control plane on Tuesday and roll nodes across the following two weeks, one node group at a time, in daylight.

This single fact removes most of the pressure from the process. The official version skew policy has the complete rules.

The audit checklist

In-Place vs Blue-Green Cluster Upgrades

In-place

Upgrade the control plane, then roll nodes within the same cluster. This is the default on every managed platform and the right choice most of the time.

Strengths: no data migration, no DNS changes, no duplicate cost, and the process is well-trodden.

Weaknesses: rollback is limited — control plane downgrades are generally not supported, so your recovery path is restore-from-backup rather than switch-back. The cluster is in a mixed-version state during the node rollout.

Blue-green

Stand up a new cluster at the target version, deploy workloads to it, shift traffic, then decommission the old one.

Strengths: rollback is a traffic switch. You can validate the new cluster under real load before committing. It also cleans up accumulated cluster drift — the CRDs nobody remembers installing, the orphaned ConfigMaps.

Weaknesses: double infrastructure cost for the overlap period, and every piece of state has to be dealt with. Persistent volumes do not move between clusters without deliberate work. External dependencies — IAM roles, security groups, database allow-lists, webhook endpoints — all need to point at the new cluster.

Which to choose

SituationApproach
Routine minor version, stateless workloadsIn-place
Catching up several versions at onceBlue-green — avoids repeated sequential upgrades
Regulated environment needing proven rollbackBlue-green
Changing something structural (CNI, node OS, control plane config)Blue-green
Heavy stateful workloads with large PVsIn-place — volume migration usually outweighs the benefit
Cost-constrained, single environmentIn-place

The middle path most teams land on: in-place for the control plane, blue-green at the node group level. Create a new node group at the target version, cordon the old one, migrate workloads across, delete the old group. You get switchable rollback on the risky part without duplicating the cluster.

This works particularly well with Karpenter, where node pools can be versioned and drained progressively — an approach we cover in Karpenter and EKS autoscaling.

Node Rollouts: Surge, Drain and PodDisruptionBudgets

This is the phase that drops traffic. Four things determine whether it hurts.

1. Surge before you drain

Set maxSurge to at least one node. The rollout then adds a new node at the target version, waits for it to become Ready, and only then drains an old one.

Without surge, drained pods have nowhere to go and sit Pending until a replacement node joins — which, including provisioning and image pull, can be three to six minutes of reduced capacity per node.

Set maxUnavailable to zero where the platform allows it. On EKS managed node groups, maxUnavailable: 1 with maxSurge: 1 gives a steady one-in-one-out rollout; raising surge to 25% or 33% of the group finishes faster at the cost of temporary over-provisioning.

2. Make PodDisruptionBudgets permit motion

A PDB is a promise about voluntary disruption. If it permits zero, kubectl drain will wait indefinitely rather than break it — which is correct behaviour and a common cause of stalled upgrades.

PDB settingEffect during drain
minAvailable: 1 on a 1-replica deploymentBlocks forever. No eviction is ever permitted.
maxUnavailable: 0Blocks forever. Same problem, different spelling.
minAvailable: 50% on 2 replicasAllows one at a time. Workable.
minAvailable: 1 on 3+ replicasAllows two at once. Fast, possibly too fast.
maxUnavailable: 1 on 3+ replicasPredictable, one at a time. Usually the right answer.

Audit every PDB before the rollout, not during it. The pattern that causes most incidents is a single-replica Deployment protected by minAvailable: 1 — created with good intentions, and guaranteed to block.

3. Give pods time to leave cleanly, but cap it

Three settings interact here:

  • terminationGracePeriodSeconds — how long the pod has to shut down after SIGTERM. The default of 30 seconds is too short for anything holding long connections.
  • A preStop hook with a short sleep — gives the endpoint controller and load balancer time to deregister the pod before the process starts shutting down. Without it, traffic is still arriving at a terminating pod.
  • Drain timeout — how long the rollout waits before giving up on a node. Unbounded means an indefinite hang; too short means pods get killed mid-request.

The sequence that produces clean drains: readiness probe fails → endpoint removed → preStop sleep → SIGTERM → in-flight requests complete → SIGKILL at grace period. Getting the ordering right here eliminates most of the 502s people attribute to "the upgrade."

4. Canary one node group first

Never roll every node group at once. Upgrade the least critical group, watch for 30 to 60 minutes, then proceed.

What to watch during and after the rollout:

  • Pod restart count and CrashLoopBackOff across all namespaces
  • Error rate and P99 latency at the ingress
  • Pending pods (indicates scheduling pressure or taint mismatch)
  • Node Ready status and kubelet errors
  • PVC attach failures
  • DNS resolution errors — CoreDNS problems often surface as unrelated application errors

If a group behaves, the rest usually will. If it doesn't, you have lost one node group instead of a cluster. Good monitoring and observability makes this phase a decision rather than a guess.

Stateful workloads need explicit handling

StatefulSets roll one pod at a time by ordinal and wait for each to become Ready. That is deliberate and safe, but it means an eight-replica database StatefulSet takes eight sequential restarts.

Before rolling nodes carrying stateful workloads:

  • Confirm the volume type supports the move — a zonal EBS volume cannot follow a pod to another AZ
  • For quorum systems, verify quorum is maintained at every point in the sequence
  • For databases, fail over deliberately rather than letting the drain trigger it
  • Consider draining stateful node groups last, after everything else has proven fine

Building an Upgrade Cadence That Holds

The reason upgrades feel dangerous is usually that they happen rarely. A team that upgrades twice a year has a hard, memorable, high-stakes event. A team that upgrades every release has a routine.

What that looks like in practice:

  • Upgrade non-production within two weeks of a release reaching your platform. Let it soak. Real workloads find real problems.
  • Production one release behind the newest is a reasonable standing target — recent enough to have runway, mature enough to avoid early-release surprises.
  • Patch versions monthly, without ceremony. They are low-risk and carry the security fixes.
  • Treat the add-on matrix as a living document, updated at every upgrade rather than rebuilt from scratch each time.
  • Automate the audit. Deprecated API scanning belongs in CI, not in a checklist someone works through at 10pm.

A note on managed platforms: EKS Auto Mode, GKE Autopilot and AKS auto-upgrade channels will handle node upgrades for you. They are genuinely useful and they do not remove your responsibility for the audit — an automated node rollout still stalls on a PDB that permits zero disruptions. Automation makes a correct process faster; it does not make an incorrect one safe.

If upgrades are the thing that keeps getting deferred, our managed Kubernetes service and 24×7 SRE teams run this cadence continuously. The related Kubernetes production readiness checklist covers the controls that make upgrades boring in the first place.

The teams for whom upgrades are boring are the ones who do them often. Cadence is the strategy.

If upgrades keep slipping, or you are on a version approaching end of life, our managed Kubernetes service handles this as routine work rather than a quarterly event.

Talk to our team → We will start with your current version, your add-on matrix and your PDBs — which is where the surprises are.