Kubernetes networking comes down to three decisions, and they have wildly different reversal costs. The CNI is close to permanent. The ingress controller is moderately painful to swap. The service mesh is the easiest to remove and the one teams agonise over most — which is exactly backwards.

This guide covers what each layer actually does, how the current options compare, and a framework for choosing based on team size and compliance requirements rather than on what is fashionable.

The Three Networking Decisions Every Cluster Makes

Every cluster answers three questions whether or not anyone writes them down.

LayerHandlesReversal cost
CNIPod-to-pod connectivity, IP allocation, network policy enforcementHigh — swapping requires draining every node
IngressNorth-south traffic entering the cluster, TLS, routingModerate — run both, then move DNS
Service meshEast-west mTLS, L7 policy, retries, traffic shiftingLow — scope per namespace and back out

Why the order of attention matters

Teams routinely spend three weeks debating service meshes and ten minutes accepting whatever CNI the managed service defaulted to. That is the wrong allocation.

Changing a CNI in a running cluster means every node drains and every pod restarts, because pod networking is reconfigured at the node level. In a cluster with stateful workloads that is a planned maintenance window with tested rollback, not a Saturday morning task. Some teams find it genuinely easier to build a new cluster and migrate.

A service mesh, by contrast, can be enabled on one namespace, evaluated, and removed. The reversible decision deserves less deliberation than the irreversible one.

The layers overlap more than they used to

The neat separation of "CNI does L3/L4, mesh does L7" has largely dissolved. Modern eBPF-based CNIs handle identity, encryption, L7-aware policy and flow-level observability in the kernel. Gateway API implementations increasingly handle east-west routing as well as north-south.

The practical consequence: the question is no longer "which mesh" but "do I need a mesh on top of what my CNI already does". For a growing number of teams, the answer is no.

Choosing a CNI: Cilium, Calico and Cloud-Native Options

What the CNI is actually responsible for

  • Allocating pod IP addresses and defining the IP model
  • Routing packets between pods, nodes and the outside world
  • Enforcing NetworkPolicy (a spec that requires a CNI to implement — apply a policy on a CNI that doesn't and nothing happens, silently)
  • Optionally: transparent encryption, load balancing, and flow telemetry

That third point deserves emphasis. NetworkPolicy resources are accepted by the API server regardless of whether anything enforces them. Confirm your CNI implements them before assuming you have isolation.

The three realistic choices

Cilium — eBPF-based, CNCF graduated, and effectively the direction the industry has moved. It handles routing, policy, load balancing and encryption in the kernel, with Hubble providing flow-level observability without sidecars. It is also increasingly the default underneath managed platforms: GKE Dataplane V2 is built on Cilium and enabled by default in Autopilot, and Azure CNI Powered by Cilium is the default on AKS Automatic.

Trade-offs: requires a reasonably modern kernel, has a larger surface area than simpler CNIs, and its feature breadth means more to learn and more to misconfigure.

Calico — the mature, well-understood option with a long operational track record. Offers both an iptables dataplane and an eBPF dataplane, so you can adopt eBPF incrementally. Its policy model is rich and predictable, and Calico Enterprise adds compliance reporting that some regulated environments want.

Trade-offs: the iptables dataplane doesn't match eBPF performance at scale, and observability requires more assembly than Hubble gives you out of the box.

Cloud-native CNIs (AWS VPC CNI, Azure CNI, GKE default) — pods receive real VPC IP addresses, which means security groups, VPC flow logs and existing network tooling work without translation. Least to operate, because the cloud provider owns it.

Trade-offs: IP address consumption is the big one. Pods consume VPC addresses, and per-node pod density is capped by ENI and IP limits on the instance type. Run a large cluster on the AWS VPC CNI without an IP allocation plan and you will exhaust your subnets — which is why address space needs planning before the first cluster exists, not after.

Comparison

 CiliumCalicoCloud CNI
DataplaneeBPFiptables or eBPFCloud-native
NetworkPolicyYes, plus L7-aware CRDsYes, rich policy modelYes on major clouds
EncryptionWireGuard or IPsecWireGuardVPC-level
ObservabilityHubble, built inRequires assemblyCloud flow logs
IP modelOverlay or native routingOverlay or BGPVPC IPs per pod
Pod densityHighHighENI/IP limited
Operational loadMediumMediumLowest

A defensible default

For a new cluster, Cilium is the reasonable default — particularly if your managed platform already offers it, in which case you get the capability without owning the upgrade path. Enable Hubble on day one; flow visibility is the thing teams most regret not having during an incident.

Choose Calico if your team already knows it well, if you need its specific enterprise compliance features, or if kernel-level eBPF raises audit questions in your environment.

Choose the cloud CNI if your cluster is small, your team is small, and VPC-native IPs simplify your existing security tooling — as long as you have planned the address space.

Whatever you choose, decide before the cluster exists. This is the layer where "we'll revisit it later" carries a real bill.

Choosing an Ingress Controller: Ingress vs Gateway API

The decision was made for you in March 2026

The community ingress-nginx controller — used by roughly half of all Kubernetes clusters — was retired in March 2026. Kubernetes SIG Network and the Security Response Committee announced the retirement in November 2025, citing maintainer burnout and accumulated technical debt around features like configuration-snippet annotations. Best-effort maintenance ended on schedule and the repository is archived and read-only.

What that means concretely:

  • No further releases, bug fixes, or security patches. Existing deployments keep routing traffic; they just stop receiving CVE fixes.
  • Existing artifacts remain available. Helm charts and images are still downloadable.
  • The Ingress API itself is not removed. It remains GA but feature-frozen — active development moved to Gateway API.
  • nginxinc/kubernetes-ingress, F5's separate commercial controller, is a different project and is not affected.

If you are still running community ingress-nginx in production, you are operating an unmaintained internet-facing component. The official retirement notice recommends migrating to a Gateway API implementation.

Why Gateway API replaced Ingress

Ingress had three structural problems that could not be fixed within the API:

  • Vendor-specific annotations. Anything beyond basic path routing required controller-specific annotations, making manifests non-portable.
  • No role separation. Platform and application teams both edited the same resource, with no way to delegate safely.
  • HTTP only. No first-class support for TCP, UDP, gRPC or TLS passthrough.

Gateway API fixes all three with a role-oriented resource model:

ResourceOwned byPurpose
GatewayClassInfrastructure providerThe implementation itself
GatewayPlatform teamListeners, ports, TLS certificates
HTTPRoute / GRPCRoute / TCPRouteApplication teamRouting rules for their own services

Application teams create routes and attach them to a Gateway the platform team owns. Traffic splitting, header-based routing and request mirroring are first-class fields rather than annotations. Gateway API reached GA in October 2023 and has continued to add capability since.

Picking an implementation

ImplementationBest for
Cilium GatewayClusters already running Cilium — no additional component to operate
Envoy GatewayTeams wanting the Envoy ecosystem without full Istio
NGINX Gateway FabricTeams with existing NGINX expertise and config to carry over
Cloud LB controllers (AWS Load Balancer Controller, GKE)Teams preferring managed L7 with cloud-native integration
Traefik, HAProxy, KongExisting users — all have Gateway API support

The strongest consolidation argument: if you run Cilium, its Gateway API implementation means one fewer component. Fewer moving parts is a real operational benefit, not a minor one.

Migrating without a maintenance window

The path is straightforward and does not require downtime:

  1. Install a Gateway API implementation alongside the existing controller.
  2. Convert existing Ingress resources with the ingress2gateway tool, then review the output — annotation-heavy configs need manual attention.
  3. Create a Gateway with its own load balancer address.
  4. Move services across one at a time, shifting DNS per service and validating each.
  5. Decommission the old controller once nothing routes through it.

Budget most of the effort for step 2. Annotations are where the custom behaviour hides — rate limits, auth snippets, rewrite rules — and they do not convert automatically.

When You Actually Need a Service Mesh — and When You Don't

Start from none

The default answer is no mesh. A service mesh is a distributed system with its own control plane, its own failure modes, its own upgrade cadence, and its own debugging surface. It should earn its place.

The three questions

Add a mesh only if at least one of these is a clear yes:

1. Do you need L7 authorisation? Not "encrypt traffic between services" — a modern CNI does transparent encryption with WireGuard or IPsec at the node level. The mesh case is per-path, per-method, per-identity rules: service A may call POST /orders on service B, but not DELETE /orders/{id}.

2. Do you need traffic shifting by request attribute? Canary releases by header value, request mirroring to a shadow environment, retries with budgets, circuit breaking, fault injection. If you deploy with straightforward percentage rollouts, Gateway API HTTPRoute weights already cover you.

3. Do you need uniform telemetry you cannot get otherwise? Consistent golden-signal metrics for every service without instrumenting each one. Note that eBPF-based CNIs give much of this via flow telemetry — the gap is narrower than it was.

If all three are no, a modern CNI covers your requirements and a mesh adds cost without adding capability.

If the answer is yes

Istio ambient mode is the current default for teams that need a full mesh. Ambient replaces per-pod sidecars with ztunnel, a per-node DaemonSet handling L4 mTLS, plus optional per-namespace waypoint proxies deployed only where L7 features are actually needed. It reached GA in Istio 1.24, and the resource reduction relative to sidecars is substantial.

The architectural benefit is that you pay for L7 only where you use it, rather than forcing an Envoy sidecar onto every pod in the cluster.

Cilium Service Mesh makes sense if Cilium is already your CNI. mTLS, L7-aware policy and observability without a separate control plane. The caveat worth knowing: Cilium's mutual authentication uses eventual consistency for policy sync, so environments requiring instantaneous and fully auditable policy propagation may prefer Istio's synchronous model.

Linkerd remains the lightest-weight option operationally, built on a purpose-made Rust proxy rather than general-purpose Envoy. Check its current licensing and release model before committing, as the distribution of stable builds has changed.

The sidecar tax, quantified

Traditional sidecars add roughly 50–100 MB of memory per pod and 1–3 ms of latency per hop. At 1,000 pods that is 50–100 GB of memory spent on proxies. This is why every major mesh has moved away from the sidecar model, and why "we'll add a mesh later if we need it" is a more reasonable position than it was three years ago.

The anti-patterns

  • Adopting a mesh to satisfy a compliance checkbox. If the requirement is encryption in transit, CNI-level WireGuard satisfies it with a fraction of the complexity.
  • Running two meshes across overlapping namespaces. Fine as a migration pattern, scoped per namespace. As a permanent state, overlapping mTLS and policy layers produce failures that are genuinely hard to diagnose.
  • Adopting a mesh before you have observability. A mesh generates a great deal of telemetry. Without somewhere to put it and someone reading it, you have added complexity and gained nothing.

Decision Framework by Team Size and Compliance Needs

By team size

Cluster / team sizeCNIIngressMesh
Under 10 nodes, one teamCloud CNI or Cilium via managed defaultCloud LB controller or Cilium GatewayNone
10–50 nodes, few teamsCiliumGateway API — Cilium Gateway or Envoy GatewayNone, unless L7 policy is a hard requirement
50–200 nodes, multiple teamsCilium with HubbleGateway API with role separation enforcedEvaluate — usually yes if you need L7 authorisation
200+ nodes, platform teamCilium, possibly Calico EnterpriseGateway API, likely multiple Gateways per tenantIstio ambient or Cilium mesh

The consistent pattern: the mesh decision correlates with organisational complexity, not cluster size. Fifty services owned by one team rarely need a mesh. Fifteen services owned by six teams with different security requirements often do.

By compliance requirement

RequirementWhat actually satisfies it
Encryption in transitCNI-level WireGuard or IPsec. A mesh is not required.
Workload identity and mTLSMesh, or Cilium's mutual authentication
Network segmentationNetworkPolicy, enforced by any competent CNI
L7 authorisation and auditMesh — this is the strongest genuine case
Flow logging for auditHubble, Calico flow logs, or cloud VPC flow logs
FIPS-validated cryptographyCheck specific builds; not every implementation offers them

The useful discipline here is separating what the auditor actually asks for from what a vendor says satisfies it. "Encryption between services" is a CNI feature. "Demonstrate that service A cannot call endpoint X on service B, with an audit trail" is a mesh feature.

Two rules that hold across all sizes

  • Decide the CNI before the cluster exists. Everything else can be added, swapped or removed later. This one cannot, cheaply.
  • Add the mesh last, if at all. Get the CNI right, get ingress onto Gateway API, get observability working. Then ask whether a mesh still adds something. Frequently it no longer does.

Conclusion

Three decisions, and the effort should be distributed inversely to how much they get debated.

The CNI is the one that matters. It determines your IP model, your policy enforcement point, your encryption story and your flow visibility, and changing it means draining every node in the cluster. For most new clusters Cilium is the defensible default — increasingly because the managed platforms have already chosen it. Calico remains a sound choice for teams with existing expertise or specific compliance tooling needs. Cloud-native CNIs are fine for smaller clusters with a real IP allocation plan.

Ingress is no longer a preference question. With community ingress-nginx retired and unpatched since March 2026, the only live question is which Gateway API implementation you move to and how quickly. If Cilium is your CNI, its Gateway implementation removes a component rather than adding one. Migration is incremental — run both, move services across, shift DNS per service — and the effort concentrates in converting annotation-heavy configuration, not in the mechanics.

The service mesh is optional, and the bar has risen. Modern eBPF CNIs cover encryption, identity and flow telemetry in the kernel. What remains genuinely mesh-shaped is L7 authorisation, header-based traffic shifting and uniform golden-signal telemetry. If none of those is a requirement, adding a mesh buys complexity rather than capability. If they are, Istio ambient or Cilium's own mesh are the current answers, and a namespace-scoped rollout keeps the decision reversible.

Get the first one right before the cluster exists. Migrate the second on a schedule you control rather than one a CVE chooses for you. Defer the third until you can name the specific capability you are buying.

If you are standing up clusters, planning an ingress-nginx migration, or trying to work out whether a mesh is worth the operational cost, our managed Kubernetes service and 24×7 SRE team do this work continuously. Related reading: our Kubernetes production readiness checklist and EKS vs GKE vs AKS comparison, which covers how the managed platforms differ on networking defaults.

Talk to our team → We will start with what your CNI is and whether your NetworkPolicies are actually being enforced — the answer surprises people more often than it should.