Kubernetes resource quotas exist to make that impossible, and they only work when paired with LimitRange, which supplies the per-pod values quotas depend on.

The two objects are constantly confused. One caps a namespace total; the other governs individual containers and fills in what developers omit. This guide covers what each actually enforces, the order they run in, and the specific behaviours that produce confusing failures.

The Problem: Noisy Neighbours on Shared Clusters

Kubernetes defaults are permissive

A pod with no resources block is not constrained in any way. It requests nothing, so the scheduler treats it as free to place, and it is limited to nothing, so it can consume whatever the node has.

That produces three failure modes on any shared cluster:

  • Memory exhaustion. One container leaks, the node runs out, and the kubelet evicts. It does not evict the offender preferentially — it evicts by QoS class.
  • CPU starvation. A pod with no CPU request gets the minimum share under contention. Latency-sensitive services degrade while a batch job saturates the cores.
  • Scheduling distortion. Pods without requests appear free to the scheduler, so it packs them densely onto nodes that are actually full.

QoS classes decide who dies first

Kubernetes assigns every pod one of three quality-of-service classes, derived entirely from its requests and limits:

QoS classConditionEviction order
GuaranteedEvery container has requests equal to limits, for both CPU and memoryLast
BurstableAt least one container has a request, but not requests equal to limitsMiddle
BestEffortNo requests or limits anywhere in the podFirst

This is the mechanism behind the opening scenario. The payments API had no resource block, which made it BestEffort, which made it the kubelet's first choice under memory pressure. The team that omitted the values believed they were being flexible; they had, in fact, volunteered for eviction.

LimitRange exists to make this failure impossible by ensuring no pod is ever accidentally BestEffort.

What each object solves

ObjectScopeQuestion it answers
LimitRangeIndividual containers and PVCs"Is this one object reasonably sized, and what happens if the developer said nothing?"
ResourceQuotaThe namespace as a whole"Has this team used more than its allocation?"

You need both. A quota without a LimitRange rejects every pod that omits requests. A LimitRange without a quota constrains individual pods but lets a team run ten thousand of them.

What ResourceQuota Controls and How It Is Enforced

What it can cap

ResourceQuota is namespace-scoped and enforces aggregate limits across three categories.

Compute and memory:

apiVersion: v1
kind: ResourceQuota
metadata:
  name: team-alpha
  namespace: team-alpha
spec:
  hard:
    requests.cpu: "40"
    requests.memory: 80Gi
    limits.cpu: "80"
    limits.memory: 160Gi

Storage:

    requests.storage: 2Ti
    persistentvolumeclaims: "50"
    gold.storageclass.storage.k8s.io/requests.storage: 500Gi

Object counts — the one most teams forget, and the one that prevents a runaway controller from filling etcd:

    pods: "200"
    services: "30"
    services.loadbalancers: "4"
    secrets: "100"
    configmaps: "100"
    count/deployments.apps: "50"

Capping services.loadbalancers is worth calling out separately. Each one provisions a cloud load balancer with a monthly bill attached, and it is a common source of unexplained cost growth — the kind of thing that surfaces in a cloud cost review months later.

How enforcement actually works

ResourceQuota is an admission controller. When a create or update request arrives, it computes what the namespace total would become and rejects the request with 403 Forbidden if that exceeds the hard limit.

Three consequences follow, and all three surprise people.

1. Enabling a compute quota makes requests mandatory. Once requests.cpu or requests.memory appears in a quota, every pod in that namespace must specify the corresponding value. Pods that don't are rejected outright. This is precisely why LimitRange has to be applied first — otherwise you break every existing deployment the moment the quota lands.

2. Rejections are invisible if you only look at pods. A Deployment does not create pods directly; its ReplicaSet does. When quota rejects the pod, the error appears on the ReplicaSet, not on the Deployment and not as a Pending pod. The symptom is a Deployment stuck at 2/5 replicas with no obvious reason.

kubectl describe replicaset -n team-alpha | grep -A5 Events
kubectl get events -n team-alpha --field-selector reason=FailedCreate

This is the single most common quota support ticket, and the fix is knowing where to look.

3. Quota is not capacity. It governs admission, not scheduling. A pod can pass quota and still sit Pending because no node has room. Conversely, the sum of all namespace quotas can — and usually should — exceed the cluster's actual capacity, because no team runs at its ceiling simultaneously.

Checking consumption

kubectl describe resourcequota team-alpha -n team-alpha

This prints used versus hard for every tracked resource and is the first command to run when a team reports that deployments have stopped working.

Scopes: quotas that apply selectively

Scopes let one namespace carry different rules for different workload classes.

ScopeMatches
TerminatingPods with activeDeadlineSeconds set
NotTerminatingLong-running pods
BestEffortPods with no requests or limits
NotBestEffortEverything else
PriorityClassPods with a matching priority class

The PriorityClass scope is the most useful in practice. It lets you give a team a generous allocation for normal workloads and a small, separate allocation for anything claiming high priority — which stops teams from escalating their way out of the constraint.

spec:
  hard:
    requests.cpu: "10"
  scopeSelector:
    matchExpressions:
    - operator: In
      scopeName: PriorityClass
      values: ["high-priority"]

Setting the numbers

Rather than guessing:

  1. Measure 30 days of actual usage per namespace, taking the peak rather than the average.
  2. Set requests quota at roughly 1.5× observed peak — enough for growth without inviting waste.
  3. Set limits quota at roughly 2× the requests quota, permitting reasonable burst.
  4. Set object counts at 3–5× current, which stops runaway loops without blocking normal work.
  5. Review quarterly and adjust.

Deliberately over-subscribe. The sum of namespace quotas should exceed cluster capacity, because teams do not all peak together. Sizing quotas to exactly match the cluster wastes a large fraction of what you paid for — the same headroom logic that governs capacity planning.

What LimitRange Controls and Why Defaults Matter

The four things it does

LimitRange operates on individual objects within a namespace, and it is the only one of the two that can change your spec.

apiVersion: v1
kind: LimitRange
metadata:
  name: team-alpha-limits
  namespace: team-alpha
spec:
  limits:
  - type: Container
    default:                  # limits, if unspecified
      cpu: 500m
      memory: 512Mi
    defaultRequest:           # requests, if unspecified
      cpu: 100m
      memory: 128Mi
    min:                      # floor per container
      cpu: 50m
      memory: 64Mi
    max:                      # ceiling per container
      cpu: "4"
      memory: 8Gi
    maxLimitRequestRatio:     # burst ceiling
      cpu: 10
      memory: 4

Four distinct behaviours, easily conflated:

  • default sets limits when the container specifies none
  • defaultRequest sets requests when the container specifies none
  • min / max reject containers outside the permitted range
  • maxLimitRequestRatio caps how far a container may burst above its request

maxLimitRequestRatio is underused and valuable. A container requesting 100m CPU with a 4-core limit looks cheap to the scheduler and behaves expensively on the node. A ratio of 10 keeps the gap honest.

Why defaults are the important part

The defaultRequest field is what guarantees no pod in the namespace is ever BestEffort. Every container gets requests whether the developer thought about it or not, which means every pod is at least Burstable and no longer first in the eviction queue.

This also has a practical politics benefit. Rolling out quotas without LimitRange means every team must add resource blocks before anything deploys. Rolling out LimitRange first means existing manifests keep working, and teams tune values when they care to.

The catch that causes confusion

Defaults apply at object creation only. Changing a LimitRange does not retroactively modify running pods. Existing pods keep the values they were admitted with until they are recreated.

The practical consequence: after updating a LimitRange, nothing changes until the next rollout. Teams report that "the new defaults aren't working" when in fact they will apply to the next deployed pod and no earlier.

Other types LimitRange supports

Beyond Container, two more types are useful:

  • Pod — caps the total across all containers in a pod, including sidecars. Useful where a mesh or logging sidecar would otherwise push a pod past what you intended.
  • PersistentVolumeClaim — enforces min and max PVC size, preventing a 10 TiB claim from an accidental zero.
  - type: PersistentVolumeClaim
    min:
      storage: 1Gi
    max:
      storage: 500Gi

The order they run in

The sequence matters and explains most confusing errors:

  1. LimitRange mutating admission — missing requests and limits are filled in from defaults
  2. LimitRange validating admission — min, max and ratio are checked; violations are rejected here
  3. ResourceQuota admission — the namespace total is recalculated; the pod is rejected if it would exceed
  4. Scheduling — only now does the scheduler look for a node with room

A rejection at step 2 never reaches the quota check, so the error message mentions LimitRange rather than quota even though the team was near its quota. And a pod that clears all three can still sit Pending, because step 4 is a separate problem entirely.

A sensible rollout order

Applying these to a live cluster in the wrong order breaks deployments. The safe sequence:

  1. Measure first. Collect per-namespace usage for 30 days before writing any numbers.
  2. Apply LimitRange with generous defaults. Nothing breaks; pods that specified nothing start getting reasonable values on their next rollout.
  3. Wait for a deployment cycle. Let existing pods be recreated so they pick up the defaults.
  4. Apply quota in a non-production namespace first. Watch for FailedCreate events.
  5. Roll quota to production namespaces, starting with the least critical.
  6. Alert on quota utilisation above 80%, so teams get warning before deployments start failing.

That last point is worth doing properly. kube_resourcequota metrics from kube-state-metrics give you used-versus-hard per namespace, and an alert at 80% converts a confusing outage into a planned conversation. It is a small addition to any monitoring and observability setup and removes most of the support burden these objects create.

What do quotas not do?

Worth stating explicitly, because teams assume otherwise:

  • They do not isolate the network. That is NetworkPolicy.
  • They do not stop a pod exceeding its limit on a node. That is cgroup enforcement via limits.
  • They do not provide security isolation. A namespace with a quota is still a soft tenancy boundary. Hard multi-tenancy needs separate clusters.
  • They do not guarantee capacity. A team within quota can still fail to schedule if the cluster is full.

Conclusion

ResourceQuota and LimitRange are two halves of one control, and using either alone produces a predictable failure.

LimitRange is the one that prevents outages. Its most valuable field is defaultRequest, because it guarantees no pod in the namespace is accidentally BestEffort and therefore first in line when the kubelet starts evicting. Add min, max and maxLimitRequestRatio to keep individual containers sensible, and a PersistentVolumeClaim type to stop mis-sized storage claims. Apply it first, always — it is non-breaking, and it prepares every manifest in the namespace for the quota that follows.

ResourceQuota is the one that enforces fairness. Cap compute, storage and — the part teams forget — object counts, especially services.loadbalancers, where each unit carries a monthly bill. Size it from 30 days of measured peak rather than a guess, at roughly 1.5× for requests and 2× that for limits, and deliberately over-subscribe across namespaces, because no two teams peak at the same moment.

Most of the pain comes from three behaviours. A compute quota silently makes requests mandatory, so LimitRange must land first. Quota rejections appear on the ReplicaSet rather than the Deployment, which is why a stuck rollout looks like a mystery until you check FailedCreate events. And quota is admission, not capacity — passing it does not mean a node has room, and the sum of quotas exceeding cluster size is correct rather than a bug.

Get the rollout order right, alert at 80% utilisation, and review the numbers quarterly. Done that way, these objects stop being a source of support tickets and become the thing that makes a shared cluster survivable.

If you are setting up multi-tenancy, working out quota numbers for existing teams, or dealing with clusters where one namespace keeps affecting the others, our managed Kubernetes service and 24×7 SRE team do this work as routine. Related reading: the Kubernetes production readiness checklist and our capacity planning guide, which covers how to derive the numbers these objects enforce.

Talk to our team → We will start with 30 days of per-namespace usage, because quota numbers pulled from a template are the ones that cause incidents.