Kubernetes resource quotas exist to make that impossible, and they only work when paired with LimitRange, which supplies the per-pod values quotas depend on.
The two objects are constantly confused. One caps a namespace total; the other governs individual containers and fills in what developers omit. This guide covers what each actually enforces, the order they run in, and the specific behaviours that produce confusing failures.
The Problem: Noisy Neighbours on Shared Clusters
Kubernetes defaults are permissive
A pod with no resources block is not constrained in any way. It requests nothing, so the scheduler treats it as free to place, and it is limited to nothing, so it can consume whatever the node has.
That produces three failure modes on any shared cluster:
- Memory exhaustion. One container leaks, the node runs out, and the kubelet evicts. It does not evict the offender preferentially — it evicts by QoS class.
- CPU starvation. A pod with no CPU request gets the minimum share under contention. Latency-sensitive services degrade while a batch job saturates the cores.
- Scheduling distortion. Pods without requests appear free to the scheduler, so it packs them densely onto nodes that are actually full.
QoS classes decide who dies first
Kubernetes assigns every pod one of three quality-of-service classes, derived entirely from its requests and limits:
| QoS class | Condition | Eviction order |
|---|---|---|
| Guaranteed | Every container has requests equal to limits, for both CPU and memory | Last |
| Burstable | At least one container has a request, but not requests equal to limits | Middle |
| BestEffort | No requests or limits anywhere in the pod | First |
This is the mechanism behind the opening scenario. The payments API had no resource block, which made it BestEffort, which made it the kubelet's first choice under memory pressure. The team that omitted the values believed they were being flexible; they had, in fact, volunteered for eviction.
LimitRange exists to make this failure impossible by ensuring no pod is ever accidentally BestEffort.
What each object solves
| Object | Scope | Question it answers |
|---|---|---|
| LimitRange | Individual containers and PVCs | "Is this one object reasonably sized, and what happens if the developer said nothing?" |
| ResourceQuota | The namespace as a whole | "Has this team used more than its allocation?" |
You need both. A quota without a LimitRange rejects every pod that omits requests. A LimitRange without a quota constrains individual pods but lets a team run ten thousand of them.
What ResourceQuota Controls and How It Is Enforced
What it can cap
ResourceQuota is namespace-scoped and enforces aggregate limits across three categories.
Compute and memory:
apiVersion: v1
kind: ResourceQuota
metadata:
name: team-alpha
namespace: team-alpha
spec:
hard:
requests.cpu: "40"
requests.memory: 80Gi
limits.cpu: "80"
limits.memory: 160GiStorage:
requests.storage: 2Ti
persistentvolumeclaims: "50"
gold.storageclass.storage.k8s.io/requests.storage: 500GiObject counts — the one most teams forget, and the one that prevents a runaway controller from filling etcd:
pods: "200"
services: "30"
services.loadbalancers: "4"
secrets: "100"
configmaps: "100"
count/deployments.apps: "50"Capping services.loadbalancers is worth calling out separately. Each one provisions a cloud load balancer with a monthly bill attached, and it is a common source of unexplained cost growth — the kind of thing that surfaces in a cloud cost review months later.
How enforcement actually works
ResourceQuota is an admission controller. When a create or update request arrives, it computes what the namespace total would become and rejects the request with 403 Forbidden if that exceeds the hard limit.
Three consequences follow, and all three surprise people.
1. Enabling a compute quota makes requests mandatory. Once requests.cpu or requests.memory appears in a quota, every pod in that namespace must specify the corresponding value. Pods that don't are rejected outright. This is precisely why LimitRange has to be applied first — otherwise you break every existing deployment the moment the quota lands.
2. Rejections are invisible if you only look at pods. A Deployment does not create pods directly; its ReplicaSet does. When quota rejects the pod, the error appears on the ReplicaSet, not on the Deployment and not as a Pending pod. The symptom is a Deployment stuck at 2/5 replicas with no obvious reason.
kubectl describe replicaset -n team-alpha | grep -A5 Events
kubectl get events -n team-alpha --field-selector reason=FailedCreateThis is the single most common quota support ticket, and the fix is knowing where to look.
3. Quota is not capacity. It governs admission, not scheduling. A pod can pass quota and still sit Pending because no node has room. Conversely, the sum of all namespace quotas can — and usually should — exceed the cluster's actual capacity, because no team runs at its ceiling simultaneously.
Checking consumption
kubectl describe resourcequota team-alpha -n team-alphaThis prints used versus hard for every tracked resource and is the first command to run when a team reports that deployments have stopped working.
Scopes: quotas that apply selectively
Scopes let one namespace carry different rules for different workload classes.
| Scope | Matches |
|---|---|
Terminating | Pods with activeDeadlineSeconds set |
NotTerminating | Long-running pods |
BestEffort | Pods with no requests or limits |
NotBestEffort | Everything else |
PriorityClass | Pods with a matching priority class |
The PriorityClass scope is the most useful in practice. It lets you give a team a generous allocation for normal workloads and a small, separate allocation for anything claiming high priority — which stops teams from escalating their way out of the constraint.
spec:
hard:
requests.cpu: "10"
scopeSelector:
matchExpressions:
- operator: In
scopeName: PriorityClass
values: ["high-priority"]Setting the numbers
Rather than guessing:
- Measure 30 days of actual usage per namespace, taking the peak rather than the average.
- Set
requestsquota at roughly 1.5× observed peak — enough for growth without inviting waste. - Set
limitsquota at roughly 2× the requests quota, permitting reasonable burst. - Set object counts at 3–5× current, which stops runaway loops without blocking normal work.
- Review quarterly and adjust.
Deliberately over-subscribe. The sum of namespace quotas should exceed cluster capacity, because teams do not all peak together. Sizing quotas to exactly match the cluster wastes a large fraction of what you paid for — the same headroom logic that governs capacity planning.
What LimitRange Controls and Why Defaults Matter
The four things it does
LimitRange operates on individual objects within a namespace, and it is the only one of the two that can change your spec.
apiVersion: v1
kind: LimitRange
metadata:
name: team-alpha-limits
namespace: team-alpha
spec:
limits:
- type: Container
default: # limits, if unspecified
cpu: 500m
memory: 512Mi
defaultRequest: # requests, if unspecified
cpu: 100m
memory: 128Mi
min: # floor per container
cpu: 50m
memory: 64Mi
max: # ceiling per container
cpu: "4"
memory: 8Gi
maxLimitRequestRatio: # burst ceiling
cpu: 10
memory: 4Four distinct behaviours, easily conflated:
defaultsets limits when the container specifies nonedefaultRequestsets requests when the container specifies nonemin/maxreject containers outside the permitted rangemaxLimitRequestRatiocaps how far a container may burst above its request
maxLimitRequestRatio is underused and valuable. A container requesting 100m CPU with a 4-core limit looks cheap to the scheduler and behaves expensively on the node. A ratio of 10 keeps the gap honest.
Why defaults are the important part
The defaultRequest field is what guarantees no pod in the namespace is ever BestEffort. Every container gets requests whether the developer thought about it or not, which means every pod is at least Burstable and no longer first in the eviction queue.
This also has a practical politics benefit. Rolling out quotas without LimitRange means every team must add resource blocks before anything deploys. Rolling out LimitRange first means existing manifests keep working, and teams tune values when they care to.
The catch that causes confusion
Defaults apply at object creation only. Changing a LimitRange does not retroactively modify running pods. Existing pods keep the values they were admitted with until they are recreated.
The practical consequence: after updating a LimitRange, nothing changes until the next rollout. Teams report that "the new defaults aren't working" when in fact they will apply to the next deployed pod and no earlier.
Other types LimitRange supports
Beyond Container, two more types are useful:
Pod— caps the total across all containers in a pod, including sidecars. Useful where a mesh or logging sidecar would otherwise push a pod past what you intended.PersistentVolumeClaim— enforces min and max PVC size, preventing a 10 TiB claim from an accidental zero.
- type: PersistentVolumeClaim
min:
storage: 1Gi
max:
storage: 500GiThe order they run in
The sequence matters and explains most confusing errors:
- LimitRange mutating admission — missing requests and limits are filled in from defaults
- LimitRange validating admission — min, max and ratio are checked; violations are rejected here
- ResourceQuota admission — the namespace total is recalculated; the pod is rejected if it would exceed
- Scheduling — only now does the scheduler look for a node with room
A rejection at step 2 never reaches the quota check, so the error message mentions LimitRange rather than quota even though the team was near its quota. And a pod that clears all three can still sit Pending, because step 4 is a separate problem entirely.
A sensible rollout order
Applying these to a live cluster in the wrong order breaks deployments. The safe sequence:
- Measure first. Collect per-namespace usage for 30 days before writing any numbers.
- Apply LimitRange with generous defaults. Nothing breaks; pods that specified nothing start getting reasonable values on their next rollout.
- Wait for a deployment cycle. Let existing pods be recreated so they pick up the defaults.
- Apply quota in a non-production namespace first. Watch for
FailedCreateevents. - Roll quota to production namespaces, starting with the least critical.
- Alert on quota utilisation above 80%, so teams get warning before deployments start failing.
That last point is worth doing properly. kube_resourcequota metrics from kube-state-metrics give you used-versus-hard per namespace, and an alert at 80% converts a confusing outage into a planned conversation. It is a small addition to any monitoring and observability setup and removes most of the support burden these objects create.
What do quotas not do?
Worth stating explicitly, because teams assume otherwise:
- They do not isolate the network. That is NetworkPolicy.
- They do not stop a pod exceeding its limit on a node. That is cgroup enforcement via limits.
- They do not provide security isolation. A namespace with a quota is still a soft tenancy boundary. Hard multi-tenancy needs separate clusters.
- They do not guarantee capacity. A team within quota can still fail to schedule if the cluster is full.
Conclusion
ResourceQuota and LimitRange are two halves of one control, and using either alone produces a predictable failure.
LimitRange is the one that prevents outages. Its most valuable field is defaultRequest, because it guarantees no pod in the namespace is accidentally BestEffort and therefore first in line when the kubelet starts evicting. Add min, max and maxLimitRequestRatio to keep individual containers sensible, and a PersistentVolumeClaim type to stop mis-sized storage claims. Apply it first, always — it is non-breaking, and it prepares every manifest in the namespace for the quota that follows.
ResourceQuota is the one that enforces fairness. Cap compute, storage and — the part teams forget — object counts, especially services.loadbalancers, where each unit carries a monthly bill. Size it from 30 days of measured peak rather than a guess, at roughly 1.5× for requests and 2× that for limits, and deliberately over-subscribe across namespaces, because no two teams peak at the same moment.
Most of the pain comes from three behaviours. A compute quota silently makes requests mandatory, so LimitRange must land first. Quota rejections appear on the ReplicaSet rather than the Deployment, which is why a stuck rollout looks like a mystery until you check FailedCreate events. And quota is admission, not capacity — passing it does not mean a node has room, and the sum of quotas exceeding cluster size is correct rather than a bug.
Get the rollout order right, alert at 80% utilisation, and review the numbers quarterly. Done that way, these objects stop being a source of support tickets and become the thing that makes a shared cluster survivable.
If you are setting up multi-tenancy, working out quota numbers for existing teams, or dealing with clusters where one namespace keeps affecting the others, our managed Kubernetes service and 24×7 SRE team do this work as routine. Related reading: the Kubernetes production readiness checklist and our capacity planning guide, which covers how to derive the numbers these objects enforce.
Talk to our team → We will start with 30 days of per-namespace usage, because quota numbers pulled from a template are the ones that cause incidents.