Policy as code turns those rules into executable artifacts that are version-controlled, tested, and evaluated automatically at the moment a change is made. Open Policy Agent is the general-purpose engine most teams reach for, because the same policy can run in a developer's terminal, in CI, and as a Kubernetes admission controller.
This guide covers how OPA actually works, where to put enforcement, and — the part that determines whether the rollout succeeds — how to switch it on without breaking everyone's deployments.
What Policy as Code Solves That Manual Review Doesn't
Manual review fails in four specific ways
Not because reviewers are careless, but because the review process has structural limits:
- It is inconsistent. The same misconfiguration gets caught on Tuesday and missed on Thursday, depending on who reviewed it and how much else was in the diff.
- It does not scale. Fifty pull requests a day across twelve teams is not a review problem; it is an arithmetic problem.
- It produces no evidence. "We review for this" is not an answer an auditor accepts. "Here is the policy, its test suite, its git history, and the log of every decision it made" is.
- It only sees what goes through review. A change applied directly with
kubectlor from a cloud console never meets a reviewer at all.
That last one matters more than it looks. A pull request gate is a control on one path into your infrastructure, not on all of them.
What executable policy gives you instead
| Manual review | Policy as code |
|---|---|
| Inconsistent by person and by day | Identical every time |
| Scales with headcount | Scales with compute |
| Feedback in hours or days | Feedback in seconds |
| No audit trail | Git history plus decision logs |
| Covers the review path only | Covers every path you instrument |
| Rules live in a wiki | Rules live in a tested repository |
The shift that matters most is from documentation to enforcement. A rule in a Confluence page is a suggestion. A rule that fails a build is a control.
What it does not solve
Worth being honest about, because over-selling it causes the rollout to fail:
- It will not catch design problems. Policy checks structure, not judgement. It can enforce that a database is encrypted; it cannot tell you the architecture is wrong.
- It will not replace security review for genuinely novel changes.
- Badly written policy is worse than none. A rule that fires on legitimate configurations teaches people to request exceptions, and exception-granting quickly becomes the whole job.
- It needs an owner. Policies decay as platforms change. An unmaintained policy corpus generates false positives until someone disables it entirely.
How Open Policy Agent Works: Rego, Data and Decisions
The model: three inputs, one decision
OPA is a general-purpose policy engine, and its model is simple enough to state in one sentence: you give it a query, some input, and some data; it returns a decision.
- Input — the thing being evaluated, as JSON. A Kubernetes admission request, a Terraform plan, an HTTP request, a CI artifact.
- Data — reference information the policy needs. A list of approved registries, a mapping of teams to cost centres, an allowlist of regions.
- Policy — rules written in Rego that produce a decision from input and data.
OPA has no opinion about what it is evaluating. That is why the same engine handles Kubernetes manifests, Terraform plans and microservice authorisation — a property no Kubernetes-specific tool has.
Rego, briefly
Rego is declarative. You describe what a violation looks like, and OPA finds all the ways the input matches.
package kubernetes.security
# Containers must not run as root
deny contains msg if {
input.request.kind.kind == "Pod"
container := input.request.object.spec.containers[_]
not container.securityContext.runAsNonRoot
msg := sprintf("container '%s' must set runAsNonRoot", [container.name])
}Three things to notice, because they are where newcomers get stuck:
denyis a set. Every container that violates the rule adds a message. An empty set means no violations.[_]iterates. The rule evaluates once per container, automatically.nothandles absence. A missingsecurityContextand an explicitrunAsNonRoot: falseboth fail.
A syntax warning worth heeding: OPA 1.0 made Rego v1 the default, which requires the if keyword before a rule body and contains when declaring a multi-value rule. A large share of the Rego examples still on the internet were written for the older syntax and will not compile. If a snippet you found omits if, it predates the change.
Using external data
Policies that need context load it as data rather than hard-coding it:
package kubernetes.images
deny contains msg if {
container := input.request.object.spec.containers[_]
not startswith(container.image, data.registries.approved[_])
msg := sprintf("image '%s' is not from an approved registry", [container.image])
}data.registries.approved is loaded separately — from a bundle, an API, or a ConfigMap. The policy stays stable while the allowlist changes independently, which means updating the registry list is a data change rather than a code review.
Policies are code, so test them
This is the discipline that separates a policy corpus that survives from one that gets disabled after six months.
package kubernetes.security_test
test_denies_root_container if {
result := deny with input as {
"request": {
"kind": {"kind": "Pod"},
"object": {"spec": {"containers": [{"name": "app"}]}}
}
}
count(result) == 1
}Run with opa test. Every policy needs at least two tests — one that fires and one that does not — and both belong in CI. Without the negative test you will not notice when a refactor makes the rule match everything.
Bundles: how policy gets distributed
OPA loads policy from bundles — signed tarballs served over HTTP, pulled on a schedule. The practical consequence is a clean lifecycle: policies live in git, CI tests and builds them into a bundle, OPA instances pull the new version without redeployment.
Combined with decision logs — a record of every evaluation, its input and its result — this gives you the audit evidence that manual review never produced.
Deployment Patterns: Admission Control and CI/CD Gates
The same rule can be enforced in four places. They differ in how fast the feedback is and how easy they are to avoid.
1. Local — fastest feedback, no control value
conftest runs Rego against YAML, JSON and HCL from the command line:
conftest test deployment.yaml --policy ./policiesWire it into a pre-commit hook and developers find violations before pushing. Point it at the same policy repository CI uses, so there is one source of truth.
Treat this as developer experience, not enforcement. It is trivially skipped, and that is fine — its job is to stop people wasting a CI cycle.
2. CI pipeline — blocks the merge
The most valuable single integration for most teams, because it catches infrastructure changes before they exist.
terraform plan -out=tfplan
terraform show -json tfplan > plan.json
conftest test plan.json --policy ./policies/terraformEvaluating the plan rather than the HCL is the important detail. The plan contains fully resolved values after variables, modules and data sources are evaluated, so policy sees what will actually be created.
Surface failures as PR comments with the specific resource and the specific rule. A pipeline that fails with "policy violation" and no detail generates a support ticket rather than a fix.
The limitation: CI only sees changes that go through CI. Anything applied by hand is invisible to it.
3. Admission control — the one that cannot be routed around
Every change to a Kubernetes cluster passes through the API server, which makes admission the only enforcement point that sees everything — pipeline deployments, kubectl apply, cloud console edits, and controllers acting on their own.
OPA Gatekeeper brings OPA to Kubernetes through a two-level Constraint Framework:
- A ConstraintTemplate is a CRD defining a parameterised policy written in Rego
- A Constraint is an instance of that template with specific parameters and scope
apiVersion: templates.gatekeeper.sh/v1
kind: ConstraintTemplate
metadata:
name: k8srequiredlabels
spec:
crd:
spec:
names:
kind: K8sRequiredLabels
validation:
openAPIV3Schema:
properties:
labels:
type: array
items: {type: string}
targets:
- target: admission.k8s.gatekeeper.sh
rego: |
package k8srequiredlabels
violation contains {"msg": msg} if {
required := input.parameters.labels[_]
not input.review.object.metadata.labels[required]
msg := sprintf("missing required label: %v", [required])
}
apiVersion: constraints.gatekeeper.sh/v1beta1
kind: K8sRequiredLabels
metadata:
name: require-cost-centre
spec:
enforcementAction: dryrun
match:
kinds:
- apiGroups: ["apps"]
kinds: ["Deployment"]
parameters:
labels: ["cost-centre"]The separation is the point. Platform engineers write templates; the same template is instantiated many times with different parameters and different scopes, without anyone touching Rego again.
4. Audit — what is already running
Gatekeeper periodically evaluates existing cluster resources against constraints and records violations in each Constraint's status. This finds what was deployed before the policy existed — which, on any real cluster, is a lot.
Audit reports; it does not prevent. Its value is knowing the size of the remediation job before you switch anything to deny.
Where Gatekeeper still earns its place
Kubernetes has closed much of the gap. ValidatingAdmissionPolicy has been GA since v1.30 — CEL-based rules evaluated inside the API server with no webhook, no network hop and no external component to keep alive. MutatingAdmissionPolicy reached stable in v1.36, so mutation no longer requires a webhook either.
For a simple, high-volume rule, native admission policy is now the better default. It is one less thing to operate and it cannot fail open.
Gatekeeper remains the right answer when you need:
- Logic beyond CEL's expressiveness — complex evaluation that Rego handles cleanly
- Existing-resource evaluation via audit
- External data in policy decisions
- One engine across Kubernetes, Terraform and CI, which is OPA's genuine differentiator
- An established OPA investment you are not going to unwind
On Kyverno
Kyverno is the main alternative and the fair comparison is ergonomic. Kyverno policies are Kubernetes-native YAML, which most teams find easier to write and review than Rego. It writes results as PolicyReport CRDs, an open Policy Working Group format, and its newer CEL-based ValidatingPolicy types align closely with native admission policy.
The trade-off: Kyverno is Kubernetes-only. If you want one policy engine covering your clusters, your Terraform and your CI pipelines, OPA is the one that spans all three. If Kubernetes is the whole problem, Kyverno is usually the more pleasant tool.
Either way, note that Kyverno's legacy kyverno.io/v1 ClusterPolicy API was deprecated in 1.17 and is scheduled for removal — write new policies against the current types.
Rolling Out Enforcement: Audit Mode, Warnings, Then Blocking
This section determines whether the initiative survives. Switching policies to deny on a live cluster is the most reliable way to have the whole programme abandoned.
The ladder
Gatekeeper's enforcementAction gives you three rungs. Use all of them, weeks apart.
| Stage | Action | What happens | Duration |
|---|---|---|---|
| 1 | dryrun | Violation recorded in Constraint status; request succeeds | 2–4 weeks |
| 2 | warn | Developer sees the message; request still succeeds | 2–4 weeks |
| 3 | deny | Request rejected | Permanent |
The equivalent in native admission policy is the validationActions field on a PolicyBinding, which takes Audit, Warn and Deny and can combine them.
Stage 1 — dryrun
Deploy the policy with enforcementAction: dryrun and leave it alone.
What you are looking for:
- How many violations exist. This is the remediation backlog. It is always larger than expected.
- Which teams are affected. Ten teams with one violation each is a different conversation from one team with two hundred.
- Whether the rule is even correct. False positives surface here, where they cost nothing.
kubectl get constraint require-cost-centre -o jsonpath='{.status.totalViolations}'Run it for at least a full sprint. Weekly deployment cycles, cron jobs and quarterly batch processes all need to pass through the policy before you trust the count.
Stage 2 — warn
Switch to warn. Deployments still succeed, but the person applying sees the message immediately.
This stage is doing two things at once. It remediates the backlog through normal work, because developers fix what they can see. And it stress-tests your error messages: if a warning does not tell someone exactly what to change, you find out now rather than after enforcement.
Do not move on until violations are approaching zero and the trend is downward.
Stage 3 — deny
Only when the violation count is zero or every remaining case has a documented exception.
Two things to have in place first:
- An exception mechanism. Gatekeeper's
matchblock excludes namespaces or label selectors. Exceptions should be time-bounded, recorded in git, and reviewed — not permanent and not verbal. - A break-glass procedure. If a policy blocks a production fix during an incident, there must be a documented way to bypass it that leaves a trail. Without one, someone will delete the constraint at 3am and nobody will remember to restore it.
Practical guidance for the rollout
- Non-production first. Prove the policy in dev and staging before production sees it.
- One policy at a time. Ten policies deployed together produce ten simultaneous conversations and no clear signal.
- Start with the ones nobody argues about. Required labels, no
:latestimage tags, no privileged containers. Save the contentious rules until the mechanism has earned trust. failurePolicy: Ignoreat first. If the Gatekeeper webhook is unavailable, do you want deployments to fail? Start withIgnore, move toFailonce it has proven stable — and make sure Gatekeeper's own namespace is exempt, or a restart can deadlock the cluster.- Exempt system namespaces from the start.
kube-systemand your CNI's namespace are not where you want to discover a policy interaction.
Write error messages people can act on
The single highest-leverage thing you can do for adoption.
Poor: policy violation: k8srequiredlabels
Good: Deployment 'checkout-api' is missing required label 'cost-centre'. Add it under metadata.labels. See https://wiki/policies/cost-centre for approved values.
The second version is fixed by the developer in thirty seconds. The first becomes a Slack message to the platform team.
Operate it as a product
Policies are code, so they need the same lifecycle: version-controlled in one repository, tested with opa test in CI, delivered through a pipeline rather than applied by hand, and reviewed quarterly for rules that no longer apply.
Track two numbers. Violation count over time should fall after each rollout. Exception count should stay flat — a rising exception count means the policy is wrong, not that the teams are.
This is the operational core of any serious DevSecOps practice: controls that run automatically, fail loudly and produce their own evidence. It pairs naturally with GitOps delivery, where the policy repository and the manifests it governs move through the same pipeline, and with the controls covered in our Kubernetes production readiness checklist.
If you are introducing policy as code, have a Gatekeeper deployment stuck in audit mode that nobody trusts enough to enforce, or need to decide between OPA, Kyverno and native admission policy for your estate, our DevSecOps services and 24×7 SRE team cover this work — policy design, test suites, rollout sequencing and the exception process that keeps it maintainable.
Talk to our team → We will start by running your candidate policies in dry-run, because the violation count is the only honest starting point.