Policy as code turns those rules into executable artifacts that are version-controlled, tested, and evaluated automatically at the moment a change is made. Open Policy Agent is the general-purpose engine most teams reach for, because the same policy can run in a developer's terminal, in CI, and as a Kubernetes admission controller.

This guide covers how OPA actually works, where to put enforcement, and — the part that determines whether the rollout succeeds — how to switch it on without breaking everyone's deployments.

What Policy as Code Solves That Manual Review Doesn't

Manual review fails in four specific ways

Not because reviewers are careless, but because the review process has structural limits:

  • It is inconsistent. The same misconfiguration gets caught on Tuesday and missed on Thursday, depending on who reviewed it and how much else was in the diff.
  • It does not scale. Fifty pull requests a day across twelve teams is not a review problem; it is an arithmetic problem.
  • It produces no evidence. "We review for this" is not an answer an auditor accepts. "Here is the policy, its test suite, its git history, and the log of every decision it made" is.
  • It only sees what goes through review. A change applied directly with kubectl or from a cloud console never meets a reviewer at all.

That last one matters more than it looks. A pull request gate is a control on one path into your infrastructure, not on all of them.

What executable policy gives you instead

Manual reviewPolicy as code
Inconsistent by person and by dayIdentical every time
Scales with headcountScales with compute
Feedback in hours or daysFeedback in seconds
No audit trailGit history plus decision logs
Covers the review path onlyCovers every path you instrument
Rules live in a wikiRules live in a tested repository

The shift that matters most is from documentation to enforcement. A rule in a Confluence page is a suggestion. A rule that fails a build is a control.

What it does not solve

Worth being honest about, because over-selling it causes the rollout to fail:

  • It will not catch design problems. Policy checks structure, not judgement. It can enforce that a database is encrypted; it cannot tell you the architecture is wrong.
  • It will not replace security review for genuinely novel changes.
  • Badly written policy is worse than none. A rule that fires on legitimate configurations teaches people to request exceptions, and exception-granting quickly becomes the whole job.
  • It needs an owner. Policies decay as platforms change. An unmaintained policy corpus generates false positives until someone disables it entirely.

How Open Policy Agent Works: Rego, Data and Decisions

The model: three inputs, one decision

OPA is a general-purpose policy engine, and its model is simple enough to state in one sentence: you give it a query, some input, and some data; it returns a decision.

  • Input — the thing being evaluated, as JSON. A Kubernetes admission request, a Terraform plan, an HTTP request, a CI artifact.
  • Data — reference information the policy needs. A list of approved registries, a mapping of teams to cost centres, an allowlist of regions.
  • Policy — rules written in Rego that produce a decision from input and data.

OPA has no opinion about what it is evaluating. That is why the same engine handles Kubernetes manifests, Terraform plans and microservice authorisation — a property no Kubernetes-specific tool has.

Rego, briefly

Rego is declarative. You describe what a violation looks like, and OPA finds all the ways the input matches.

package kubernetes.security

# Containers must not run as root
deny contains msg if {
    input.request.kind.kind == "Pod"
    container := input.request.object.spec.containers[_]
    not container.securityContext.runAsNonRoot
    msg := sprintf("container '%s' must set runAsNonRoot", [container.name])
}

Three things to notice, because they are where newcomers get stuck:

  • deny is a set. Every container that violates the rule adds a message. An empty set means no violations.
  • [_] iterates. The rule evaluates once per container, automatically.
  • not handles absence. A missing securityContext and an explicit runAsNonRoot: false both fail.

A syntax warning worth heeding: OPA 1.0 made Rego v1 the default, which requires the if keyword before a rule body and contains when declaring a multi-value rule. A large share of the Rego examples still on the internet were written for the older syntax and will not compile. If a snippet you found omits if, it predates the change.

Using external data

Policies that need context load it as data rather than hard-coding it:

package kubernetes.images

deny contains msg if {
    container := input.request.object.spec.containers[_]
    not startswith(container.image, data.registries.approved[_])
    msg := sprintf("image '%s' is not from an approved registry", [container.image])
}

data.registries.approved is loaded separately — from a bundle, an API, or a ConfigMap. The policy stays stable while the allowlist changes independently, which means updating the registry list is a data change rather than a code review.

Policies are code, so test them

This is the discipline that separates a policy corpus that survives from one that gets disabled after six months.

package kubernetes.security_test

test_denies_root_container if {
    result := deny with input as {
        "request": {
            "kind": {"kind": "Pod"},
            "object": {"spec": {"containers": [{"name": "app"}]}}
        }
    }
    count(result) == 1
}

Run with opa test. Every policy needs at least two tests — one that fires and one that does not — and both belong in CI. Without the negative test you will not notice when a refactor makes the rule match everything.

Bundles: how policy gets distributed

OPA loads policy from bundles — signed tarballs served over HTTP, pulled on a schedule. The practical consequence is a clean lifecycle: policies live in git, CI tests and builds them into a bundle, OPA instances pull the new version without redeployment.

Combined with decision logs — a record of every evaluation, its input and its result — this gives you the audit evidence that manual review never produced.

Deployment Patterns: Admission Control and CI/CD Gates

policy as a code enforcement

The same rule can be enforced in four places. They differ in how fast the feedback is and how easy they are to avoid.

1. Local — fastest feedback, no control value

conftest runs Rego against YAML, JSON and HCL from the command line:

conftest test deployment.yaml --policy ./policies

Wire it into a pre-commit hook and developers find violations before pushing. Point it at the same policy repository CI uses, so there is one source of truth.

Treat this as developer experience, not enforcement. It is trivially skipped, and that is fine — its job is to stop people wasting a CI cycle.

2. CI pipeline — blocks the merge

The most valuable single integration for most teams, because it catches infrastructure changes before they exist.

terraform plan -out=tfplan
terraform show -json tfplan > plan.json
conftest test plan.json --policy ./policies/terraform

Evaluating the plan rather than the HCL is the important detail. The plan contains fully resolved values after variables, modules and data sources are evaluated, so policy sees what will actually be created.

Surface failures as PR comments with the specific resource and the specific rule. A pipeline that fails with "policy violation" and no detail generates a support ticket rather than a fix.

The limitation: CI only sees changes that go through CI. Anything applied by hand is invisible to it.

3. Admission control — the one that cannot be routed around

Every change to a Kubernetes cluster passes through the API server, which makes admission the only enforcement point that sees everything — pipeline deployments, kubectl apply, cloud console edits, and controllers acting on their own.

OPA Gatekeeper brings OPA to Kubernetes through a two-level Constraint Framework:

  • A ConstraintTemplate is a CRD defining a parameterised policy written in Rego
  • A Constraint is an instance of that template with specific parameters and scope
apiVersion: templates.gatekeeper.sh/v1
kind: ConstraintTemplate
metadata:
  name: k8srequiredlabels
spec:
  crd:
    spec:
      names:
        kind: K8sRequiredLabels
      validation:
        openAPIV3Schema:
          properties:
            labels:
              type: array
              items: {type: string}
  targets:
  - target: admission.k8s.gatekeeper.sh
    rego: |
      package k8srequiredlabels
       violation contains {"msg": msg} if {
        required := input.parameters.labels[_]
        not input.review.object.metadata.labels[required]
        msg := sprintf("missing required label: %v", [required])
      }

 
apiVersion: constraints.gatekeeper.sh/v1beta1
kind: K8sRequiredLabels
metadata:
  name: require-cost-centre
spec:
  enforcementAction: dryrun
  match:
    kinds:
    - apiGroups: ["apps"]
      kinds: ["Deployment"]
  parameters:
    labels: ["cost-centre"]

The separation is the point. Platform engineers write templates; the same template is instantiated many times with different parameters and different scopes, without anyone touching Rego again.

4. Audit — what is already running

Gatekeeper periodically evaluates existing cluster resources against constraints and records violations in each Constraint's status. This finds what was deployed before the policy existed — which, on any real cluster, is a lot.

Audit reports; it does not prevent. Its value is knowing the size of the remediation job before you switch anything to deny.

Where Gatekeeper still earns its place

Kubernetes has closed much of the gap. ValidatingAdmissionPolicy has been GA since v1.30 — CEL-based rules evaluated inside the API server with no webhook, no network hop and no external component to keep alive. MutatingAdmissionPolicy reached stable in v1.36, so mutation no longer requires a webhook either.

For a simple, high-volume rule, native admission policy is now the better default. It is one less thing to operate and it cannot fail open.

Gatekeeper remains the right answer when you need:

  • Logic beyond CEL's expressiveness — complex evaluation that Rego handles cleanly
  • Existing-resource evaluation via audit
  • External data in policy decisions
  • One engine across Kubernetes, Terraform and CI, which is OPA's genuine differentiator
  • An established OPA investment you are not going to unwind

On Kyverno

Kyverno is the main alternative and the fair comparison is ergonomic. Kyverno policies are Kubernetes-native YAML, which most teams find easier to write and review than Rego. It writes results as PolicyReport CRDs, an open Policy Working Group format, and its newer CEL-based ValidatingPolicy types align closely with native admission policy.

The trade-off: Kyverno is Kubernetes-only. If you want one policy engine covering your clusters, your Terraform and your CI pipelines, OPA is the one that spans all three. If Kubernetes is the whole problem, Kyverno is usually the more pleasant tool.

Either way, note that Kyverno's legacy kyverno.io/v1 ClusterPolicy API was deprecated in 1.17 and is scheduled for removal — write new policies against the current types.

Rolling Out Enforcement: Audit Mode, Warnings, Then Blocking

This section determines whether the initiative survives. Switching policies to deny on a live cluster is the most reliable way to have the whole programme abandoned.

The ladder

Gatekeeper's enforcementAction gives you three rungs. Use all of them, weeks apart.

StageActionWhat happensDuration
1dryrunViolation recorded in Constraint status; request succeeds2–4 weeks
2warnDeveloper sees the message; request still succeeds2–4 weeks
3denyRequest rejectedPermanent

The equivalent in native admission policy is the validationActions field on a PolicyBinding, which takes Audit, Warn and Deny and can combine them.

Stage 1 — dryrun

Deploy the policy with enforcementAction: dryrun and leave it alone.

What you are looking for:

  • How many violations exist. This is the remediation backlog. It is always larger than expected.
  • Which teams are affected. Ten teams with one violation each is a different conversation from one team with two hundred.
  • Whether the rule is even correct. False positives surface here, where they cost nothing.
kubectl get constraint require-cost-centre -o jsonpath='{.status.totalViolations}'

Run it for at least a full sprint. Weekly deployment cycles, cron jobs and quarterly batch processes all need to pass through the policy before you trust the count.

Stage 2 — warn

Switch to warn. Deployments still succeed, but the person applying sees the message immediately.

This stage is doing two things at once. It remediates the backlog through normal work, because developers fix what they can see. And it stress-tests your error messages: if a warning does not tell someone exactly what to change, you find out now rather than after enforcement.

Do not move on until violations are approaching zero and the trend is downward.

Stage 3 — deny

Only when the violation count is zero or every remaining case has a documented exception.

Two things to have in place first:

  • An exception mechanism. Gatekeeper's match block excludes namespaces or label selectors. Exceptions should be time-bounded, recorded in git, and reviewed — not permanent and not verbal.
  • A break-glass procedure. If a policy blocks a production fix during an incident, there must be a documented way to bypass it that leaves a trail. Without one, someone will delete the constraint at 3am and nobody will remember to restore it.

Practical guidance for the rollout

  • Non-production first. Prove the policy in dev and staging before production sees it.
  • One policy at a time. Ten policies deployed together produce ten simultaneous conversations and no clear signal.
  • Start with the ones nobody argues about. Required labels, no :latest image tags, no privileged containers. Save the contentious rules until the mechanism has earned trust.
  • failurePolicy: Ignore at first. If the Gatekeeper webhook is unavailable, do you want deployments to fail? Start with Ignore, move to Fail once it has proven stable — and make sure Gatekeeper's own namespace is exempt, or a restart can deadlock the cluster.
  • Exempt system namespaces from the start. kube-system and your CNI's namespace are not where you want to discover a policy interaction.

Write error messages people can act on

The single highest-leverage thing you can do for adoption.

Poor: policy violation: k8srequiredlabels

Good: Deployment 'checkout-api' is missing required label 'cost-centre'. Add it under metadata.labels. See https://wiki/policies/cost-centre for approved values.

The second version is fixed by the developer in thirty seconds. The first becomes a Slack message to the platform team.

Operate it as a product

Policies are code, so they need the same lifecycle: version-controlled in one repository, tested with opa test in CI, delivered through a pipeline rather than applied by hand, and reviewed quarterly for rules that no longer apply.

Track two numbers. Violation count over time should fall after each rollout. Exception count should stay flat — a rising exception count means the policy is wrong, not that the teams are.

This is the operational core of any serious DevSecOps practice: controls that run automatically, fail loudly and produce their own evidence. It pairs naturally with GitOps delivery, where the policy repository and the manifests it governs move through the same pipeline, and with the controls covered in our Kubernetes production readiness checklist.

If you are introducing policy as code, have a Gatekeeper deployment stuck in audit mode that nobody trusts enough to enforce, or need to decide between OPA, Kyverno and native admission policy for your estate, our DevSecOps services and 24×7 SRE team cover this work — policy design, test suites, rollout sequencing and the exception process that keeps it maintainable.

Talk to our team → We will start by running your candidate policies in dry-run, because the violation count is the only honest starting point.