Developers / Benchmark

Policy-engine latency, with the method that produced it.

Policy evaluation in the sidecar deployment model is CPU-bound: the decision never leaves the host. What this page measures is the policy engine itself, in process. End-to-end sidecar percentiles are a separate measurement and are still outstanding; the title says which of the two you are reading, because a latency figure without its method is not evidence.

In-process engine, sidecar topology Method stated before the number Not a platform-wide SLA

SCOPE

What this figure is, and is not.

Stated first, deliberately, because the same number is misleading one topology over.

It is

Policy evaluation latency measured in the sidecar deployment model, where the engine runs alongside the agent on the same host and the decision path involves no network round-trip.

It is not

  • A service-level guarantee. It is a measurement under stated conditions, not a commitment.
  • A figure that transfers to other seams. A gateway adapter, a transparent MCP proxy or an HTTP forward-auth hop each add their own cost, on their own path.
  • A comparison against other products. Nothing here is measured against another vendor, so nothing here claims to beat one.

RESULTS · POLICY EVALUATION

The decision itself costs about 0.13 of a microsecond.

This is the first figure published on this page, and it is deliberately the narrower of the two this page owes you: the cost of evaluating a policy, measured in process. The conditions are in the second table, beside the number rather than after it, because a latency figure whose conditions arrive afterwards is a marketing claim wearing a lab coat. The end-to-end sidecar percentiles from the load harness are still outstanding, and the distinction matters enough to keep them apart.

Policy evaluation
Mean, nanoseconds per evaluation
Allow, tool in allow-list
62 – 68 ns
Deny, tool not permitted
83 – 89 ns
Full engine evaluation
124 – 139 ns
With rules and an argument condition
147 – 150 ns
Under concurrency
147 – 155 ns
Large rule set
1,530 – 1,627 ns

A microsecond is a thousand nanoseconds. A typical evaluation is roughly an eight-thousandth of a millisecond, and a large rule set is still under two microseconds. The useful statement is not that this is fast — it is that the decision is cheap enough that anything you can measure in a deployment is the hop around it, not the decision. That is why this page will not print one number and call it latency.

Condition
As run
What was measured
Policy evaluation only, in process. No network, no adapter hop, no serialisation, no I/O, no audit write.
Statistic
Arithmetic mean nanoseconds per operation, as Go's benchmark tooling reports it. Not p50, p95 or p99 — a microbenchmark does not produce a distribution, and this page said the percentile belongs in the headline. It is absent here because it has not been measured, not because it is inconvenient.
Hardware
Apple M4, 10 cores. A developer machine, not the production target. The production stack runs on Amazon EKS, and a cloud vCPU will not match an M4 core. Expect a larger figure there.
Deployment topology
Sidecar. The policy engine runs on the same host as the agent, so no policy decision crosses the network. This page describes no other topology — a gateway adapter, a transparent MCP proxy or a forward-auth hop each add their own cost on their own path.
Policy set
The benchmark fixtures in internal/policy: a small allow-list for the simple cases, a rule set with an argument condition for the rule cases, and a deliberately large rule set for the last row. Engine embedded, not a remote OPA.
Runtime
go1.26.4, darwin/arm64.
Method
go test -bench -benchtime=3s -count=3 against internal/policy. Three runs per case; the spread between them is the range printed above.
Sample size
Between 2.2 million and 165 million evaluations per case, chosen by the benchmark tool to fill three seconds.
Date
26 September 2026.

What is still owed. Two things, and this page will not pretend either is done. A run on the production instance type rather than a laptop — the number above is honest about its hardware but it is not the hardware you would deploy onto. And p50, p95, p99 and max from the load harness, covering the sidecar path end to end rather than the evaluation alone. Those are the figures an SRE will ask for, and they are not these.

WHAT WAS WITHDRAWN

The figures that used to be here.

Continue