Prove it is safe to deploy
The agent is scored across six domains before it ships. Clear the bar and it goes; miss it and the pipeline stops rather than a dashboard turning amber somewhere nobody is looking.
AssurancePLATFORM / ASSURANCE
GovernorAI evaluates an immutable snapshot of an agent against preregistered acceptance bars across six assurance domains, before the agent reaches production. All six domains have an implemented evaluator. Whether a domain can actually be measured is a property of how your environment is wired — so GovernorAI computes that per deployment and names the prerequisite that is missing.
Six domains, six implemented evaluators. A domain scores against a bar that was preregistered before the run, so the bar cannot be moved to fit the result.
BEFORE PRODUCTION, THEN IN IT
Two halves, and the order is the point. You build confidence that an agent should ship. Then you keep control of what it does once it has.
The agent is scored across six domains before it ships. Clear the bar and it goes; miss it and the pipeline stops rather than a dashboard turning amber somewhere nobody is looking.
AssuranceOnce it is live, the agent acts. Each action that would change a system of record is decided against your policy at the point it would take effect — allow, deny, or hold for a human — and written down where an assessor can check it.
EnforcementMost assure the model and guard the prompt. These two assess the agent’s authority before release, and decide the action once it is live — across the clouds, SaaS and tools you already run, under one policy and one evidence trail. See where that lands against what you run →
THE DISTINCTION EVERYONE ELSE ROUNDS UP
Most assurance pages answer only the first one, then present the answer as a score. GovernorAI answers both, separately, and shows you which one it is answering.
Permissions, Grounding, Robustness, Security, Privacy and Efficiency each have a real evaluator: its own preregistered scenario battery, deterministic scoring, and its own test suite. None of them is a placeholder waiting to be written.
86 source files · 101 test filesA domain that scores the agent's own outputs needs something able to execute the agent-under-test. The one that probes cross-tenant isolation needs the database whose policies it exercises. Where a prerequisite is absent, the code is fine and the measurement is impossible.
a prerequisite, not a gap in the productA capability manifest answers “could this unit have been measured here?” from how the deployment is wired — deliberately outside the run itself, because a run that quietly degrades would otherwise issue its own excuse alongside its own gap.
GET /api/v1/assurance/ci/capabilitiesAn unmeasurable domain returns not_assessed with an operator-readable reason — no runtime policy decider configured, no agent runner configured — and still records the preregistered bar it would have been held to.
Measurability is derived from how the deployment is wired, never from what a run reported. A domain whose prerequisite is absent reports not_assessed — not a pass, not a rounded-up score, and not a silent omission from the list. Five of the six domains are critical, and a critical unit that reports not_assessed blocks the pre-deployment gate rather than clearing it.
COVERAGE, STATED EXACTLY
Read the last column as the answer GovernorAI actually returns when the prerequisite is not wired. It is a result, not a placeholder for a score we intend to show you later.
| Domain | What the evaluator exercises | Prerequisite in your deployment | Result when the prerequisite is absent |
|---|---|---|---|
| Permissionscritical | Preregistered permission probes replayed through the real runtime policy path — instruction adherence, delegated-authority boundary, unsafe-tool refusal, approval escalation — with the resulting runtime decision evidence linked to the scenario. Identical probes are deduplicated by probe identity, so repetition cannot inflate the denominator. | A runtime policy decider on the evaluation path. | not_assessed “no runtime policy decider configured” |
| Groundingcritical | Source attribution and citation validity, abstention where support is absent, rejection of stale documents, escalation on conflicting sources. Scored deterministically — no model is ever the sole blocking signal. | An agent runner able to execute the agent-under-test and capture its outputs. | not_assessed “grounding cannot be measured without executing the agent-under-test” |
| Robustnesscritical | Repeat consistency, degradation under padded context, duplicate and conflicting input handling, schema conformance, behaviour under bounded concurrent load. | An agent runner able to execute the agent-under-test. | not_assessed “robustness cannot be measured without executing the agent-under-test” |
| Securitycritical | The same deterministic inline detector battery the gateway enforces with — prompt injection, jailbreak, indirect injection, prompt extraction, secret and regulated-identifier exfiltration, unsafe destination — run read-only over the execute seam, with no dispatch and no side effect. It resolves the account's effective outcome overrides fresh on every probe, so it scores the posture you actually run, not the default one. | A shared store the control plane can read the account's effective outcome overrides from. | not_assessed “no inline security inspector configured” |
| Privacycritical · composite | Isolation: a bounded, serialized, unconditionally rolled-back transaction per crown-jewel table that seeds a row as one synthetic tenant and asserts a second cannot read, update, delete, or write into it — exercising the real database isolation policies, not a re-implementation. PII: whether the agent's own outputs disclose sensitive data. | Isolation needs a database carrying the real isolation policies. PII needs an agent runner. | composite not_assessed The domain cannot pass on half its evidence. The isolation component's own counts and records are still preserved and surfaced, and the reason names the status of each half. |
| Efficiencynon-critical | Latency, provider-metered token and cost consumption, task-success-adjusted cost, bounded consumption. Only usage carrying explicit provenance is counted. | An agent runner supplying metered usage with provenance. | not_assessed Projected or unprovenanced usage is never promoted to a measurement. Any unassessable probe caps the domain at |
Two domains are scored with different instruments, and the report says which. A confidence bound answers “given a sample, what can I claim about the population?” — the right question for adversarial probes standing in for an unbounded space, and the wrong one for an enumerated set of tables where the population is the sample. Privacy's isolation component is therefore scored by coverage, passing only at N of N, and no confidence interval or acceptance bar is reported for it. A reader is expected to key off the instrument, not off a number that would be meaningless for it.
THE PRE-DEPLOYMENT GATE
A result that outlives the thing it described is worse than no result. Assurance pins the identity it measured, and invalidates itself when that identity — or the policy path around it — moves.
A snapshot fixes what was evaluated: model provider, name and version; a hash over the prompt manifest; a hash over the tool manifest; the artifact digest and the runtime-config digest. The row is immutable — whether it is the current snapshot is derived at read time, never written back onto it.
config_hashEvery measured domain reports its raw numerator and denominator, sample size, a confidence interval, and a bar that was preregistered rather than chosen after seeing the result. A domain reporting not_assessed still carries its bar, so what it would have been held to stays visible.
A change to the evaluated configuration mints a new snapshot. A change that alters the policy path without touching the configuration — a namespace move, an assigned-policy edit, a rule change — stales every prior run for that agent. So does a superseded evaluator version. The gate treats a stale run as a block, not a warning.
--on-stale=failPass, blocked, malformed input, control plane unreachable, and authentication failure are five distinct exit codes. Unreachable never exits zero — a control plane that read no evidence cannot produce a green gate — and auth failure stays separate, so a workflow tolerating a network blip can never silently tolerate a deleted key.
cmd/governor-assurance-gateinit, register, assess, gate, deploy-begin, attest, promote — invoked as separate steps so your own deployment sits between them. Each finalising step demands the deployment step's own outcome, because a skipped deploy that still attested would bind evidence to bytes nobody shipped.
A run exports as a self-describing evidence bundle whose staleness statement is made as of the moment it was generated. A separate operator client retains captured bundles in the repository and re-verifies them, and reports the retention condition honestly — including when it is not met.
cmd/governor-assurance-evidence# the identity being evaluated, at the commit being deployed
version: 1
agent_id: refund-agent
namespace: governor.prod
model:
provider: anthropic
name: claude-opus-4-5
version: "20260514"
prompts:
- prompts/system.md
- prompts/refund-policy.md
tools:
- path: tools/refund.schema.json
artifact:
image_repository: ghcr.io/acme/refund-agent
runtime_config:
values:
TEMPERATURE: "0.2"
files: []
secrets:
- env: PROVIDER_API_KEY
provider: aws-secretsmanager
ref: prod/agent/provider-key $ governor-assurance-gate gate --format text
GovernorAI assurance gate: blocked
Measured 3/6 critical units
Blocking (3):
- grounding [not_assessed]
no agent runner configured — grounding cannot
be measured without executing the
agent-under-test
- robustness [not_assessed]
no agent runner configured — robustness cannot
be measured without executing the
agent-under-test
- privacy/pii [not_assessed]
no agent runner configured — PII leakage cannot
be measured without executing the
agent-under-test
# exit 1 — a real verdict about the candidate Any tool can return a score. What decides whether a reviewer should believe this one is that the subject is pinned rather than a moving branch — the thing that passed and the thing that ships are provably the same thing — that the bar was set before the run, that an unmeasurable domain says so instead of rounding up, and that the verdict expires when the ground moves. The result leaves as a re-verifiable artifact rather than a dashboard state.
The check conclusion is deliberately not the same thing as the step outcome. Only a clean pass reports success; an outcome that was degraded, excused, or reached over unmeasured units reports neutral or failure. There is no configuration in which a degraded result can be reported as a green check.
SCOPE BOUNDARY
Precision is the point. A number that survives contact with an auditor is worth more than six of them that do not.
Assurance evaluates an agent you already built, running on a framework you already chose. It does not host the agent, replace its harness, or require a rewrite to be evaluated — the snapshot is derived from a manifest checked in beside your code.
no agent-builder replacementWhere a domain is measurable, it exercises the production mechanism rather than a model of it: Permissions replays through the real runtime policy path, Security over the real detector battery, Privacy's isolation half against the real database policies.
not a simulation of the controlA measured pass, a measured failure, and an absence of evidence are three different statements, and a malformed probe battery is a fourth — a configuration error, never a measured failure of the agent. Collapsing any of them into the others is how a green board stops meaning anything.
pass · fail · not_assessed · config error