USE CASES / MODEL CONTEXT PROTOCOL

Govern tools/call. Leave the rest of the protocol alone.

GovernorAI ships a transparent reverse proxy that speaks MCP Streamable-HTTP and sits in front of a real MCP server. One method is governed — the one that makes something happen — and every other message is forwarded untouched. Point your client's server URL at the proxy and nothing in the agent changes.

Transparent JSON-RPC proxy Argument-aware inspection Fail closed by contract
FOUR VERDICTS, ON THE WIRE allow forwarded upstream byte for byte allow + shaped params.arguments replaced pause JSON-RPC error −32001 deny JSON-RPC error −32002 initialize · tools/list · ping — forwarded untouched. One method is governed.

One method is governed — the one that makes something happen. initialize, tools/list and ping are forwarded untouched, so the proxy is transparent for everything that is not a tool call.

tools/call−32001 pause−32002 denyeverything else forwarded

TRANSPARENT PROXY · governor-mcp-proxy

One call in. Four things that can come back out.

The proxy never re-implements policy, data controls, approvals or the kill switch. It translates the MCP tool call into the canonical gateway request, asks for a verdict, and enforces the answer on the wire. There is one decision brain and this is a protocol translator in front of it.

The client sends
{ "jsonrpc": "2.0", "id": 7, "method": "tools/call",
  "params": { "name": "refund_customer",
              "arguments": { "account": "…", "amount": 8200 } } }
allow Forwarded verbatim

The original request reaches the upstream MCP server unchanged. The proxy adds nothing to the payload.

→ upstream, byte for byte
allow + shaped Arguments rewritten

params.arguments is replaced with the shaped arguments, then the rewritten request is forwarded.

→ upstream, with redacted args
pause Held for a human

A JSON-RPC error carrying the approval id and URL. The call is never forwarded while it waits.

← error −32001
deny Held, and recorded

Policy deny, kill switch, or a fail-closed denial. Never forwarded; the decision still produces evidence.

← error −32002

Everything else passes through. initialize, tools/list, ping, notifications, resources and prompts are forwarded untouched, and the Streamable-HTTP session headers are preserved in both directions so the transport handshake keeps working. Fail-closed is the contract: an unparseable body, a tools/call with no tool name, an oversized payload or any decision error resolves to a denial — the upstream server is reached only for calls the gateway explicitly allowed.

No gateway at all Straight in front of the server

The simplest insertion: the proxy is the URL the MCP client is configured with, and the real server sits behind it.

Chained Inside a gateway you already run

The proxy is chainable inside Docker MCP Gateway or ContextForge, so an existing MCP gateway keeps its role and gains a decision point.

Elsewhere in the stack Other insertion adapters

The same decision is available as an Envoy external processor and as an HTTP forward-auth check for teams whose traffic already goes through one.

SEAM · mcp_invocation

An MCP call is a first-class decision target, not a prompt with extra steps.

Most AI controls are built to read text. An MCP tool call is often the least textual thing in the trace and the most consequential — a tool name and a structured argument object. GovernorAI registers mcp_invocation as its own enforcement point, and argument-aware inspection runs on it even where the same call would be a no-op for prompt or response classification.

Declared capability of the mcp_invocation enforcement point
DeclaredWhat the enforcement point carries
Interaction kinds MCP calls and tool calls. Plus retrieved content: when an operator declares — or a caller hints — that an MCP call is a retrieval fetch, the returned body is treated as untrusted retrieved content and inspected for indirect prompt injection.
Outcomes allow · deny · pause · constrain · redact · mask — one of the richest outcome sets in the registry.
Not carried Ordinary model-answer responses are not inspected on this enforcement point. Response inspection lives on the central gateway path and on the Bedrock adapter.
Unknown enforcement point behaviour A enforcement point not in the registry falls back to allow, deny and approval only. Asking for a richer outcome than an enforcement point declares fails closed to a deny.
Why the argument matters The question is not whether this agent may call this tool. It is whether it may call it with these arguments.

refund_customer is not a risk. refund_customer with an amount above a threshold, against an account outside a scope, is. Policy is evaluated against the action and its arguments, an argument can be narrowed before dispatch, and the constrained call is re-verified against the exact payload rather than trusted to have been applied.

MCP CATALOG

Know which servers exist before you argue about which are allowed.

The catalog builds a register of publicly published MCP servers from three sources, scores each entry, and gives it a lifecycle state a security team can act on. It is the artifact behind the sentence "these are the servers we permit" — and its limits are as important as its contents.

Tier 1 Official registry sync

Sync the official MCP servers repository and record what it publishes, with the repository metadata that goes into scoring.

Tier 2 GitHub discovery

Search GitHub for MCP server repositories, optionally restricted to an allowlist of organisations you actually trust.

Tier 3 Curated manifest repo

Read manifests from a repository you control, in a catalog-only, catalog-plus-approval, or full sync mode.

discovered→ validated→ assessed→ approved· blocked· quarantined
The lifecycle An entry moves from discovered, through validated and assessed, to approved — or to blocked or quarantined.

The states are the point of the catalog. A discovered server is not a permitted one, and the record of who moved it and when is what makes an allowlist reviewable rather than folkloric. Entries carry their source, so "approved" always answers the follow-up question: approved on the strength of what?

ASSESSMENT · FOUR DIMENSIONS, THREE OF THEM REAL

The composite score, and the part of it that is a constant.

Each entry is assessed on four weighted dimensions and reduced to a composite risk score. Three of them are computed from evidence. The fourth is not implemented, and pretending otherwise would be the exact failure mode this product exists to argue against — so it is on the page, at the same size as the rest.

0.25 Provenance

Where the server came from: official registry listing, repository age, popularity signals. A judgement about the source, not the code.

0.20 Supply chain

Distribution and packaging: whether a digest is pinned, how it is distributed, community signals around the repository.

0.35 Declared capability

The heaviest weight. Dangerous verbs in the name and description, and manifest permissions — PII access, data egress, secrets access, tool count, sandbox level, network allowlist.

0.20 Behaviour

Not assessed. The behavioural stage returns a fixed placeholder score for every entry, labelled "behavioural sandbox assessment is not available yet". It contributes an identical amount to every composite and distinguishes nothing.

Placeholder · not measured
So what is the catalog for The register earns its place at the moment a call is made, not at the moment a score is computed.

A catalog entry gives a policy something to point at: this MCP id, this approval state, this declared capability. The decision that actually protects a system of record is the one at the action boundary — this agent, this tool, these arguments, now — and that decision is made against the live call whatever the catalog thought of the server in the abstract. Use the score to triage a review queue. Do not use it as a substitute for a control.

BEFORE THE TOOL IS EXPOSED

Governing tools/call assumes the tool surface was worth exposing.

Runtime mediation decides each invocation. Assurance asks the earlier question — whether this agent, with this tool surface and these permissions, should reach production — and answers it against a bar fixed before the run.

Six domains carry six implemented evaluators, each with its own preregistered scenario battery and deterministic scoring.

See what each domain needs →

Continue