Standards — applied when an agent is created
When an agent is created, your organization’s standards pack applies automatically. The point is that a new agent cannot quietly ship without an owner, a budget, or grading, because those are not steps somebody has to remember.
Each standard is listed with the value it applies, and the pack is scored out of seven — 7/7 when
everything applies. The standards scorecard sits beside the agent summary, and that is what you
confirm before creating.
Every standard except Docs carries a Waive switch. Waiving one drops the score — 7/7 becomes
6/7 — and the row shows Waived in place of its value rather than disappearing. The governance
value here is legibility, not enforcement: an agent with a waiver is fine, an agent whose waiver
nobody can see is not.
Guardrails — enforced at the action boundary
A guardrail is a rule you configure that decides whether a specific action may proceed. It is evaluated outside the agent: an agent cannot skip a guardrail, because the check does not run in the agent’s own process. It runs at the boundary the agent’s traffic has to cross. You author them under Connect → Integrations, in the Govern → Guardrails section, where New guardrail opens the authoring flow and the list gives each rule its gate, its matcher, and what it does On a match. Policies, beside it under Govern, is a different control and not where guardrails are written. Reading the rules and running a backtest is its own capability, held by the built-in Viewer, Owner, and Admin roles; creating, editing, enabling, disabling or deleting one is an administrator action, held by Owner and Admin. Without the read capability the section does not appear at all, so grant it before pointing someone at this screen.The five gates
A guardrail stands at one gate, and the gate is fixed for the life of the rule.Changing a rule that is already live
A rule opens from its name in the list as Edit guardrail, and every part of it except its gate can change. The Enabled switch stops a rule without removing it, which is usually what you want: deleting asks first and offers to disable it instead, because a delete takes the control away with no undo. A rule declared by a configuration cluster carries a Cluster-managed badge, and while that cluster owns it the switch and the delete are inactive: change the manifest instead.What a guardrail does on a match
Every outcome is recorded, block and record only alike, so a rule that let an action through leaves
the same trail as one that stopped it. Recent decisions, below the rule list, is that trail:
when, which agent, the outcome, the rule that decided, and the tool it was deciding about.
When more than one rule covers the same action, the most severe outcome decides: block, then hold
for approval, then redact, then record only. A rule that only records can never soften one that stops
the action. Redaction is the exception to a single rule deciding — every covering redaction applies,
so two rules removing two different values both take effect.
Redaction reads differently in each direction. On the way in — a tool result, a model answer — the
agent can see that something was removed and reason about it. On the way out the agent composed
the text and is not shown the rewrite, so it may go on to interpret a reply to a question it did not
actually ask. When the agent needs to act on what comes back, block is the honest choice.
Who a rule covers
Target a guardrail at explicit agents, or at a label bag the calling agent’s own labels must contain. Labels are the durable option: a rule scoped toenv = production keeps covering the right
agents as the fleet changes. Set both and the rule covers an agent that matches either one.
A rule with no target at all is the account-wide rule. It covers every agent in the account,
including ones added later, and it covers calls that carry no agent identity — a tool call made from
a person’s own session, or a model request sent with the account key rather than an agent’s.
Naming any target excludes those, so narrowing a rule to a few agents also stops it covering them.
Matching a call
The authoring form is a picker rather than a syntax. It lists the tools your account actually has, grouped by server, so you are choosing from reality instead of guessing at tool names. A glob pattern is available as an alternative — a selection covers the tools that exist now, a pattern covers the ones that arrive later — and the form reports how many of your current tools a pattern matches. You can also condition on a tool’s arguments, chosen from the real argument names of the tools you picked, and on named value classes — a payment card number, a cloud access key — which match a kind of value rather than a string you had to write yourself. A rule can also match on judgment rather than on text. Choose Model judgment as the matcher, state a Policy in plain language — no customer identifiers leave the account — and pick which of your account’s models is Judged by it. That model is shown the intercepted action and answers whether it breaks the policy. A judged policy reads the same at any of the five gates. A rule can carry both shapes at once, and they combine in opposite directions: any criterion matching fires the guardrail, while the conditions within a single criterion must all hold. An either/or is two criteria; a narrowing is two conditions on one.A judged policy cannot drive a redaction. It answers whether a rule was broken, not where the
offending value sits, so there is nothing to excise; the form refuses the pairing. Use a judged
policy to block, hold, or record only.
Test a rule before it is live
A draft rule at the Tool call gate, matching on patterns, can be backtested: Check the last 7 days runs it against the tool calls your agents recorded in that window and reports how many it would have caught. That is what makes it reasonable to write a blocking rule against production: backtesting shows how the rule would have affected the recorded traffic you tested it against, rather than leaving you to discover it live. It reads what was recorded, so treat it as a sample of what happened rather than a complete account of it, and the panel says so when that history cannot be read or when it reached only the most recent calls in the window. The other gates have nothing to replay, and neither does a condition judged by a model. For those, save the rule as Record only, watch what it catches on live traffic, then switch it to block.Which one to reach for
Next steps
Approvals
When the right answer is a human decision rather than a rule.
Roles & permissions
The permission layer a guardrail sits on top of.
Audit log
Where guardrail decisions are recorded.
Data handling & redaction
What is masked without any rule at all.