Skip to content

Enforcement actions

Detection settings decide what counts as sensitive. Enforcement actions decide what happens when something is caught. Both live on the Guardrails page at app.sanitized.ai.

The three actions

Every guardrail has one:

  • Nudge — the warning banner appears with Sanitize and Edit options, plus a Send Anyway button. The person can proceed, and the override is recorded.
  • Block — the same banner without Send Anyway. The prompt cannot be sent until the sensitive content is removed. Sanitize and Edit remain, so nobody is ever stuck.
  • Monitor — no banner at all. The prompt sends and the detection is logged silently. Useful when you want visibility on something before you decide to act on it.

Defaults by severity

You don’t have to set an action on all your guardrails one at a time. Each guardrail carries a severity, and severity picks the default:

SeverityDefault action
CriticalBlock
HighBlock
Medium-HighNudge
MediumNudge
LowNudge

The Enforcement Policy panel at the top of the Guardrails page shows this table alongside how many active guardrails sit in each severity — your whole policy in five rows.

Overriding a single guardrail

The Enforcement column in the guardrail table lets you override any one guardrail. Leave it on Default to follow the severity rule, or pick Nudge, Block, or Monitor explicitly.

Anything you override appears in the Overrides card in the Enforcement Policy panel, with a one-click Reset to put it back on the default. That way exceptions to your own policy stay visible instead of getting lost among dozens of guardrails.

Before you enable a block

When you switch a guardrail to Block, a confirmation dialog shows what that block would have caught recently — for example, “In the last 30 days, this would have blocked 142 prompts from 37 users.”

Treat that number as a warning sign, not just a statistic. A high count usually means one of two things: the data really is being pasted into AI tools a lot (in which case the block is doing its job), or the guardrail is over-firing (in which case fix it first). If you’re unsure, use a soft launch instead of blocking immediately.

Explaining a guardrail to your team

Each guardrail has a Warning Banner Note field, editable from the pencil icon. Whatever you write appears directly in the warning your team sees.

Use it to answer “why is this blocked?” before anyone has to ask you:

Patient identifiers can never enter AI tools, per our privacy obligations.

A note is far more effective than a generic warning, and it’s the cheapest way to cut down on questions.

Stricter rules on riskier tools

Not every AI tool carries the same risk. Risk-aware enforcement, in the Enforcement Policy panel, lets you set one threshold: when a prompt is headed for a tool at or above that risk band, nudges become blocks.

The result is that the same data can nudge on a sanctioned enterprise tool and hard-block on a risky free one, without you maintaining two sets of guardrails.

Two things to know:

  • Only nudges escalate. A guardrail you set to Monitor stays silent — that’s a deliberate “don’t interrupt this person” decision, and escalation won’t override it. A guardrail already set to Block is unaffected.
  • Risk bands come from each tool’s report card, which you can see on the Guarded Sites page.

Seeing whether it’s working

The Events page has an Overrides card showing the share of flags people sent anyway, plus the guardrails overridden most often.

That list is your tuning queue. A guardrail near the top is one your team routinely works around, which means either it’s firing on things that aren’t sensitive — fix the guardrail — or the data genuinely needs to be sent and the guardrail should become an exception rather than a daily obstacle.