Guardrails & moderation
Relay includes lightweight, deterministic (BCL-only) safety features you can turn on without any external service.
#PII redaction
Guardrails.RedactPii masks obvious personal data — emails, phone numbers, and long card-like digit runs — before text is sent upstream. Enable it per agent (the Redact PII option). When on, it applies to both the user's messages and any retrieved RAG context.
Use it for agents that touch customer data but shouldn't leak raw identifiers to the model provider.
#Content moderation
ContentModerator is a rule-based moderator plus prompt-injection / jailbreak detection, using deterministic keyword and pattern heuristics. Control it with config:
Gateway:Moderation:Enabled(defaultfalse) — turn moderation on.Gateway:Moderation:Block(defaulttrue) — block flagged content vs. just flag it.Gateway:Moderation:FlagJailbreak(defaulttrue) — detect prompt-injection / jailbreak attempts.
When enabled and blocking, a flagged request is rejected before it reaches a provider; otherwise the flag is recorded and the call proceeds.
#What these are (and aren't)
These are fast, local, best-effort guards — no model calls, no packages, no latency tax. They catch common cases (obvious PII, blatant injection, banned keywords). They are not a replacement for a dedicated moderation model where regulatory-grade filtering is required. Combine them with provider-side safety settings for defence in depth.