Content Filtering
Opt-in harmful-content checks. Enable per project; output enforcement needs a policy.
Content filtering is explicit opt-in and off by default. Enable in Project > Security before expecting blocks.
Categories
When enabled, input scoring covers harmful/injection keywords unified across providers (not provider-native safety settings):
There are no fixed Low (0.8) / Medium (0.5) / High (0.2) gateway cutoffs. You get one Safety Threshold slider in Project > Security (higher = stricter). Input score is compared against that value; output uses threshold − 0.1 internally when a governance policy enables output evaluation.
Output note
The gateway does not auto-block toxic output via the master switch. For output categories, either:
- Call the standalone Moderation API explicitly, or
- Install a Governance policy with output rules (block / redact / require_approval).
This avoids false-positive stream terminations from broad substring heuristics (e.g. normal coding answers mentioning "vulnerability").
Webhooks & Alerts
Receive notifications when a guard actually fires (no events fire while scanning is off):
- Go to Dashboard > Settings > Webhooks.
- Add an endpoint URL (e.g., your Slack webhook or pagerduty).
- Subscribe to
security.content_flagged.
Payload Example:
{
"event": "security.content_flagged",
"project_id": "proj_123",
"timestamp": "2024-03-20T10:00:00Z",
"data": {
"prompt": "...",
"category": "violence",
"score": 0.95,
"user_id": "user_456"
}
}Moderation Endpoint
You can also use the standalone Moderation API to check text without generating a response:
const result = await cencori.ai.moderation({
input: "Text to check"
});