Dashboard|

Content Filtering

Opt-in harmful-content checks. Enable per project; output enforcement needs a policy.

Content filtering is explicit opt-in and off by default. Enable in Project > Security before expecting blocks.

Categories

When enabled, input scoring covers harmful/injection keywords unified across providers (not provider-native safety settings):

CategoryWhat fires
Hate / Violence / Self-Harm / SexualProvider-mapped moderation via the standalone Moderation API when called explicitly; gateway input guard scores injection/PII keywords and jailbreak signals toward the Safety Threshold
Jailbreak / InjectionSee Prompt Injection
PIISee PII Detection

There are no fixed Low (0.8) / Medium (0.5) / High (0.2) gateway cutoffs. You get one Safety Threshold slider in Project > Security (higher = stricter). Input score is compared against that value; output uses threshold − 0.1 internally when a governance policy enables output evaluation.

Output note

The gateway does not auto-block toxic output via the master switch. For output categories, either:

  • Call the standalone Moderation API explicitly, or
  • Install a Governance policy with output rules (block / redact / require_approval).

This avoids false-positive stream terminations from broad substring heuristics (e.g. normal coding answers mentioning "vulnerability").

Webhooks & Alerts

Receive notifications when a guard actually fires (no events fire while scanning is off):

  1. Go to Dashboard > Settings > Webhooks.
  2. Add an endpoint URL (e.g., your Slack webhook or pagerduty).
  3. Subscribe to security.content_flagged.

Payload Example:

{
  "event": "security.content_flagged",
  "project_id": "proj_123",
  "timestamp": "2024-03-20T10:00:00Z",
  "data": {
    "prompt": "...",
    "category": "violence",
    "score": 0.95,
    "user_id": "user_456"
  }
}

Moderation Endpoint

You can also use the standalone Moderation API to check text without generating a response:

const result = await cencori.ai.moderation({
  input: "Text to check"
});