Dashboard|

Prompt Injection

Opt-in jailbreak and injection blocking. Enable per project; tune threshold and filters.

Injection blocking is explicit opt-in and off by default. Direct injections (Ignore previous instructions…) and role-spoofing that return 200 OK mean the project has no active guard — enable scanning first. Same behavior for BYOK and managed keys.

How enabling works

  1. Go to Project > Security.
  2. Turn on Enable security scanning.
  3. Keep filter_jailbreaks and filter_prompt_injection on.
  4. Tune the Safety Threshold slider (higher = stricter). There are no separate Standard / Strict / Zero-Trust tiers and no per-request Llama Guard routing in the gateway path — those names do not map to a gateway setting.

Custom data rules and Governance policies (e.g. Prompt-Injection & Jailbreak Defense template) are separate opt-ins. Creating one enforces it even if the master switch is off.

What is actually checked

When enabled, checkInputSecurity runs on the last user message:

  1. Keyword heuristics: ignore previous instructions, disregard…, you are now, pretend you are, DAN, developer mode, override security, etc.
  2. Jailbreak heuristics: social-engineering framing (writing a story, hypothetically, roleplay), system-extraction (reveal your system, show me your instructions), behavioral probes, indirect-PII prompts, multi-question / topic-switch signals. Agent tool contexts (<tool_call>, function_call, etc.) are whitelisted to avoid flagging legitimate framework traffic.
  3. Intent patterns: indirect requests to share contact info subtly.

There is no 50k-vector jailbreak database or mandatory LLM classifier in the request path. Scores combine into a risk value compared against your threshold (isJailbreakRisky defaults to risk ≥ threshold and confidence ≥ 0.3).

What you get on block

Input block returns 403:

{
  "error": "Security violation detected",
  "message": "Security violation detected",
  "code": "security_violation",
  "reasons": ["[Input] Potential prompt injection keyword detected: \"ignore previous instructions\"", "[Jailbreak] system_extraction: \"reveal your system\""]
}

There is no security_injection_detected code with metadata.confidence. Parse the flat error / code / reasons shape per Error Reference. The incident is logged to security_incidents and optionally fanned out via webhook.

Common attacks

When enabled, these score toward a block (exact outcome depends on threshold + surrounding text):

  • DAN / Mongo Tom: "You are going to pretend to be DAN…"
  • Direct override: "Ignore previous instructions…"
  • Virtualization: "Imagine you are a Linux terminal…"
  • Role spoofing via history: extra system or assistant turns claiming new privileges — evaluated as part of conversation history signals.

With scanning off, all of the above return 200 OK. That is expected default behavior.

Output and bypass notes

  • Output injection / exfiltration teaching is not auto-blocked by the master switch. Add a governance output policy if you need it.
  • passthrough: true / fast_lane: true (body) or x-cencori-passthrough: true / x-cencori-fast-lane: true (headers) skips input and output guards entirely, even when enabled. Use only for callers running their own safety layers.