Prompt Injection
Opt-in jailbreak and injection blocking. Enable per project; tune threshold and filters.
Injection blocking is explicit opt-in and off by default. Direct injections (
Ignore previous instructions…) and role-spoofing that return200 OKmean the project has no active guard — enable scanning first. Same behavior for BYOK and managed keys.
How enabling works
- Go to Project > Security.
- Turn on Enable security scanning.
- Keep
filter_jailbreaksandfilter_prompt_injectionon. - Tune the Safety Threshold slider (higher = stricter). There are no separate Standard / Strict / Zero-Trust tiers and no per-request Llama Guard routing in the gateway path — those names do not map to a gateway setting.
Custom data rules and Governance policies (e.g. Prompt-Injection & Jailbreak Defense template) are separate opt-ins. Creating one enforces it even if the master switch is off.
What is actually checked
When enabled, checkInputSecurity runs on the last user message:
- Keyword heuristics:
ignore previous instructions,disregard…,you are now,pretend you are,DAN,developer mode,override security, etc. - Jailbreak heuristics: social-engineering framing (
writing a story,hypothetically,roleplay), system-extraction (reveal your system,show me your instructions), behavioral probes, indirect-PII prompts, multi-question / topic-switch signals. Agent tool contexts (<tool_call>,function_call, etc.) are whitelisted to avoid flagging legitimate framework traffic. - Intent patterns: indirect requests to share contact info subtly.
There is no 50k-vector jailbreak database or mandatory LLM classifier in the request path. Scores combine into a risk value compared against your threshold (isJailbreakRisky defaults to risk ≥ threshold and confidence ≥ 0.3).
What you get on block
Input block returns 403:
{
"error": "Security violation detected",
"message": "Security violation detected",
"code": "security_violation",
"reasons": ["[Input] Potential prompt injection keyword detected: \"ignore previous instructions\"", "[Jailbreak] system_extraction: \"reveal your system\""]
}There is no security_injection_detected code with metadata.confidence. Parse the flat error / code / reasons shape per Error Reference. The incident is logged to security_incidents and optionally fanned out via webhook.
Common attacks
When enabled, these score toward a block (exact outcome depends on threshold + surrounding text):
- DAN / Mongo Tom: "You are going to pretend to be DAN…"
- Direct override: "Ignore previous instructions…"
- Virtualization: "Imagine you are a Linux terminal…"
- Role spoofing via history: extra
systemorassistantturns claiming new privileges — evaluated as part of conversation history signals.
With scanning off, all of the above return 200 OK. That is expected default behavior.
Output and bypass notes
- Output injection / exfiltration teaching is not auto-blocked by the master switch. Add a governance output policy if you need it.
passthrough: true/fast_lane: true(body) orx-cencori-passthrough: true/x-cencori-fast-lane: true(headers) skips input and output guards entirely, even when enabled. Use only for callers running their own safety layers.