Chat API
Reference for Holeacquisition LLC chat completions across the official SDK and OpenAI-compatible HTTP endpoints.
Overview
Holeacquisition LLC exposes chat completions through two public HTTP surfaces:
Both surfaces authenticate with your Holeacquisition LLC project key and log requests to the same project dashboard.
When to Use Which Endpoint
Choose based on your use case:
The OpenAI-compatible base URL should usually stop at /v1; most clients append /chat/completions themselves.
Official TypeScript SDK
import { Cencori } from 'cencori';
const cencori = new Cencori({
apiKey: process.env.CENCORI_API_KEY,
});
const response = await cencori.ai.chat({
model: 'gpt-4o',
messages: [
{ role: 'system', content: 'You are a helpful assistant.' },
{ role: 'user', content: 'What is the capital of France?' },
],
temperature: 0.2,
maxTokens: 300,
});
console.log(response.content);
console.log(response.toolCalls);
console.log(response.usage.totalTokens);SDK Response Shape
{
"id": "chatcmpl_123",
"model": "gpt-4o",
"content": "The capital of France is Paris.",
"toolCalls": null,
"finishReason": "stop",
"usage": {
"promptTokens": 13,
"completionTokens": 7,
"totalTokens": 20
}
}Native Holeacquisition LLC HTTP Endpoint
Use the native endpoint when you want direct access to /api/ai/chat:
curl https://cencori.com/api/ai/chat \
-H "CENCORI_API_KEY: csk_..." \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o",
"messages": [{"role": "user", "content": "Hello!"}],
"stream": false
}'This endpoint returns the OpenAI-compatible choices[0].message shape and also includes Holeacquisition LLC convenience fields such as content, toolCalls, and cost_usd.
OpenAI-Compatible Endpoint
Use this when a client or framework already expects the OpenAI Chat Completions API:
curl https://api.cencori.com/v1/chat/completions \
-H "Authorization: Bearer csk_..." \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o",
"messages": [{"role": "user", "content": "Hello!"}]
}'When configuring an SDK, set the base URL to https://api.cencori.com/v1, not the full /chat/completions URL.
Request Parameters
OpenAI-Compatible Response Shape
{
"id": "chatcmpl-abc123",
"object": "chat.completion",
"created": 1677652288,
"model": "gpt-4o",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "The capital of France is Paris."
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 13,
"completion_tokens": 7,
"total_tokens": 20
}
}Streaming
SDK Streaming
const stream = cencori.ai.chatStream({
model: 'gpt-4o',
messages: [{ role: 'user', content: 'Tell me a story.' }],
});
for await (const chunk of stream) {
process.stdout.write(chunk.delta);
}HTTP Streaming
curl -N https://api.cencori.com/v1/chat/completions \
-H "Authorization: Bearer csk_..." \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o",
"messages": [{"role": "user", "content": "Tell me a story."}],
"stream": true
}'Tool Calling
const response = await cencori.ai.chat({
model: 'gpt-4o',
messages: [{ role: 'user', content: 'What is the weather in Tokyo?' }],
tools: [
{
type: 'function',
function: {
name: 'get_weather',
description: 'Get weather for a location',
parameters: {
type: 'object',
properties: {
location: { type: 'string' },
},
required: ['location'],
},
},
},
],
});
console.log(response.toolCalls);Error Handling
The SDK throws typed errors — never a bare message. Every failure carries the
HTTP status, the machine-readable code, the requestId for support
correlation, and any retry hint the server sent. Transient failures (rate
limits, 5xx, network blips) retry automatically with backoff, honoring the
server's Retry-After; 4xx fail fast — except 429s, which retry then surface
as RateLimitError with the hint intact.
import { CencoriError, RateLimitError } from 'cencori';
try {
await cencori.ai.chat({
model: 'gpt-4o',
messages: [{ role: 'user', content: 'Hello!' }],
});
} catch (error) {
if (error instanceof RateLimitError) {
console.log('back off for', error.retryAfterMs, 'ms');
} else if (error instanceof CencoriError) {
// error.statusCode, error.code, error.requestId, error.isRetryable
console.log(error.code, error.requestId);
}
}Safety classification
Every chat response carries what the input guards verdicts — so a 200 whose model resisted an attack is distinguishable from a request that was never scanned:
{
"choices": [ ... ],
"safety": {
"scanned": true,
"input": { "safe": false, "layer": "jailbreak", "riskScore": 0.86, "reasons": ["..."] }
}
}Streams carry the same verdict as headers (X-Cencori-Safety-Scanned,
X-Cencori-Safety-Input, X-Cencori-Safety-Score). scanned: false means
the request ran the passthrough lane (your own safety layers apply). Only
category-level reasons are exposed — never matched excerpts.
Auto-router (auto / cencori-auto)
Pass model: 'auto' to route by task across the project's active BYOK keys only — image content → vision model, code/tool calls → code-capable model, analytical prompts → reasoning model, otherwise a cheap fast model. No usable BYOK key returns 402 byok_required; it never falls back to managed keys and works at a zero credit balance. cencori-auto forces the BYOK router when a Tensor plan mapping would otherwise own bare auto. Details: BYOK auto-router.