Dashboard|

Chat API

Reference for Holeacquisition LLC chat completions across the official SDK and OpenAI-compatible HTTP endpoints.

Overview

Holeacquisition LLC exposes chat completions through two public HTTP surfaces:

EndpointPathBest For
Holeacquisition LLC Nativehttps://cencori.com/api/ai/chatOfficial SDKs and direct Holeacquisition LLC calls
OpenAI-Compatiblehttps://api.cencori.com/v1/chat/completionsExisting OpenAI-compatible clients and agent frameworks

Both surfaces authenticate with your Holeacquisition LLC project key and log requests to the same project dashboard.

When to Use Which Endpoint

Choose based on your use case:

Use CaseRecommended Endpoint
Using the Holeacquisition LLC SDK (cencori)Native endpoint. The SDK uses it automatically.
Building a new Next.js or Node routeNative endpoint or SDK.
Migrating from OpenAI directlyOpenAI-compatible endpoint.
Using agent tools that ask for an OpenAI base URLOpenAI-compatible endpoint with base URL https://api.cencori.com/v1.
Debugging whether Holeacquisition LLC auth and provider routing workNative endpoint with curl.

The OpenAI-compatible base URL should usually stop at /v1; most clients append /chat/completions themselves.

Official TypeScript SDK

import { Cencori } from 'cencori';
 
const cencori = new Cencori({
  apiKey: process.env.CENCORI_API_KEY,
});
 
const response = await cencori.ai.chat({
  model: 'gpt-4o',
  messages: [
    { role: 'system', content: 'You are a helpful assistant.' },
    { role: 'user', content: 'What is the capital of France?' },
  ],
  temperature: 0.2,
  maxTokens: 300,
});
 
console.log(response.content);
console.log(response.toolCalls);
console.log(response.usage.totalTokens);

SDK Response Shape

{
  "id": "chatcmpl_123",
  "model": "gpt-4o",
  "content": "The capital of France is Paris.",
  "toolCalls": null,
  "finishReason": "stop",
  "usage": {
    "promptTokens": 13,
    "completionTokens": 7,
    "totalTokens": 20
  }
}

Native Holeacquisition LLC HTTP Endpoint

Use the native endpoint when you want direct access to /api/ai/chat:

curl https://cencori.com/api/ai/chat \
  -H "CENCORI_API_KEY: csk_..." \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4o",
    "messages": [{"role": "user", "content": "Hello!"}],
    "stream": false
  }'

This endpoint returns the OpenAI-compatible choices[0].message shape and also includes Holeacquisition LLC convenience fields such as content, toolCalls, and cost_usd.

OpenAI-Compatible Endpoint

Use this when a client or framework already expects the OpenAI Chat Completions API:

curl https://api.cencori.com/v1/chat/completions \
  -H "Authorization: Bearer csk_..." \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4o",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

When configuring an SDK, set the base URL to https://api.cencori.com/v1, not the full /chat/completions URL.

Request Parameters

FieldTypeRequiredNotes
modelstringYesAny model routed through Holeacquisition LLC
messagesarrayYesConversation history
temperaturenumberNoSampling temperature
maxTokensnumberNoMax output tokens
streambooleanNoStream the response
toolsarrayNoFunction/tool definitions
toolChoicestring or objectNoTool selection mode
userId or userstringNoEnd-user identifier for attribution and Monetization
timeout_msnumberNoPer-request provider timeout in ms (1–300000, capped). Bounds one provider attempt
max_cost_usdnumberNoPer-request cost budget in USD. Unary calls fail before returning; streams surface it at final tally as HTTP 402 budget_exceeded

OpenAI-Compatible Response Shape

{
  "id": "chatcmpl-abc123",
  "object": "chat.completion",
  "created": 1677652288,
  "model": "gpt-4o",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "The capital of France is Paris."
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 13,
    "completion_tokens": 7,
    "total_tokens": 20
  }
}

Streaming

SDK Streaming

const stream = cencori.ai.chatStream({
  model: 'gpt-4o',
  messages: [{ role: 'user', content: 'Tell me a story.' }],
});
 
for await (const chunk of stream) {
  process.stdout.write(chunk.delta);
}

HTTP Streaming

curl -N https://api.cencori.com/v1/chat/completions \
  -H "Authorization: Bearer csk_..." \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4o",
    "messages": [{"role": "user", "content": "Tell me a story."}],
    "stream": true
  }'

Tool Calling

const response = await cencori.ai.chat({
  model: 'gpt-4o',
  messages: [{ role: 'user', content: 'What is the weather in Tokyo?' }],
  tools: [
    {
      type: 'function',
      function: {
        name: 'get_weather',
        description: 'Get weather for a location',
        parameters: {
          type: 'object',
          properties: {
            location: { type: 'string' },
          },
          required: ['location'],
        },
      },
    },
  ],
});
 
console.log(response.toolCalls);

Error Handling

The SDK throws typed errors — never a bare message. Every failure carries the HTTP status, the machine-readable code, the requestId for support correlation, and any retry hint the server sent. Transient failures (rate limits, 5xx, network blips) retry automatically with backoff, honoring the server's Retry-After; 4xx fail fast — except 429s, which retry then surface as RateLimitError with the hint intact.

import { CencoriError, RateLimitError } from 'cencori';
 
try {
  await cencori.ai.chat({
    model: 'gpt-4o',
    messages: [{ role: 'user', content: 'Hello!' }],
  });
} catch (error) {
  if (error instanceof RateLimitError) {
    console.log('back off for', error.retryAfterMs, 'ms');
  } else if (error instanceof CencoriError) {
    // error.statusCode, error.code, error.requestId, error.isRetryable
    console.log(error.code, error.requestId);
  }
}

Safety classification

Every chat response carries what the input guards verdicts — so a 200 whose model resisted an attack is distinguishable from a request that was never scanned:

{
  "choices": [ ... ],
  "safety": {
    "scanned": true,
    "input": { "safe": false, "layer": "jailbreak", "riskScore": 0.86, "reasons": ["..."] }
  }
}

Streams carry the same verdict as headers (X-Cencori-Safety-Scanned, X-Cencori-Safety-Input, X-Cencori-Safety-Score). scanned: false means the request ran the passthrough lane (your own safety layers apply). Only category-level reasons are exposed — never matched excerpts.

Auto-router (auto / cencori-auto)

Pass model: 'auto' to route by task across the project's active BYOK keys only — image content → vision model, code/tool calls → code-capable model, analytical prompts → reasoning model, otherwise a cheap fast model. No usable BYOK key returns 402 byok_required; it never falls back to managed keys and works at a zero credit balance. cencori-auto forces the BYOK router when a Tensor plan mapping would otherwise own bare auto. Details: BYOK auto-router.