One API for every model.
Route across hundreds of frontier models through a single API.
Everything in one gateway
AI Gateway combines six essential capabilities into one unified solution.
Built for developers
One endpoint.Every model.
Chat, images, voice, embeddings, and tools through one OpenAI-compatible API. Change providers without rewriting your stack.
- 150+
- Models
- 14+
- Providers
- <50 ms
- Overhead
import { Cencori } from "cencori"; const cencori = new Cencori({ apiKey: process.env.CENCORI_API_KEY,}); const response = await cencori.ai.chat({ model: "gpt-5.6-sol", messages: [{ role: "user", content: "Hello" }],}); console.log(response.content);“Take my moneyyy!!!!!!”
Universal AI Gateway
Connect any client to any model with a single, secure API.
Controls
Set limits before costs drift.
Not after the spike.
Hard spend caps
Set limits that actually stop traffic when a budget is reached, so one broken workflow does not become an expensive incident.
Budget alerts
Get warned before you hit the threshold, with alerts designed for teams managing production and staging separately.
Monetization
Allocate AI costs downstream with cleaner billing controls for products that need customer-level usage accounting.
Project-level control
Set different rules for production, staging, and internal tools so your budget policy matches how the organization actually works.
And much more
Everything behind your AI product, available through one gateway.
Responses API
Run stateful agent workflows through one OpenAI-compatible endpoint.
Built-in web search
Ground answers in first-party hybrid search results with citations.
File search
Retrieve relevant context from uploaded documents inside a response.
Code interpreter
Generate and execute code within the same agentic tool loop.
Function calling
Use provider-neutral tools with OpenAI-compatible schemas.
Structured outputs
Return JSON objects or strict schemas your application can trust.
Stateful responses
Continue conversations without rebuilding the full context each turn.
Prompt caching
Reuse exact and semantically similar prompts to cut cost and latency.
Embeddings
Generate vectors through OpenAI, Google, or Cohere from one API.
Vision inputs
Send images with text for multimodal understanding and extraction.
Custom providers
Connect Ollama, vLLM, LM Studio, or any compatible endpoint.
Streaming
Deliver token-by-token SSE responses across supported providers.
Get started in minutes
Set up your project
mkdir cencori-democd cencori-demonpm init -y


