Memory API
Give any chat a memory. Retrieve what you know about a user before each reply, persist new facts after — with PII redaction and per-org isolation built in.
Holeacquisition LLC Memory gives any model persistent memory. Two calls — recall() to pull what you already know about a user, remember() to store what's new — wrap around your own OpenAI, Anthropic, or local inference; nothing has to route through our gateway. And when your chat already does, the same behavior collapses into a single memory field on the request. One engine, either path — with PII redaction, per-org isolation, audit logs, and forget built in.
New here? Start with the Holeacquisition LLC Memory overview — positioning, scopes, and pricing in one page.
A new chat is not a memory reset. It's a fresh conversation with a model that already knows the user.
Works with any model
recall returns a ready-to-inject system string (empty when there's nothing yet, so it's safe to spread unconditionally); remember extracts and stores the durable facts from a completed exchange. Any provider — no gateway required.
// 1. Pull what you know about this user.
const context = await cencori.memory.recall(userId, message);
// 2. Your own model call — any provider, no gateway required.
const reply = await openai.chat.completions.create({
model: 'gpt-4o',
messages: [
...(context ? [{ role: 'system', content: context }] : []),
{ role: 'user', content: message },
],
});
// 3. Store the new facts from the exchange.
await cencori.memory.remember(userId, {
user: message,
assistant: reply.choices[0].message.content,
});Inference stays wherever it already runs — this is the drop-in path off Mem0/Zep.
On the gateway — one field
If your chat already routes through Holeacquisition LLC, recall and remember fuse into a single field. Retrieval and writeback both default to on; omit the field for a stateless chat.
const response = await cencori.chat.completions.create({
model: 'gpt-4o',
messages: [{ role: 'user', content: 'What did we agree about pricing?' }],
memory: { userId: session.user.id },
});
console.log(response.choices[0].message.content);
console.log(response.memory?.retrieved); // facts injected into this replyThe memory field
The response carries a memory block:
{
"choices": [ ... ],
"memory": {
"retrieved": [{ "id": "mem_xxx", "score": 0.82, "content": "Prefers TypeScript" }],
"written": [],
"write_status": "pending",
"write_request_id": "f2f55531-0a92-4d45-8e78-6624d390c75d"
}
}Writeback runs asynchronously after the response, so written is empty and write_status is pending on the same request — the facts are available on the next turn. write_request_id is the receipt: poll GET /v1/memory/writes/:requestId until status leaves pending (success carries extracted/written counts, error carries the reason). The streaming response also sets X-Cencori-Memory-Retrieved and X-Cencori-Memory-Write-Request headers with the injected count and the same receipt.
Direct memory endpoints
The full REST surface. remember and search back the recall/remember path above; the rest cover settings pages, GDPR panels, and manual seeding.
// Seed a fact
await cencori.memory.write({
userId: session.user.id,
content: 'Prefers dark mode. Uses TypeScript primarily.',
});
// Search
const { results } = await cencori.memory.searchUser({
userId: session.user.id,
query: 'ui preferences',
topK: 3,
});
// Forget
await cencori.memory.forget(results[0].id);Write paths report what they spent: remember / write / write/batch
responses include costUsd (plus model and the actual provider), and
search reports its embedding cost. A 201 with extracted: 0 means the
exchange held nothing durable — not a failure. On the rare empty first
completion the extractor retries once before accepting that verdict.
Batch writes, forget-by-filter, GDPR export
// Seed up to 50 facts in one call — one quota check, one embedding call.
await cencori.memory.writeBatch({
userId: session.user.id,
memories: [{ content: 'Prefers dark mode' }, { content: 'Uses TypeScript', importance: 0.9 }],
});
// Forget everything before a date (namespace / ids filters also work).
// Real row removal, audit-logged — not an annotation.
await cencori.memory.forgetByFilter({ userId: session.user.id, before: '2026-01-01T00:00:00Z' });
// GDPR export — walk nextCursor until truncated is false, hand the dump over.
const dump = await cencori.memory.export({ userId: session.user.id });Quotas — stored count and operation allowances
Two independent gates. Stored count caps rows per project (free 1,000 / pro
100,000 / team 500,000); writes 429 memory_quota_exceeded while reads keep
working. Operation allowances cap managed-LLM volume per project per month
and per end-user per day (free 10k/2k + 200/100 daily, pro 500k/100k +
2000/1000 daily); exhaustion 429s memory_ops_quota_exceeded. Both carry an
upgradeUrl, used/limit, and (for ops) retryAfterMs. Chat retrieval
fails open past either gate — the reply still sends, memory is just skipped.
Pilot and enterprise contracts can pin a project below (or above) its tier
with custom monthly allowances (maxSearchesMonthly / maxWritesMonthly in
project memory settings, dashboard Controls tab). An explicit custom cap
always wins — including over unlimited tiers.
Temporal recall — asOf
Memory is bi-temporal: when a fact changes, the old one isn't deleted, it's superseded. Current recall returns only what's true now, but you can ask what was true at any past instant with asOf.
// What does the user prefer now? → "Rust"
const now = await cencori.memory.recall(userId, 'preferred language');
// What did they prefer back in January? → "Python" (the superseded fact)
const then = await cencori.memory.searchUser({
userId,
query: 'preferred language',
asOf: '2026-01-15T00:00:00Z',
});asOf works on the memory field, searchUser, and recall. Superseded facts never surface in normal recall — only when you explicitly ask for a point in time.
Progressive disclosure — index mode
Injecting the full text of every recalled memory on every turn buries the signal. Index mode shows the model a compact table of contents — id + one-line summary — and lets it pull the full note only when it needs it. Best for agents and long-running sessions.
// Inject mode (default): full contents in context.
const block = await cencori.memory.recall(userId, message);
// Index mode: a table of contents. The model fetches full notes on demand.
const toc = await cencori.memory.recall(userId, message, { mode: 'index' });
// → "Memory index — what you know about this user (summaries only):
// - [mem_abc123] User switched their primary language to Rust
// - [mem_def456] User is building a fintech app called Ledgerkit"
// Fetch a full note by id when a summary is relevant.
const memory = await cencori.memory.fetch('mem_abc123');For tool-calling agents, register MEMORY_FETCH_TOOL so the model fetches autonomously:
import { Cencori, MEMORY_FETCH_TOOL } from 'cencori';
const response = await cencori.chat.completions.create({
model: 'gpt-4o',
messages,
memory: { userId, mode: 'index' }, // inject the index
tools: [MEMORY_FETCH_TOOL], // let the model pull full notes
});
// When the model calls memory_fetch({ id }), resolve it with cencori.memory.fetch(id).Forgetting — suggestions, never silent
A memory store is only as good as its forgetting, but Holeacquisition LLC never deletes a
user's memories on its own. forgetSuggestions returns candidates — stale,
low-strength, long-unused memories — ranked weakest-first. You decide what to
actually forget.
const { suggestions } = await cencori.memory.forgetSuggestions({
userId: user.id,
minIdleDays: 90, // only memories untouched for 90+ days
});
// Review, then forget explicitly.
for (const s of suggestions) {
await cencori.memory.forget(s.id);
}Strength is query-independent: it rewards importance, recent use, and repeated usefulness, and decays neglect. Contradicted facts are already superseded (see Temporal recall), so they never show up here — this is for genuine dead weight.
Entity graph — multi-hop recall
Flat facts answer "what is Sarah's role?" but not "who does Sarah report to, and where do they work?" — that's a walk across relationships. The graph layer extracts entities (people, orgs, projects, places) and typed relations from exchanges, resolves them (so "John from Zap" and "John Smith" are one node), and lets you traverse.
// Build the graph from an exchange (entities are resolved + merged).
await cencori.memory.rememberGraph({
userId: user.id,
user: 'Sarah moved to Northwind and now reports to Marcus.',
assistant: 'Got it.',
});
// Traverse outward from an entity.
const { seed, nodes, edges } = await cencori.memory.graph({
userId: user.id,
entity: 'Sarah',
hops: 2,
});
// nodes: [{ name: 'Sarah', hops: 0 }, { name: 'Marcus', hops: 1, path: ['reports_to'] }, ...]
// edges: [{ source: 'Sarah', relation: 'reports_to', target: 'Marcus' }, ...]
// List what the graph knows about a user.
const { entities } = await cencori.memory.entities({ userId: user.id });Entity extraction runs on Holeacquisition LLC's managed model. The graph shares the same tenant boundary as every other memory — no entity or edge ever crosses orgs.
The graph maintains itself
You don't have to call rememberGraph to get a graph. Every memory writeback —
memory on a chat completion, a session turn, or memory.remember() — extracts
entities and relations from the same exchange and links them to the facts it just
stored. rememberGraph is for building the graph from text you aren't storing as
memories.
Recall walks it for you
Graph traversal isn't a separate query you have to make. When a recall names an entity we know, retrieval walks outward from it and adds connected facts that the query vector would never have matched:
// Stored across two different sessions:
// "Sarah works at Zap Corp." → Sarah —works_at→ Zap Corp
// "Zap Corp's office is Berlin." → Zap Corp —located_in→ Berlin
const { results } = await cencori.memory.searchUser({
userId: user.id,
query: 'where does Sarah work from?',
});
// [
// { content: 'Sarah works at Zap Corp.', score: 0.71, source: 'vector' },
// { content: "Zap Corp's office is Berlin.", score: 0, source: 'graph', hops: 2 },
// ]The Berlin fact shares no words with the question — similarity alone never
returns it. source tells you how each memory was reached, and hops how far
the walk went. Graph hits carry a score of 0 because they were reached by
relation, not by match; they're capped so a walk supplements recall instead of
flooding it.
Set graph: false on the memory field (or on searchUser / recall) for
strictly vector recall. Turning the graph off for a whole project — which also
skips the extraction call on write — is a project-level setting.
Use with any provider — recall + remember
You don't have to route inference through Holeacquisition LLC to use memory. Keep your existing OpenAI / Anthropic / Bedrock / LangChain setup exactly as-is and bolt memory on beside it with two helpers:
import { Cencori } from 'cencori';
const cencori = new Cencori({ apiKey: process.env.CENCORI_API_KEY });
// 1. Recall — returns an inject-ready system string ('' if nothing yet)
const context = await cencori.memory.recall(userId, userMessage);
// 2. Your existing model call — untouched
const completion = await openai.chat.completions.create({
model: 'gpt-4o',
messages: [
...(context ? [{ role: 'system', content: context }] : []),
{ role: 'user', content: userMessage },
],
});
const reply = completion.choices[0].message.content;
// 3. Remember — we extract the durable facts from the exchange and store them
await cencori.memory.remember(userId, { user: userMessage, assistant: reply });remember runs the same fact extraction the gateway uses (customizable via extract), redacts before writing, and enforces org isolation and quota — all without any inference going through Holeacquisition LLC. recall is a thin wrapper over search that formats results the same way the gateway injects them.
Under the hood these map to POST /v1/memory/remember and POST /v1/memory/search. Embeddings use your project's OpenAI key if configured (BYOK), or Holeacquisition LLC's managed key otherwise — so zero provider setup is required to start.
React — memory-aware chat
cencori/react ships a drop-in <Chat> component. Pass memory and the conversation is stateful across sessions:
import { Chat } from 'cencori/react';
<Chat
model="gpt-4o"
apiKey={process.env.NEXT_PUBLIC_CENCORI_KEY!}
memory={{ userId: session.user.id }}
/>It streams by default and shows a subtle "recalled N memories" indicator so users know the model remembers them.
The useMemory hook
For custom UI — a "what we remember about you" panel with right-to-be-forgotten:
import { useMemory } from 'cencori/react';
function MemorySettings({ userId }: { userId: string }) {
const { memories, forget, exportAll, loading } = useMemory({
userId,
apiKey: process.env.NEXT_PUBLIC_CENCORI_KEY,
});
if (loading) return <p>Loading…</p>;
return (
<div>
<h2>What we remember about you</h2>
{memories.map((m) => (
<div key={m.id}>
<span>{m.content}</span>
<button onClick={() => forget(m.id)}>Forget</button>
</div>
))}
<button onClick={exportAll}>Download my data</button>
</div>
);
}useMemory returns { memories, loading, error, refresh, write, search, forget, exportAll }.
Governance
- PII redaction runs before writeback. The raw value never persists — the gateway's PII pipeline tokenizes SSNs, redacts names, etc. on any string headed for storage.
- Injection screening runs before writeback. Stored facts are scanned for instruction-override language ("ignore previous instructions" and kin). Hits are dropped, never persisted — critical for shared workspace/org scopes, where one writer's poison would otherwise reach every reader. Recall screens again as defense in depth for older rows.
- Injected context is marked untrusted. Every recalled block labels stored notes UNTRUSTED data-never-instructions with an explicit conflict rule: notes that fight system instructions are ignored by the model.
- Hard boundary at the organization. Under no code path can a memory written by one org be read by another. Enforced in SQL and property-tested against the live database.
- Forget is real deletion.
DELETE /v1/memory/:idremoves the row and its embedding. - Opt-in per request. No
memoryfield means no memory. Nothing is stored unless you ask for it.
Limits
Memory is metered by count per project (each memory has a 10KB content soft cap). Free projects get 1,000 memories with 30-day session / 90-day user retention. At 100% the gateway returns 429 memory_quota_exceeded on writes only — reads keep working, so your product never silently forgets its users.