Dashboard|

Memory API

Give any chat a memory. Retrieve what you know about a user before each reply, persist new facts after — with PII redaction and per-org isolation built in.

Holeacquisition LLC Memory gives any model persistent memory. Two calls — recall() to pull what you already know about a user, remember() to store what's new — wrap around your own OpenAI, Anthropic, or local inference; nothing has to route through our gateway. And when your chat already does, the same behavior collapses into a single memory field on the request. One engine, either path — with PII redaction, per-org isolation, audit logs, and forget built in.

New here? Start with the Holeacquisition LLC Memory overview — positioning, scopes, and pricing in one page.

A new chat is not a memory reset. It's a fresh conversation with a model that already knows the user.

Works with any model

recall returns a ready-to-inject system string (empty when there's nothing yet, so it's safe to spread unconditionally); remember extracts and stores the durable facts from a completed exchange. Any provider — no gateway required.

// 1. Pull what you know about this user.
const context = await cencori.memory.recall(userId, message);
 
// 2. Your own model call — any provider, no gateway required.
const reply = await openai.chat.completions.create({
  model: 'gpt-4o',
  messages: [
    ...(context ? [{ role: 'system', content: context }] : []),
    { role: 'user', content: message },
  ],
});
 
// 3. Store the new facts from the exchange.
await cencori.memory.remember(userId, {
  user: message,
  assistant: reply.choices[0].message.content,
});

Inference stays wherever it already runs — this is the drop-in path off Mem0/Zep.

On the gateway — one field

If your chat already routes through Holeacquisition LLC, recall and remember fuse into a single field. Retrieval and writeback both default to on; omit the field for a stateless chat.

const response = await cencori.chat.completions.create({
  model: 'gpt-4o',
  messages: [{ role: 'user', content: 'What did we agree about pricing?' }],
  memory: { userId: session.user.id },
});
 
console.log(response.choices[0].message.content);
console.log(response.memory?.retrieved); // facts injected into this reply

The memory field

FieldTypeDefaultDescription
userIdstring—End-user this memory belongs to. Scope defaults to user.
sessionIdstring—Pass with scope: "session" for ephemeral, per-chat memory.
scope"session" | "user" | "workspace" | "org""user"session clears on session end (Redis); user persists across sessions, devices, and new chats (Postgres + pgvector); workspace partitions a caller-supplied workspaceId (team memory); org keys the authenticated organization (company playbook, defaults automatically).
retrievebooleantruePull top-K relevant memories and inject them ahead of the user turn.
writebooleantrueAfter the reply, extract facts from the exchange and persist them.
topKnumber5How many memories to inject.
thresholdnumber0.7Minimum similarity for a memory to be injected.
namespacestring—Sub-scope partition. Maps cleanly to per-project sidebars for multi-project users.
asOfstring (ISO 8601)—Temporal recall — retrieve memory as it was valid at this instant, including facts later superseded. Omit for current state.
mode"inject" | "index""inject"How recalled memories are surfaced. inject = full contents in context. index = a compact table of contents the model fetches from on demand. See Progressive disclosure.
graphbooleantrueGraph-aware recall — when the query names an entity we know, also walk its relations and pull in connected facts similarity would miss. See Entity graph.
extractobject—Override fact extraction: { model, prompt, minImportance }.

The response carries a memory block:

{
  "choices": [ ... ],
  "memory": {
    "retrieved": [{ "id": "mem_xxx", "score": 0.82, "content": "Prefers TypeScript" }],
    "written": [],
    "write_status": "pending",
    "write_request_id": "f2f55531-0a92-4d45-8e78-6624d390c75d"
  }
}

Writeback runs asynchronously after the response, so written is empty and write_status is pending on the same request — the facts are available on the next turn. write_request_id is the receipt: poll GET /v1/memory/writes/:requestId until status leaves pending (success carries extracted/written counts, error carries the reason). The streaming response also sets X-Cencori-Memory-Retrieved and X-Cencori-Memory-Write-Request headers with the injected count and the same receipt.

Direct memory endpoints

The full REST surface. remember and search back the recall/remember path above; the rest cover settings pages, GDPR panels, and manual seeding.

PathPurpose
POST /v1/memory/writeWrite a single scoped memory (raw content, no extraction)
POST /v1/memory/write/batchWrite up to 50 memories in one call (one quota check, one embedding call)
POST /v1/memory/forgetForget memories by filter (namespace / before / ids) — hard delete, audit-logged
POST /v1/memory/rememberExtract durable facts from a { user, assistant } exchange and store them
POST /v1/memory/searchSemantic search over a user's memories
GET /v1/memory/listPaginate a user's memories
GET /v1/memory/:idFetch one memory's full content by id (cencori.memory.fetch(id))
GET /v1/memory/forget-suggestionsPropose stale, low-value memories worth forgetting (candidates only)
POST /v1/memory/graphExtract entities + relations from an exchange into the memory graph
GET /v1/memory/graphTraverse the graph outward from an entity (multi-hop)
GET /v1/memory/entitiesList the entities in a user's memory graph
DELETE /v1/memory/:idForget a memory — a hard delete, not an annotation
GET /v1/memory/writes/:requestIdConfirm an async chat writeback (pending → success / error)
POST /v1/memory/exportGDPR export — portable dump of a scope key (cursor-paginated)
// Seed a fact
await cencori.memory.write({
  userId: session.user.id,
  content: 'Prefers dark mode. Uses TypeScript primarily.',
});
 
// Search
const { results } = await cencori.memory.searchUser({
  userId: session.user.id,
  query: 'ui preferences',
  topK: 3,
});
 
// Forget
await cencori.memory.forget(results[0].id);

Write paths report what they spent: remember / write / write/batch responses include costUsd (plus model and the actual provider), and search reports its embedding cost. A 201 with extracted: 0 means the exchange held nothing durable — not a failure. On the rare empty first completion the extractor retries once before accepting that verdict.

Batch writes, forget-by-filter, GDPR export

// Seed up to 50 facts in one call — one quota check, one embedding call.
await cencori.memory.writeBatch({
  userId: session.user.id,
  memories: [{ content: 'Prefers dark mode' }, { content: 'Uses TypeScript', importance: 0.9 }],
});
 
// Forget everything before a date (namespace / ids filters also work).
// Real row removal, audit-logged — not an annotation.
await cencori.memory.forgetByFilter({ userId: session.user.id, before: '2026-01-01T00:00:00Z' });
 
// GDPR export — walk nextCursor until truncated is false, hand the dump over.
const dump = await cencori.memory.export({ userId: session.user.id });

Quotas — stored count and operation allowances

Two independent gates. Stored count caps rows per project (free 1,000 / pro 100,000 / team 500,000); writes 429 memory_quota_exceeded while reads keep working. Operation allowances cap managed-LLM volume per project per month and per end-user per day (free 10k/2k + 200/100 daily, pro 500k/100k + 2000/1000 daily); exhaustion 429s memory_ops_quota_exceeded. Both carry an upgradeUrl, used/limit, and (for ops) retryAfterMs. Chat retrieval fails open past either gate — the reply still sends, memory is just skipped.

Pilot and enterprise contracts can pin a project below (or above) its tier with custom monthly allowances (maxSearchesMonthly / maxWritesMonthly in project memory settings, dashboard Controls tab). An explicit custom cap always wins — including over unlimited tiers.

Temporal recall — asOf

Memory is bi-temporal: when a fact changes, the old one isn't deleted, it's superseded. Current recall returns only what's true now, but you can ask what was true at any past instant with asOf.

// What does the user prefer now? → "Rust"
const now = await cencori.memory.recall(userId, 'preferred language');
 
// What did they prefer back in January? → "Python" (the superseded fact)
const then = await cencori.memory.searchUser({
  userId,
  query: 'preferred language',
  asOf: '2026-01-15T00:00:00Z',
});

asOf works on the memory field, searchUser, and recall. Superseded facts never surface in normal recall — only when you explicitly ask for a point in time.

Progressive disclosure — index mode

Injecting the full text of every recalled memory on every turn buries the signal. Index mode shows the model a compact table of contents — id + one-line summary — and lets it pull the full note only when it needs it. Best for agents and long-running sessions.

// Inject mode (default): full contents in context.
const block = await cencori.memory.recall(userId, message);
 
// Index mode: a table of contents. The model fetches full notes on demand.
const toc = await cencori.memory.recall(userId, message, { mode: 'index' });
// → "Memory index — what you know about this user (summaries only):
//    - [mem_abc123] User switched their primary language to Rust
//    - [mem_def456] User is building a fintech app called Ledgerkit"
 
// Fetch a full note by id when a summary is relevant.
const memory = await cencori.memory.fetch('mem_abc123');

For tool-calling agents, register MEMORY_FETCH_TOOL so the model fetches autonomously:

import { Cencori, MEMORY_FETCH_TOOL } from 'cencori';
 
const response = await cencori.chat.completions.create({
  model: 'gpt-4o',
  messages,
  memory: { userId, mode: 'index' }, // inject the index
  tools: [MEMORY_FETCH_TOOL],         // let the model pull full notes
});
// When the model calls memory_fetch({ id }), resolve it with cencori.memory.fetch(id).

Forgetting — suggestions, never silent

A memory store is only as good as its forgetting, but Holeacquisition LLC never deletes a user's memories on its own. forgetSuggestions returns candidates — stale, low-strength, long-unused memories — ranked weakest-first. You decide what to actually forget.

const { suggestions } = await cencori.memory.forgetSuggestions({
  userId: user.id,
  minIdleDays: 90, // only memories untouched for 90+ days
});
 
// Review, then forget explicitly.
for (const s of suggestions) {
  await cencori.memory.forget(s.id);
}

Strength is query-independent: it rewards importance, recent use, and repeated usefulness, and decays neglect. Contradicted facts are already superseded (see Temporal recall), so they never show up here — this is for genuine dead weight.

Entity graph — multi-hop recall

Flat facts answer "what is Sarah's role?" but not "who does Sarah report to, and where do they work?" — that's a walk across relationships. The graph layer extracts entities (people, orgs, projects, places) and typed relations from exchanges, resolves them (so "John from Zap" and "John Smith" are one node), and lets you traverse.

// Build the graph from an exchange (entities are resolved + merged).
await cencori.memory.rememberGraph({
  userId: user.id,
  user: 'Sarah moved to Northwind and now reports to Marcus.',
  assistant: 'Got it.',
});
 
// Traverse outward from an entity.
const { seed, nodes, edges } = await cencori.memory.graph({
  userId: user.id,
  entity: 'Sarah',
  hops: 2,
});
// nodes: [{ name: 'Sarah', hops: 0 }, { name: 'Marcus', hops: 1, path: ['reports_to'] }, ...]
// edges: [{ source: 'Sarah', relation: 'reports_to', target: 'Marcus' }, ...]
 
// List what the graph knows about a user.
const { entities } = await cencori.memory.entities({ userId: user.id });

Entity extraction runs on Holeacquisition LLC's managed model. The graph shares the same tenant boundary as every other memory — no entity or edge ever crosses orgs.

The graph maintains itself

You don't have to call rememberGraph to get a graph. Every memory writeback — memory on a chat completion, a session turn, or memory.remember() — extracts entities and relations from the same exchange and links them to the facts it just stored. rememberGraph is for building the graph from text you aren't storing as memories.

Recall walks it for you

Graph traversal isn't a separate query you have to make. When a recall names an entity we know, retrieval walks outward from it and adds connected facts that the query vector would never have matched:

// Stored across two different sessions:
//   "Sarah works at Zap Corp."      → Sarah —works_at→ Zap Corp
//   "Zap Corp's office is Berlin."  → Zap Corp —located_in→ Berlin
 
const { results } = await cencori.memory.searchUser({
  userId: user.id,
  query: 'where does Sarah work from?',
});
 
// [
//   { content: 'Sarah works at Zap Corp.',     score: 0.71, source: 'vector' },
//   { content: "Zap Corp's office is Berlin.", score: 0,    source: 'graph', hops: 2 },
// ]

The Berlin fact shares no words with the question — similarity alone never returns it. source tells you how each memory was reached, and hops how far the walk went. Graph hits carry a score of 0 because they were reached by relation, not by match; they're capped so a walk supplements recall instead of flooding it.

Set graph: false on the memory field (or on searchUser / recall) for strictly vector recall. Turning the graph off for a whole project — which also skips the extraction call on write — is a project-level setting.

Use with any provider — recall + remember

You don't have to route inference through Holeacquisition LLC to use memory. Keep your existing OpenAI / Anthropic / Bedrock / LangChain setup exactly as-is and bolt memory on beside it with two helpers:

import { Cencori } from 'cencori';
const cencori = new Cencori({ apiKey: process.env.CENCORI_API_KEY });
 
// 1. Recall — returns an inject-ready system string ('' if nothing yet)
const context = await cencori.memory.recall(userId, userMessage);
 
// 2. Your existing model call — untouched
const completion = await openai.chat.completions.create({
  model: 'gpt-4o',
  messages: [
    ...(context ? [{ role: 'system', content: context }] : []),
    { role: 'user', content: userMessage },
  ],
});
const reply = completion.choices[0].message.content;
 
// 3. Remember — we extract the durable facts from the exchange and store them
await cencori.memory.remember(userId, { user: userMessage, assistant: reply });

remember runs the same fact extraction the gateway uses (customizable via extract), redacts before writing, and enforces org isolation and quota — all without any inference going through Holeacquisition LLC. recall is a thin wrapper over search that formats results the same way the gateway injects them.

Under the hood these map to POST /v1/memory/remember and POST /v1/memory/search. Embeddings use your project's OpenAI key if configured (BYOK), or Holeacquisition LLC's managed key otherwise — so zero provider setup is required to start.

React — memory-aware chat

cencori/react ships a drop-in <Chat> component. Pass memory and the conversation is stateful across sessions:

import { Chat } from 'cencori/react';
 
<Chat
  model="gpt-4o"
  apiKey={process.env.NEXT_PUBLIC_CENCORI_KEY!}
  memory={{ userId: session.user.id }}
/>

It streams by default and shows a subtle "recalled N memories" indicator so users know the model remembers them.

The useMemory hook

For custom UI — a "what we remember about you" panel with right-to-be-forgotten:

import { useMemory } from 'cencori/react';
 
function MemorySettings({ userId }: { userId: string }) {
  const { memories, forget, exportAll, loading } = useMemory({
    userId,
    apiKey: process.env.NEXT_PUBLIC_CENCORI_KEY,
  });
 
  if (loading) return <p>Loading…</p>;
 
  return (
    <div>
      <h2>What we remember about you</h2>
      {memories.map((m) => (
        <div key={m.id}>
          <span>{m.content}</span>
          <button onClick={() => forget(m.id)}>Forget</button>
        </div>
      ))}
      <button onClick={exportAll}>Download my data</button>
    </div>
  );
}

useMemory returns { memories, loading, error, refresh, write, search, forget, exportAll }.

Governance

  • PII redaction runs before writeback. The raw value never persists — the gateway's PII pipeline tokenizes SSNs, redacts names, etc. on any string headed for storage.
  • Injection screening runs before writeback. Stored facts are scanned for instruction-override language ("ignore previous instructions" and kin). Hits are dropped, never persisted — critical for shared workspace/org scopes, where one writer's poison would otherwise reach every reader. Recall screens again as defense in depth for older rows.
  • Injected context is marked untrusted. Every recalled block labels stored notes UNTRUSTED data-never-instructions with an explicit conflict rule: notes that fight system instructions are ignored by the model.
  • Hard boundary at the organization. Under no code path can a memory written by one org be read by another. Enforced in SQL and property-tested against the live database.
  • Forget is real deletion. DELETE /v1/memory/:id removes the row and its embedding.
  • Opt-in per request. No memory field means no memory. Nothing is stored unless you ask for it.

Limits

Memory is metered by count per project (each memory has a 10KB content soft cap). Free projects get 1,000 memories with 30-day session / 90-day user retention. At 100% the gateway returns 429 memory_quota_exceeded on writes only — reads keep working, so your product never silently forgets its users.