Agents

How do you give an AI agent memory?

Agent memory has layers: the context window holds the working conversation, summarisation compresses long runs, and a vector or key-value store carries facts across sessions. Write selectively, retrieve by relevance rather than recency alone, attach expiry and deletion rules, and scope memory per user. Memory is a product decision with privacy consequences, not just a database choice.

Key facts

Working memoryMessage history held inside the context window
CompressionSummarise older turns to stay within token budgets
Long-term storeVector or key-value store for facts across sessions
RetrievalHybrid search with metadata filters per user and tenant
ForgettingExpiry, correction and deletion paths for stored memories
PrivacyMemory inherits the retention and residency rules of its store
EmbeddingsThe embeddings API is live for retrieval layers
RoadmapFile and batch endpoints are coming soon

TL;DR

  • Separate working memory, summaries and long-term facts — they have different rules.
  • Write memories deliberately; never store every turn by default.
  • Retrieve by relevance and recency together, filtered by user and tenant.
  • Build expiry, correction and deletion in from day one.
  • Measure memory's cost and latency impact, not just its helpfulness.

How it works, step by step

  1. Define what the agent must remember: preferences, decisions, entities and prior outcomes.
  2. Keep recent turns in the context window and summarise older ones into a running brief.
  3. Store durable facts in a vector or key-value store keyed by user and tenant.
  4. Retrieve on demand with hybrid search and metadata filters, not by loading everything.
  5. Add expiry, correction and deletion paths, and expose them to users where required.
  6. Cap memory size per user and measure how much retrieved context improves task success.
  7. Audit what was stored and why, and apply the same residency rules as your other data.
1Define what theagent mustremember:2Keep recent turnsin the contextwindow and3Store durable factsin a vector orkey-value store4Retrieve on demandwith hybrid searchand metadata5Add expiry,correction anddeletion paths, and6Cap memory size peruser and measurehow much retrieved

Try it yourself

Open the vector database comparison →

The three layers of agent memory

Working memory is the conversation: messages, tool results and recent context inside the model's window. It is fast and precise but finite, and every turn re-sends it. Compression is the next layer: summarising or extracting older turns into a compact brief so long runs stay affordable and coherent.

Long-term memory sits outside the model in a store you own: facts, preferences, entities and outcomes that should survive beyond the session. Retrieval decides what comes back per turn. Confusing these layers is the most common memory bug — teams either stuff everything into the prompt until it overflows, or store everything forever and retrieve noise.

What to remember and what to forget

Store what changes future behaviour: stable preferences, confirmed facts, decisions and their rationale. Do not store what can be looked up fresh, what is sensitive by default, or what is merely conversational. Every stored item should have a reason, a source and an expiry class.

  • Write criteria: would this fact change a future action? If not, skip it.
  • Scoping: memory belongs to a user and tenant, never to the agent globally.
  • Correction: users must be able to see and fix what is remembered about them.
  • Deletion: removal must propagate to indexes and caches, not only the primary store.

Cost, privacy and failure modes

Memory is not free. Every retrieved token raises prompt size on every turn, so retrieval quality directly affects cost per task. Track hit rate and relevance, and prune aggressively. Stale or contradictory memories are worse than none: a wrong preference remembered confidently will degrade every future run.

Because memories are user data, they carry retention, residency and access requirements. On Plugsky, the embeddings API and RAG are live, so retrieval layers run against the same OpenAI-compatible API as your agent, with 30+ models behind one key. Regulated teams can keep stores and endpoints inside their chosen region or on-prem using VPC, on-prem or air-gapped deployment. Chat, streaming and JSON mode are live; files and batch endpoints are coming soon. Plans are on the live pricing page.

Honest comparison

Memory layerHoldsLifetimeMain risk
Context windowRecent messages and tool resultsOne runOverflow and cost
Summary bufferCompressed older turnsOne long runLossy compression
Vector storeFacts, entities, past outcomesWeeks to foreverStale or wrong recall
Key-value profileExplicit preferences and settingsUntil changedDrift and inconsistency
Audit logWhat happened, for reviewPolicy-definedNot a memory source

Frequently asked questions

Do agents need long-term memory?

Only when continuity across sessions changes outcomes. Preferences, recurring entities and past decisions justify it; most conversations do not.

How much history should I keep in context?

Keep recent turns verbatim and compress older ones. Watch token growth per turn and test whether the extra context actually improves task success.

What database should I use for memory?

A vector store for semantic recall plus a relational or key-value store for structured facts and metadata. The right choice depends on scale, filtering needs and residency requirements.

How do I stop memory from becoming wrong?

Write deliberately, attach sources and timestamps, expire stale entries, and provide a correction path. Verify recalled facts rather than trusting them blindly.

Is memory a privacy risk?

Yes. Memories are personal data. Scope them per user, honour deletion and access requests, and apply the same residency and retention rules as the rest of your systems.

How does memory affect cost?

Retrieved context increases tokens on every turn, so poor retrieval quality raises cost per task. Measure relevance and prune regularly.

Can I build memory with Plugsky?

Yes. The embeddings API and RAG are live, so you can build retrieval memory on the same OpenAI-compatible API that serves your agent, with 30+ models on one key.