Key facts
| Working memory | Message history held inside the context window |
| Compression | Summarise older turns to stay within token budgets |
| Long-term store | Vector or key-value store for facts across sessions |
| Retrieval | Hybrid search with metadata filters per user and tenant |
| Forgetting | Expiry, correction and deletion paths for stored memories |
| Privacy | Memory inherits the retention and residency rules of its store |
| Embeddings | The embeddings API is live for retrieval layers |
| Roadmap | File and batch endpoints are coming soon |
TL;DR
- Separate working memory, summaries and long-term facts — they have different rules.
- Write memories deliberately; never store every turn by default.
- Retrieve by relevance and recency together, filtered by user and tenant.
- Build expiry, correction and deletion in from day one.
- Measure memory's cost and latency impact, not just its helpfulness.
How it works, step by step
- Define what the agent must remember: preferences, decisions, entities and prior outcomes.
- Keep recent turns in the context window and summarise older ones into a running brief.
- Store durable facts in a vector or key-value store keyed by user and tenant.
- Retrieve on demand with hybrid search and metadata filters, not by loading everything.
- Add expiry, correction and deletion paths, and expose them to users where required.
- Cap memory size per user and measure how much retrieved context improves task success.
- Audit what was stored and why, and apply the same residency rules as your other data.
Try it yourself
Open the vector database comparison →
The three layers of agent memory
Working memory is the conversation: messages, tool results and recent context inside the model's window. It is fast and precise but finite, and every turn re-sends it. Compression is the next layer: summarising or extracting older turns into a compact brief so long runs stay affordable and coherent.
Long-term memory sits outside the model in a store you own: facts, preferences, entities and outcomes that should survive beyond the session. Retrieval decides what comes back per turn. Confusing these layers is the most common memory bug — teams either stuff everything into the prompt until it overflows, or store everything forever and retrieve noise.
What to remember and what to forget
Store what changes future behaviour: stable preferences, confirmed facts, decisions and their rationale. Do not store what can be looked up fresh, what is sensitive by default, or what is merely conversational. Every stored item should have a reason, a source and an expiry class.
- Write criteria: would this fact change a future action? If not, skip it.
- Scoping: memory belongs to a user and tenant, never to the agent globally.
- Correction: users must be able to see and fix what is remembered about them.
- Deletion: removal must propagate to indexes and caches, not only the primary store.
Cost, privacy and failure modes
Memory is not free. Every retrieved token raises prompt size on every turn, so retrieval quality directly affects cost per task. Track hit rate and relevance, and prune aggressively. Stale or contradictory memories are worse than none: a wrong preference remembered confidently will degrade every future run.
Because memories are user data, they carry retention, residency and access requirements. On Plugsky, the embeddings API and RAG are live, so retrieval layers run against the same OpenAI-compatible API as your agent, with 30+ models behind one key. Regulated teams can keep stores and endpoints inside their chosen region or on-prem using VPC, on-prem or air-gapped deployment. Chat, streaming and JSON mode are live; files and batch endpoints are coming soon. Plans are on the live pricing page.
Honest comparison
| Memory layer | Holds | Lifetime | Main risk |
|---|---|---|---|
| Context window | Recent messages and tool results | One run | Overflow and cost |
| Summary buffer | Compressed older turns | One long run | Lossy compression |
| Vector store | Facts, entities, past outcomes | Weeks to forever | Stale or wrong recall |
| Key-value profile | Explicit preferences and settings | Until changed | Drift and inconsistency |
| Audit log | What happened, for review | Policy-defined | Not a memory source |
Frequently asked questions
Do agents need long-term memory?
Only when continuity across sessions changes outcomes. Preferences, recurring entities and past decisions justify it; most conversations do not.
How much history should I keep in context?
Keep recent turns verbatim and compress older ones. Watch token growth per turn and test whether the extra context actually improves task success.
What database should I use for memory?
A vector store for semantic recall plus a relational or key-value store for structured facts and metadata. The right choice depends on scale, filtering needs and residency requirements.
How do I stop memory from becoming wrong?
Write deliberately, attach sources and timestamps, expire stale entries, and provide a correction path. Verify recalled facts rather than trusting them blindly.
Is memory a privacy risk?
Yes. Memories are personal data. Scope them per user, honour deletion and access requests, and apply the same residency and retention rules as the rest of your systems.
How does memory affect cost?
Retrieved context increases tokens on every turn, so poor retrieval quality raises cost per task. Measure relevance and prune regularly.
Can I build memory with Plugsky?
Yes. The embeddings API and RAG are live, so you can build retrieval memory on the same OpenAI-compatible API that serves your agent, with 30+ models on one key.