Key facts
| Pattern | Retrieval exposed as a callable tool, not a fixed pre-step |
| Retrieval quality | Hybrid keyword plus vector search, then rerank |
| Chunking | Structure-aware chunks with metadata and stable IDs |
| Citations | Return source IDs with passages and surface them in answers |
| Freshness | Index timestamps and re-index triggers for changing sources |
| Embeddings | The embeddings API and RAG are live on Plugsky |
| Cost | Each retrieval adds tokens and vector queries on relevant turns |
| Roadmap | File and batch endpoints are coming soon |
TL;DR
- Make retrieval a tool so the agent can search again when the first result misses.
- Hybrid search plus reranking beats vector-only search on real questions.
- Chunk on document structure and keep stable IDs for citations.
- Track index freshness; stale retrieval is worse than admitting ignorance.
- Require citations and score groundedness in your evaluation set.
How it works, step by step
- Index the sources you actually trust, with metadata for source, section and time.
- Chunk on structure — headings, sections, records — and keep stable IDs.
- Expose retrieval as a tool with a query parameter and filters such as date or source.
- Add hybrid search and a reranker, then return passages with IDs and snippets.
- Let the agent rewrite queries and retrieve again when coverage is poor.
- Require answers to cite source IDs, and refuse when retrieval finds nothing useful.
- Measure retrieval precision, groundedness and cost per question in your evaluations.
Try it yourself
Open the RAG architecture builder →
RAG as a tool changes the loop
Classic RAG retrieves once, then answers. Agentic RAG turns retrieval into a capability the model can invoke several times: search, read, notice a gap, refine the query, search again, compare passages and synthesise. That extra control helps multi-hop questions such as comparing two policies or tracing a value through several documents.
The cost is that retrieval now happens per relevant turn, increasing tokens and vector queries. Budgets and relevance thresholds matter more, not less, because an agent that searches compulsively will spend its way to the same answer a single well-formed query would have found.
Building retrieval that survives production
- Hybrid search: combine keyword and vector scores; exact terms still matter for codes, names and numbers.
- Reranking: a cross-encoder reranker over the top candidates materially improves precision.
- Metadata filters: restrict by tenant, date, source and access level before ranking.
- Chunk design: structure-aware chunks with overlap, sized to the question type, not a global default.
- Freshness: re-index on source changes and expose the document date in results.
Instrument retrieval separately from generation: precision at k, recall on known questions and rerank lift tell you which layer to fix.
Grounding, citations and honesty
The point of RAG in an agent is not longer answers; it is verifiable ones. Require the agent to cite the source IDs behind each claim, and treat uncited statements as failures in evaluation. When retrieval returns nothing relevant, the correct behaviour is to say so and suggest next steps, not to improvise.
Access control belongs in retrieval, not in the prompt: filter documents the user may not see before they reach the model. On Plugsky, the embeddings API and RAG are live, so indexing and retrieval run against the same OpenAI-compatible platform as your agent loop, with 30+ models on one key. Chat, streaming, JSON mode and function calling are live; file and batch endpoints are coming soon. Plans, including the free plugsky-micro and plugsky-lite tier, are on the live pricing page.
Honest comparison
| Approach | Retrieval | Strengths | Weaknesses |
|---|---|---|---|
| Classic RAG | One query before answering | Simple, cheap, predictable | Weak on multi-hop questions |
| Agentic RAG | Tool calls, multiple searches | Handles comparisons and gaps | Higher cost, needs budgets |
| Fine-tuning alone | None | Style and format stability | Cannot cite or update facts |
| Long context only | Whole corpus in prompt | No index to maintain | Expensive and hard to control access |
| Hybrid | Retrieval plus tuning | Best quality when tuned | Most moving parts |
Frequently asked questions
Is agentic RAG always better?
No. It costs more per question because retrieval happens repeatedly. Use classic single-pass RAG for simple lookups and agentic retrieval where questions need comparison or multi-step evidence.
Why add a reranker?
Vector similarity is a coarse filter. A reranker scores the top candidates against the actual query, which usually improves precision enough to justify the extra latency.
How should I chunk documents?
Split on structure such as headings, sections and records, with a small overlap. Avoid fixed-size chunks that cut sentences and separate questions from their answers.
How do I handle stale documents?
Store a timestamp with each chunk, re-index when sources change, and let the agent filter or deprioritise old content. Surfacing dates in answers builds trust.
Can the agent cite sources?
Yes, if you return stable source IDs with passages and instruct answers to reference them. Treat uncited claims as evaluation failures.
Should access control be in the prompt?
No. Filter documents by user and tenant permissions during retrieval, before content reaches the model, so restricted text is never in the context.
Does Plugsky provide embeddings?
Yes. The embeddings API is live and supports RAG pipelines on the same OpenAI-compatible platform that runs your agent.