Key facts
| RAG endpoints | POST /v1/embeddings, POST /v1/rag/collections, POST /v1/rag/query |
| Retrieval modes | Keyword, vector and hybrid search with optional cross-encoder reranking |
| Agent building blocks | Chat, streaming, JSON mode and function calling are live |
| Orchestration | OpenAI-compatible tool calls; works with LangChain, LlamaIndex and Haystack |
| Citations | Every query returns ranked chunks with source attribution |
| Models | 30+ models behind one API, from free tiers to frontier reasoning |
| Deployment | Hosted, VPC, on-prem and air-gapped options |
| Product status | Live |
TL;DR
- Agentic RAG replaces one retrieval pass with a bounded plan, retrieve, critique loop.
- Use it for multi-hop questions, ambiguous phrasing and tasks mixing documents with tools.
- Bound the loop with a maximum step count, a token budget and request timeouts.
- Keep citations on every retrieval so each hop stays auditable.
- Plugsky exposes chat, function calling, embeddings and RAG on one OpenAI-compatible API.
How it works, step by step
- Start with one-shot retrieval and record where it fails on real questions.
- Define the toolset the model may call: collection query, filters and any external API.
- Set a maximum number of retrieval steps and a per-request token budget.
- Rewrite or decompose the query before each retrieval and log the rewritten text.
- Require the model to cite chunk IDs in the final answer.
- Evaluate the loop against a fixed question set before enabling it in production.
- Route planning steps to a fast model and the final synthesis to a stronger one.
Try it yourself
Open the RAG architecture builder →
What agentic RAG adds to a retrieval pipeline
Classical RAG runs one retrieval pass: embed the question, fetch top-k chunks, generate an answer. Agentic RAG wraps that call in a reasoning loop. The model plans the question, may rewrite or decompose it, calls the retrieval tool, inspects the returned chunks, and decides whether one more pass is needed. This matters for multi-hop questions such as comparing two policies, for ambiguous queries that need clarification through search, and for tasks that combine a document corpus with a live system call.
The cost is straightforward: every extra step is another model call and another retrieval, so latency and token use grow with loop depth. The design goal is not to let the model loop freely but to give it a small, bounded toolbox and a clear stop condition.
The loop: plan, retrieve, critique, answer
A practical agentic loop has four phases. Plan: turn the user request into a search plan, including filters and expected sources. Retrieve: call the collection query endpoint with top_k and optional reranking. Critique: check whether the returned chunks actually contain the missing facts, and rewrite the query if they do not. Answer: compose the response from cited chunks only, and refuse when the corpus does not support an answer.
Each phase is observable if you log the tool call, the query text and the chunk IDs returned. That log is what makes an agentic pipeline debuggable and auditable later, and it is also the raw material for evaluation sets.
When the loop is worth the latency
Use agentic RAG when single-pass recall is the bottleneck: questions that need two or more facts joined, corpora split across collections, or workflows where the model must choose between search and another tool. Stay with one-shot retrieval for high-volume support answers, latency-sensitive chat, and simple lookups where a reranked top-10 already contains the answer.
A useful middle ground is a single rewrite step: retrieve once, ask the model whether the evidence is sufficient, and allow at most one corrected retrieval. Most of the quality gain arrives with the second attempt, while latency stays predictable and easy to explain to stakeholders.
Building agentic RAG on Plugsky
Plugsky provides the pieces on one OpenAI-compatible API: chat completions for reasoning, live function calling for tool use, POST /v1/embeddings for vectors, and collections plus queries for retrieval with citations. Existing LangChain, LlamaIndex or Haystack agents can point at the same base URL and keep their code. Route cheap planning steps to a fast model and the final answer to a stronger one across 30+ models. Start on the free plan with plugsky-micro and plugsky-lite, or use the 14-day full-access trial; current plans are on the live pricing page.
Honest comparison
| Capability | One-shot RAG | Agentic RAG | Plugsky support |
|---|---|---|---|
| Retrieval passes | One | One or many, decided by the model | Manual loop over /v1/rag/query |
| Query handling | Original query only | Rewritten and decomposed | Any model on the chat endpoint |
| Tool use | Retrieval only | Retrieval plus external tools | Function calling is live |
| Latency profile | Lower and predictable | Higher, grows with steps | Route steps across model tiers |
| Best for | Simple lookups | Multi-hop and ambiguous tasks | Both patterns on one API |
Frequently asked questions
What is agentic RAG in one sentence?
Agentic RAG is a retrieval pipeline where a model plans, searches, inspects results and decides whether to search again before answering.
Do I need a framework to build it?
No. Any loop that calls an OpenAI-compatible chat endpoint with tools and a query endpoint works. Frameworks such as LangChain, LlamaIndex and Haystack add orchestration conveniences.
Does Plugsky support tool calls?
Yes. Function calling is live on the chat completions endpoint, so the model can call your retrieval function and receive the results.
How do I stop the loop from running away?
Set a maximum step count, a per-request token budget and a timeout, and return the best supported answer when the budget is exhausted.
Is agentic RAG always better than one-shot RAG?
No. It improves multi-hop and ambiguous questions but adds latency and cost. For simple lookups, a reranked single retrieval pass is usually enough.
Is there a free plan?
Yes. The free plan includes plugsky-micro and plugsky-lite with 2 API keys and no credit card, and a 14-day full-access trial is available.