RAG

What is agentic RAG and when should you use it?

Agentic RAG adds a reasoning loop to classical retrieval: the model plans, rewrites queries, calls tools and retrieves again until it has enough evidence. It helps multi-hop and ambiguous questions but increases latency and token use. Plugsky supports it with OpenAI-compatible chat, live function calling, embeddings and RAG collections.

Key facts

RAG endpointsPOST /v1/embeddings, POST /v1/rag/collections, POST /v1/rag/query
Retrieval modesKeyword, vector and hybrid search with optional cross-encoder reranking
Agent building blocksChat, streaming, JSON mode and function calling are live
OrchestrationOpenAI-compatible tool calls; works with LangChain, LlamaIndex and Haystack
CitationsEvery query returns ranked chunks with source attribution
Models30+ models behind one API, from free tiers to frontier reasoning
DeploymentHosted, VPC, on-prem and air-gapped options
Product statusLive

TL;DR

  • Agentic RAG replaces one retrieval pass with a bounded plan, retrieve, critique loop.
  • Use it for multi-hop questions, ambiguous phrasing and tasks mixing documents with tools.
  • Bound the loop with a maximum step count, a token budget and request timeouts.
  • Keep citations on every retrieval so each hop stays auditable.
  • Plugsky exposes chat, function calling, embeddings and RAG on one OpenAI-compatible API.

How it works, step by step

  1. Start with one-shot retrieval and record where it fails on real questions.
  2. Define the toolset the model may call: collection query, filters and any external API.
  3. Set a maximum number of retrieval steps and a per-request token budget.
  4. Rewrite or decompose the query before each retrieval and log the rewritten text.
  5. Require the model to cite chunk IDs in the final answer.
  6. Evaluate the loop against a fixed question set before enabling it in production.
  7. Route planning steps to a fast model and the final synthesis to a stronger one.
1Start with one-shotretrieval andrecord where it2Define the toolsetthe model may call:collection query,3Set a maximumnumber of retrievalsteps and a4Rewrite ordecompose the querybefore each5Require the modelto cite chunk IDsin the final6Evaluate the loopagainst a fixedquestion set before

Try it yourself

Open the RAG architecture builder →

What agentic RAG adds to a retrieval pipeline

Classical RAG runs one retrieval pass: embed the question, fetch top-k chunks, generate an answer. Agentic RAG wraps that call in a reasoning loop. The model plans the question, may rewrite or decompose it, calls the retrieval tool, inspects the returned chunks, and decides whether one more pass is needed. This matters for multi-hop questions such as comparing two policies, for ambiguous queries that need clarification through search, and for tasks that combine a document corpus with a live system call.

The cost is straightforward: every extra step is another model call and another retrieval, so latency and token use grow with loop depth. The design goal is not to let the model loop freely but to give it a small, bounded toolbox and a clear stop condition.

The loop: plan, retrieve, critique, answer

A practical agentic loop has four phases. Plan: turn the user request into a search plan, including filters and expected sources. Retrieve: call the collection query endpoint with top_k and optional reranking. Critique: check whether the returned chunks actually contain the missing facts, and rewrite the query if they do not. Answer: compose the response from cited chunks only, and refuse when the corpus does not support an answer.

Each phase is observable if you log the tool call, the query text and the chunk IDs returned. That log is what makes an agentic pipeline debuggable and auditable later, and it is also the raw material for evaluation sets.

When the loop is worth the latency

Use agentic RAG when single-pass recall is the bottleneck: questions that need two or more facts joined, corpora split across collections, or workflows where the model must choose between search and another tool. Stay with one-shot retrieval for high-volume support answers, latency-sensitive chat, and simple lookups where a reranked top-10 already contains the answer.

A useful middle ground is a single rewrite step: retrieve once, ask the model whether the evidence is sufficient, and allow at most one corrected retrieval. Most of the quality gain arrives with the second attempt, while latency stays predictable and easy to explain to stakeholders.

Building agentic RAG on Plugsky

Plugsky provides the pieces on one OpenAI-compatible API: chat completions for reasoning, live function calling for tool use, POST /v1/embeddings for vectors, and collections plus queries for retrieval with citations. Existing LangChain, LlamaIndex or Haystack agents can point at the same base URL and keep their code. Route cheap planning steps to a fast model and the final answer to a stronger one across 30+ models. Start on the free plan with plugsky-micro and plugsky-lite, or use the 14-day full-access trial; current plans are on the live pricing page.

Honest comparison

CapabilityOne-shot RAGAgentic RAGPlugsky support
Retrieval passesOneOne or many, decided by the modelManual loop over /v1/rag/query
Query handlingOriginal query onlyRewritten and decomposedAny model on the chat endpoint
Tool useRetrieval onlyRetrieval plus external toolsFunction calling is live
Latency profileLower and predictableHigher, grows with stepsRoute steps across model tiers
Best forSimple lookupsMulti-hop and ambiguous tasksBoth patterns on one API

Frequently asked questions

What is agentic RAG in one sentence?

Agentic RAG is a retrieval pipeline where a model plans, searches, inspects results and decides whether to search again before answering.

Do I need a framework to build it?

No. Any loop that calls an OpenAI-compatible chat endpoint with tools and a query endpoint works. Frameworks such as LangChain, LlamaIndex and Haystack add orchestration conveniences.

Does Plugsky support tool calls?

Yes. Function calling is live on the chat completions endpoint, so the model can call your retrieval function and receive the results.

How do I stop the loop from running away?

Set a maximum step count, a per-request token budget and a timeout, and return the best supported answer when the budget is exhausted.

Is agentic RAG always better than one-shot RAG?

No. It improves multi-hop and ambiguous questions but adds latency and cost. For simple lookups, a reranked single retrieval pass is usually enough.

Is there a free plan?

Yes. The free plan includes plugsky-micro and plugsky-lite with 2 API keys and no credit card, and a 14-day full-access trial is available.