RAG

How do you build a local RAG pipeline?

A local RAG pipeline runs three components on hardware you control: an embedding model, a vector store and a generation model, all kept offline if needed. It gives strong privacy and predictable cost, but you own GPU capacity, upgrades and evaluation. Plugsky's private deployment options cover the managed middle ground when you want retrieval without building the stack.

Key facts

Pipeline partsEmbedding model, vector store and generation model on hardware you control
Private deploymentPlugsky supports VPC, on-prem and air-gapped deployments
RAG endpointsPOST /v1/embeddings, POST /v1/rag/collections, POST /v1/rag/query
Retrieval modesKeyword, vector and hybrid search with optional reranking
Models30+ models behind one OpenAI-compatible API
Data handlingPer-collection encryption at rest; API data is not used to train models
Free tierplugsky-micro and plugsky-lite, no card required
Product statusLive

TL;DR

  • Local RAG means embeddings, vectors and generation all run on hardware you control.
  • Privacy and cost predictability are the main wins; GPU capacity and upgrades are the costs.
  • Start with a small open-weight model and a local vector store, then measure quality.
  • Keep the same pipeline shape so you can move between local and managed later.
  • Plugsky offers VPC, on-prem and air-gapped deployment for the model layer.

How it works, step by step

  1. Classify the data: what must never leave your network, and what may use a hosted API.
  2. Pick a local embedding model and a local vector store that fit your hardware.
  3. Chunk and index a representative sample, then build a golden question set.
  4. Run a local generation model and check groundedness and citation quality.
  5. Measure latency and throughput at the concurrency your users will create.
  6. Decide which layer must stay local and which can move to a private or hosted API.
  7. Document the evaluation set so upgrades can be compared against a fixed baseline.
1Classify the data:what must neverleave your network,2Pick a localembedding model anda local vector3Chunk and index arepresentativesample, then build4Run a localgeneration modeland check5Measure latency andthroughput at theconcurrency your6Decide which layermust stay local andwhich can move to a

Try it yourself

Open the local RAG stack generator →

What local RAG means in practice

A RAG pipeline has three replaceable parts: an embedding model that turns text into vectors, a vector store that finds candidates, and a generation model that writes the grounded answer. Local RAG means all three run on hardware you control, typically a workstation, an on-prem server or a private cloud account with no data egress.

Local is not binary. Many teams keep sensitive documents on local embeddings and a local vector store while using a private API for generation, or the reverse. The decision should follow data classification: embeddings leak information about content, so if the corpus is sensitive, the embedding layer is usually the first thing to keep inside the perimeter.

Choosing the components

For embeddings, open-weight families are widely available and run on CPU or GPU depending on model size; multilingual models matter if your corpus is not English. For storage, a local vector database or a Postgres extension keeps vectors on your infrastructure with metadata filtering for access control.

For generation, smaller open-weight models are often enough for extractive, citation-bound answers, while harder synthesis tasks need a larger model and more VRAM. Quantisation reduces memory at some quality cost. The reliable approach is to test candidates against your own golden question set instead of trusting general rankings.

The real trade-offs of going local

The wins are privacy, offline operation and predictable infrastructure cost. The costs are concrete: GPU capacity has to be bought or rented, models must be updated and re-evaluated, throughput is bounded by your hardware, and the team owns monitoring and on-call. Embedding a large corpus locally is also a batch workload that can take hours on modest hardware.

Quality is the subtler trade-off. A frontier API model may answer complex questions better than a model that fits your GPU. If the data allows it, a hybrid setup often wins: keep retrieval and embeddings local, and route only the final synthesis to a private or hosted endpoint.

The hybrid option with Plugsky

Plugsky gives you the same pipeline shape in managed form: POST /v1/embeddings for vectors, POST /v1/rag/collections for ingestion and indexing, and POST /v1/rag/query for ranked chunks with citations. For teams that cannot use a shared cloud, VPC, on-prem and air-gapped deployment options keep the model layer inside your boundary, with 30+ models behind one OpenAI-compatible API.

Use the local RAG stack generator to sketch the architecture, then prototype on the free plan with plugsky-micro and plugsky-lite or the 14-day full-access trial. Current plans are on the live pricing page.

Honest comparison

FactorFully localPlugsky private deploymentPlugsky managed
Data pathStays on your hardwareInside your VPC or on-premPlugsky cloud with region choice
HardwareYou buy or rent GPUsYou provide the environmentNo infrastructure to run
Model rangeWhat fits your hardware30+ models on one API30+ models on one API
OperationsYou own upgrades and on-callShared responsibilityManaged by Plugsky
Best forStrict offline or air-gapped needsResidency with less build effortFastest path to grounded answers

Frequently asked questions

Can RAG run fully offline?

Yes. With local embeddings, a local vector store and a locally hosted model, the entire pipeline can run without network access, which is what air-gapped deployments do.

What hardware does local RAG need?

It depends on model size and quantisation. Embeddings often run on CPU, while generation benefits from a GPU with enough VRAM for the chosen model and context length.

Do local models answer as well as hosted ones?

Smaller local models handle extractive, citation-bound answers well. Complex synthesis and reasoning tasks usually favour larger models, which may not fit your hardware.

Does Plugsky support on-prem deployment?

Yes. Plugsky supports VPC, on-prem and air-gapped deployments for enterprise customers, with the same OpenAI-compatible endpoints.

Is there a free plan?

Yes. The free plan includes plugsky-micro and plugsky-lite with 2 API keys and no credit card, and a 14-day full-access trial is available.

How is pricing structured?

Self-serve plans are flat monthly with unlimited fair-use usage and no per-token charges or overage fees. See the live pricing page for current plans.