Key facts
| Pipeline parts | Embedding model, vector store and generation model on hardware you control |
| Private deployment | Plugsky supports VPC, on-prem and air-gapped deployments |
| RAG endpoints | POST /v1/embeddings, POST /v1/rag/collections, POST /v1/rag/query |
| Retrieval modes | Keyword, vector and hybrid search with optional reranking |
| Models | 30+ models behind one OpenAI-compatible API |
| Data handling | Per-collection encryption at rest; API data is not used to train models |
| Free tier | plugsky-micro and plugsky-lite, no card required |
| Product status | Live |
TL;DR
- Local RAG means embeddings, vectors and generation all run on hardware you control.
- Privacy and cost predictability are the main wins; GPU capacity and upgrades are the costs.
- Start with a small open-weight model and a local vector store, then measure quality.
- Keep the same pipeline shape so you can move between local and managed later.
- Plugsky offers VPC, on-prem and air-gapped deployment for the model layer.
How it works, step by step
- Classify the data: what must never leave your network, and what may use a hosted API.
- Pick a local embedding model and a local vector store that fit your hardware.
- Chunk and index a representative sample, then build a golden question set.
- Run a local generation model and check groundedness and citation quality.
- Measure latency and throughput at the concurrency your users will create.
- Decide which layer must stay local and which can move to a private or hosted API.
- Document the evaluation set so upgrades can be compared against a fixed baseline.
Try it yourself
Open the local RAG stack generator →
What local RAG means in practice
A RAG pipeline has three replaceable parts: an embedding model that turns text into vectors, a vector store that finds candidates, and a generation model that writes the grounded answer. Local RAG means all three run on hardware you control, typically a workstation, an on-prem server or a private cloud account with no data egress.
Local is not binary. Many teams keep sensitive documents on local embeddings and a local vector store while using a private API for generation, or the reverse. The decision should follow data classification: embeddings leak information about content, so if the corpus is sensitive, the embedding layer is usually the first thing to keep inside the perimeter.
Choosing the components
For embeddings, open-weight families are widely available and run on CPU or GPU depending on model size; multilingual models matter if your corpus is not English. For storage, a local vector database or a Postgres extension keeps vectors on your infrastructure with metadata filtering for access control.
For generation, smaller open-weight models are often enough for extractive, citation-bound answers, while harder synthesis tasks need a larger model and more VRAM. Quantisation reduces memory at some quality cost. The reliable approach is to test candidates against your own golden question set instead of trusting general rankings.
The real trade-offs of going local
The wins are privacy, offline operation and predictable infrastructure cost. The costs are concrete: GPU capacity has to be bought or rented, models must be updated and re-evaluated, throughput is bounded by your hardware, and the team owns monitoring and on-call. Embedding a large corpus locally is also a batch workload that can take hours on modest hardware.
Quality is the subtler trade-off. A frontier API model may answer complex questions better than a model that fits your GPU. If the data allows it, a hybrid setup often wins: keep retrieval and embeddings local, and route only the final synthesis to a private or hosted endpoint.
The hybrid option with Plugsky
Plugsky gives you the same pipeline shape in managed form: POST /v1/embeddings for vectors, POST /v1/rag/collections for ingestion and indexing, and POST /v1/rag/query for ranked chunks with citations. For teams that cannot use a shared cloud, VPC, on-prem and air-gapped deployment options keep the model layer inside your boundary, with 30+ models behind one OpenAI-compatible API.
Use the local RAG stack generator to sketch the architecture, then prototype on the free plan with plugsky-micro and plugsky-lite or the 14-day full-access trial. Current plans are on the live pricing page.
Honest comparison
| Factor | Fully local | Plugsky private deployment | Plugsky managed |
|---|---|---|---|
| Data path | Stays on your hardware | Inside your VPC or on-prem | Plugsky cloud with region choice |
| Hardware | You buy or rent GPUs | You provide the environment | No infrastructure to run |
| Model range | What fits your hardware | 30+ models on one API | 30+ models on one API |
| Operations | You own upgrades and on-call | Shared responsibility | Managed by Plugsky |
| Best for | Strict offline or air-gapped needs | Residency with less build effort | Fastest path to grounded answers |
Frequently asked questions
Can RAG run fully offline?
Yes. With local embeddings, a local vector store and a locally hosted model, the entire pipeline can run without network access, which is what air-gapped deployments do.
What hardware does local RAG need?
It depends on model size and quantisation. Embeddings often run on CPU, while generation benefits from a GPU with enough VRAM for the chosen model and context length.
Do local models answer as well as hosted ones?
Smaller local models handle extractive, citation-bound answers well. Complex synthesis and reasoning tasks usually favour larger models, which may not fit your hardware.
Does Plugsky support on-prem deployment?
Yes. Plugsky supports VPC, on-prem and air-gapped deployments for enterprise customers, with the same OpenAI-compatible endpoints.
Is there a free plan?
Yes. The free plan includes plugsky-micro and plugsky-lite with 2 API keys and no credit card, and a 14-day full-access trial is available.
How is pricing structured?
Self-serve plans are flat monthly with unlimited fair-use usage and no per-token charges or overage fees. See the live pricing page for current plans.