Key facts
| Deployment | Single binary or container, run locally or in your VPC |
| Data model | Collections hold points: a vector plus a JSON payload |
| Filtering | Payload filters with indexed fields keep retrieval scoped |
| Quantization | Scalar and binary quantization reduce vector memory |
| Hybrid | Sparse and dense vectors can be combined in one collection |
| Snapshots | Snapshot and restore support backup and migration |
| Generation | Any local or OpenAI-compatible chat model |
| Endpoint status | Embeddings and chat are live; batch endpoints are coming soon |
TL;DR
- Qdrant runs locally but behaves like a production vector service.
- Use payload indexes so filtered search stays fast.
- Quantization cuts vector memory when the corpus grows.
- Snapshots make backup and migration straightforward.
- Combine sparse and dense vectors for hybrid retrieval.
How it works, step by step
- Run Qdrant locally with a container or a single binary.
- Create a collection with the vector size and distance your model expects.
- Upsert points with vectors and payload metadata, using stable ids.
- Create payload indexes on the fields you filter by.
- Search with top-k and filters, then rerank or trim the results.
- Generate answers with a local chat model using only retrieved context.
- Snapshot the collection on a schedule and test a restore.
Try it yourself
Open the local RAG stack generator →
Why use Qdrant for local RAG
Chroma and FAISS are excellent starting points, but Qdrant adds the features that RAG systems eventually need: indexed metadata filters, quantization to control memory, sparse vectors for keyword-aware retrieval, and snapshots for backup. It runs comfortably on one machine, so you can develop locally and keep the same data model when you move to a server.
That continuity matters. Migrating between vector stores is rarely a copy-and-paste job, because filters, ids and distance metrics behave differently. Choosing a store that already behaves like a service avoids one migration later.
Payloads, filters and quantization
Every point carries a vector and a payload. Put source, title, section, date and access level in the payload, then create indexes for the fields you filter on. Filtered vector search stays fast when the filter fields are indexed; without indexes it degrades as the collection grows.
- Distance metric must match your embedding model, typically cosine for text.
- Stable ids let you update chunks in place when documents change.
- Quantization reduces vector memory, with a recall trade-off to validate on a fixed query set.
- Sparse plus dense vectors give hybrid retrieval for names, codes and exact terms.
Backups and scaling out
Treat a local vector store as data, because it is. Schedule snapshots, test a restore, and record which embedding model and dimensions produced the collection. If you change embedding models, re-embed and rebuild rather than mixing vector spaces.
When concurrent use grows, the same collection can move to a server deployment inside your network while the application stays unchanged. For the generation step, Plugsky serves chat and embeddings over an OpenAI-compatible API, both live, with batch endpoints coming soon, so a Qdrant-based stack can route inference to 30+ models without moving data. See pricing for plans.
Honest comparison
| Concern | Qdrant local | Plugsky embeddings and host model | Check before deciding |
|---|---|---|---|
| Deployment | Container or binary on your machine | Hosted inference, your store | Operations and isolation |
| Filters | Indexed payload filters | Unchanged | Filter fields and tenancy |
| Memory | Quantization reduces vector footprint | Not applicable | Corpus size |
| Retrieval | Dense and sparse in one collection | Unchanged | Hybrid retrieval needs |
| Generation | Local chat model | 30+ models over one API | Quality bar |
Frequently asked questions
Can Qdrant run without the cloud?
Yes. It ships as a binary and a container image, so it runs entirely on your machine or inside your network with no external dependency.
What is a payload in Qdrant?
A JSON object attached to a point. It typically holds source, title, date and access fields, and it can be filtered and indexed for fast scoped search.
Should I quantize vectors locally?
Quantization reduces memory at a small recall cost. Test it when the corpus is large relative to available RAM, and validate against a fixed query set before enabling it in production.
Does Qdrant support hybrid search?
Yes. Sparse and dense vectors can live in the same collection and be combined, which helps with names, codes and exact terms that embeddings blur.
How do I back up a local Qdrant instance?
Use built-in snapshots on a schedule and test a restore, because a backup you have never restored is a guess. Snapshots also help migrate between machines.
Is Qdrant overkill for a small project?
For a few thousand chunks, Chroma or pgvector may be simpler. Qdrant makes sense when you need payload filtering, quantization or a clearer path to production.
How do I connect Qdrant to a hosted model?
Keep Qdrant as the store and point embeddings and generation at an OpenAI-compatible API. Keep dimensions and vector spaces consistent if you change embedding models.