Local AI

How do you build local RAG with Qdrant?

Qdrant is a vector database you can run locally in a container or as a single binary. Create a collection with the right vector size and distance, upsert points with payloads, then search with filters on those payloads. It adds payload indexing, quantization and snapshotting, which makes it a strong step up from a toy store.

Key facts

DeploymentSingle binary or container, run locally or in your VPC
Data modelCollections hold points: a vector plus a JSON payload
FilteringPayload filters with indexed fields keep retrieval scoped
QuantizationScalar and binary quantization reduce vector memory
HybridSparse and dense vectors can be combined in one collection
SnapshotsSnapshot and restore support backup and migration
GenerationAny local or OpenAI-compatible chat model
Endpoint statusEmbeddings and chat are live; batch endpoints are coming soon

TL;DR

  • Qdrant runs locally but behaves like a production vector service.
  • Use payload indexes so filtered search stays fast.
  • Quantization cuts vector memory when the corpus grows.
  • Snapshots make backup and migration straightforward.
  • Combine sparse and dense vectors for hybrid retrieval.

How it works, step by step

  1. Run Qdrant locally with a container or a single binary.
  2. Create a collection with the vector size and distance your model expects.
  3. Upsert points with vectors and payload metadata, using stable ids.
  4. Create payload indexes on the fields you filter by.
  5. Search with top-k and filters, then rerank or trim the results.
  6. Generate answers with a local chat model using only retrieved context.
  7. Snapshot the collection on a schedule and test a restore.
1Run Qdrant locallywith a container ora single binary.2Create a collectionwith the vectorsize and distance3Upsert points withvectors and payloadmetadata, using4Create payloadindexes on thefields you filter5Search with top-kand filters, thenrerank or trim the6Generate answerswith a local chatmodel using only

Try it yourself

Open the local RAG stack generator →

Why use Qdrant for local RAG

Chroma and FAISS are excellent starting points, but Qdrant adds the features that RAG systems eventually need: indexed metadata filters, quantization to control memory, sparse vectors for keyword-aware retrieval, and snapshots for backup. It runs comfortably on one machine, so you can develop locally and keep the same data model when you move to a server.

That continuity matters. Migrating between vector stores is rarely a copy-and-paste job, because filters, ids and distance metrics behave differently. Choosing a store that already behaves like a service avoids one migration later.

Payloads, filters and quantization

Every point carries a vector and a payload. Put source, title, section, date and access level in the payload, then create indexes for the fields you filter on. Filtered vector search stays fast when the filter fields are indexed; without indexes it degrades as the collection grows.

  • Distance metric must match your embedding model, typically cosine for text.
  • Stable ids let you update chunks in place when documents change.
  • Quantization reduces vector memory, with a recall trade-off to validate on a fixed query set.
  • Sparse plus dense vectors give hybrid retrieval for names, codes and exact terms.

Backups and scaling out

Treat a local vector store as data, because it is. Schedule snapshots, test a restore, and record which embedding model and dimensions produced the collection. If you change embedding models, re-embed and rebuild rather than mixing vector spaces.

When concurrent use grows, the same collection can move to a server deployment inside your network while the application stays unchanged. For the generation step, Plugsky serves chat and embeddings over an OpenAI-compatible API, both live, with batch endpoints coming soon, so a Qdrant-based stack can route inference to 30+ models without moving data. See pricing for plans.

Honest comparison

ConcernQdrant localPlugsky embeddings and host modelCheck before deciding
DeploymentContainer or binary on your machineHosted inference, your storeOperations and isolation
FiltersIndexed payload filtersUnchangedFilter fields and tenancy
MemoryQuantization reduces vector footprintNot applicableCorpus size
RetrievalDense and sparse in one collectionUnchangedHybrid retrieval needs
GenerationLocal chat model30+ models over one APIQuality bar

Frequently asked questions

Can Qdrant run without the cloud?

Yes. It ships as a binary and a container image, so it runs entirely on your machine or inside your network with no external dependency.

What is a payload in Qdrant?

A JSON object attached to a point. It typically holds source, title, date and access fields, and it can be filtered and indexed for fast scoped search.

Should I quantize vectors locally?

Quantization reduces memory at a small recall cost. Test it when the corpus is large relative to available RAM, and validate against a fixed query set before enabling it in production.

Does Qdrant support hybrid search?

Yes. Sparse and dense vectors can live in the same collection and be combined, which helps with names, codes and exact terms that embeddings blur.

How do I back up a local Qdrant instance?

Use built-in snapshots on a schedule and test a restore, because a backup you have never restored is a guess. Snapshots also help migrate between machines.

Is Qdrant overkill for a small project?

For a few thousand chunks, Chroma or pgvector may be simpler. Qdrant makes sense when you need payload filtering, quantization or a clearer path to production.

How do I connect Qdrant to a hosted model?

Keep Qdrant as the store and point embeddings and generation at an OpenAI-compatible API. Keep dimensions and vector spaces consistent if you change embedding models.