Local AI

How do you build local RAG with Chroma?

Chroma is an embedded vector database you run as a library or local server. Create a persistent client, define a collection with your embedding function, add chunks with metadata, then query by text or vector and filter on metadata. Paired with a local chat model, Chroma keeps the entire retrieval loop on your machine.

Key facts

TypeEmbedded vector database, library or local server
PersistenceA persistent client writes collections to local disk
EmbeddingsPluggable embedding functions, including local models
MetadataKey-value metadata per chunk enables filtered queries
DistanceCosine, L2 or inner product depending on configuration
RetrievalQuery by raw text or by a precomputed vector
Generation stepAny local or OpenAI-compatible chat model
Endpoint statusEmbeddings and chat are live on Plugsky; batch endpoints coming soon

TL;DR

  • Chroma runs as a library or local server, so there is no separate database to operate.
  • Use a persistent client to keep collections between restarts.
  • Store metadata with each chunk, then filter to avoid cross-source bleed.
  • Match the distance metric to your embedding model's training.
  • Keep generation separate so retrieval can be tested on its own.

How it works, step by step

  1. Install Chroma and choose an embedding function, local or hosted.
  2. Create a persistent client and a named collection.
  3. Parse documents, chunk them and attach source metadata.
  4. Add chunks in batches with stable ids so re-ingestion is idempotent.
  5. Query with top-k results and metadata filters.
  6. Feed retrieved chunks into a local chat model with citation instructions.
  7. Evaluate retrieval on known questions and adjust chunking.
1Install Chroma andchoose an embeddingfunction, local or2Create a persistentclient and a namedcollection.3Parse documents,chunk them andattach source4Add chunks inbatches with stableids so re-ingestion5Query with top-kresults andmetadata filters.6Feed retrievedchunks into a localchat model with

Try it yourself

Open the local RAG stack generator →

How Chroma fits a local RAG stack

Chroma occupies the storage and retrieval layer. Your application parses documents, splits them into chunks, and hands those chunks to Chroma together with an embedding function. Chroma computes or accepts vectors, stores them, and returns the closest matches when you query. Everything can run in-process on a laptop, which is why it is a common first choice for private RAG.

It also runs as a local server when more than one process needs access. That keeps the same collection available to a script, a notebook and a small internal app without moving data off the machine.

Collections, embeddings and metadata

A collection is the unit of organisation: one embedding model, one distance metric, one set of documents. Mixing models inside a collection produces meaningless similarities, so create a new collection when you change embedding models, and re-embed the corpus.

  • Metadata such as source, section, date and access level enables filters that keep retrieval scoped.
  • Stable ids let you upsert rather than duplicate when documents change.
  • Distance metric should follow the embedding model's design, typically cosine for text models.
  • Batch adds keep ingestion fast and memory reasonable on large corpora.

Retrieve with a small top-k and filters, then rerank or trim before generation.

Tuning retrieval and adding a hosted model

Most quality gains come from chunking and filtering, not from a bigger generator. Test chunk sizes and overlap against a fixed question set, and log which chunk answered each question so failures are visible.

When local generation becomes the bottleneck, keep Chroma as the store and swap only the model call. Plugsky exposes embeddings and chat through an OpenAI-compatible API, so the same client works; embeddings and RAG are live, batch endpoints are coming soon, and vector dimensions must match if you switch embedding models. See pricing for plans and start free with plugsky-micro and plugsky-lite.

Honest comparison

StageChroma localPlugsky embeddings and host modelCheck before deciding
DeploymentLibrary or local serverClient-side store with hosted inferenceOperations preference
EmbeddingsLocal embedding functionEmbeddings endpoint, liveDimensions and quality
FiltersMetadata where clausesUnchanged, inside your storeFilter complexity
GenerationLocal chat model30+ models over one APIContext length and quality bar
ScaleSingle-node focusedInference scales with the serviceCorpus growth

Frequently asked questions

Is Chroma a server or a library?

Both. It runs embedded in your process for scripts and notebooks, and it can run as a local server that multiple clients share.

How do I keep data between restarts?

Use a persistent client pointed at a storage directory instead of an in-memory client. Collections are written to local disk and reload with the client.

Which embedding function should I use?

A local embedding model keeps data on-device; an OpenAI-compatible embeddings endpoint keeps parity with production. Never change embedding models without re-embedding the collection.

How do I avoid mixing documents from different sources?

Attach metadata such as source, date and access level to every chunk, then pass a metadata filter on each query so retrieval stays scoped.

Does Chroma support hybrid search?

Chroma focuses on vector search with metadata filtering. If you need strong keyword matching alongside vectors, pair it with a keyword index or choose a database with built-in hybrid search.

Can I use Chroma in production?

It is widely used for local and small-scale deployments. For multi-tenant scale, replication and heavy concurrency, evaluate a server-oriented vector database.

How do I move Chroma RAG to a hosted model?

Keep Chroma local and point only the generation step at an OpenAI-compatible endpoint. If you also move embeddings, re-embed because vector spaces must match.