Key facts
| Type | Embedded vector database, library or local server |
| Persistence | A persistent client writes collections to local disk |
| Embeddings | Pluggable embedding functions, including local models |
| Metadata | Key-value metadata per chunk enables filtered queries |
| Distance | Cosine, L2 or inner product depending on configuration |
| Retrieval | Query by raw text or by a precomputed vector |
| Generation step | Any local or OpenAI-compatible chat model |
| Endpoint status | Embeddings and chat are live on Plugsky; batch endpoints coming soon |
TL;DR
- Chroma runs as a library or local server, so there is no separate database to operate.
- Use a persistent client to keep collections between restarts.
- Store metadata with each chunk, then filter to avoid cross-source bleed.
- Match the distance metric to your embedding model's training.
- Keep generation separate so retrieval can be tested on its own.
How it works, step by step
- Install Chroma and choose an embedding function, local or hosted.
- Create a persistent client and a named collection.
- Parse documents, chunk them and attach source metadata.
- Add chunks in batches with stable ids so re-ingestion is idempotent.
- Query with top-k results and metadata filters.
- Feed retrieved chunks into a local chat model with citation instructions.
- Evaluate retrieval on known questions and adjust chunking.
Try it yourself
Open the local RAG stack generator →
How Chroma fits a local RAG stack
Chroma occupies the storage and retrieval layer. Your application parses documents, splits them into chunks, and hands those chunks to Chroma together with an embedding function. Chroma computes or accepts vectors, stores them, and returns the closest matches when you query. Everything can run in-process on a laptop, which is why it is a common first choice for private RAG.
It also runs as a local server when more than one process needs access. That keeps the same collection available to a script, a notebook and a small internal app without moving data off the machine.
Collections, embeddings and metadata
A collection is the unit of organisation: one embedding model, one distance metric, one set of documents. Mixing models inside a collection produces meaningless similarities, so create a new collection when you change embedding models, and re-embed the corpus.
- Metadata such as source, section, date and access level enables filters that keep retrieval scoped.
- Stable ids let you upsert rather than duplicate when documents change.
- Distance metric should follow the embedding model's design, typically cosine for text models.
- Batch adds keep ingestion fast and memory reasonable on large corpora.
Retrieve with a small top-k and filters, then rerank or trim before generation.
Tuning retrieval and adding a hosted model
Most quality gains come from chunking and filtering, not from a bigger generator. Test chunk sizes and overlap against a fixed question set, and log which chunk answered each question so failures are visible.
When local generation becomes the bottleneck, keep Chroma as the store and swap only the model call. Plugsky exposes embeddings and chat through an OpenAI-compatible API, so the same client works; embeddings and RAG are live, batch endpoints are coming soon, and vector dimensions must match if you switch embedding models. See pricing for plans and start free with plugsky-micro and plugsky-lite.
Honest comparison
| Stage | Chroma local | Plugsky embeddings and host model | Check before deciding |
|---|---|---|---|
| Deployment | Library or local server | Client-side store with hosted inference | Operations preference |
| Embeddings | Local embedding function | Embeddings endpoint, live | Dimensions and quality |
| Filters | Metadata where clauses | Unchanged, inside your store | Filter complexity |
| Generation | Local chat model | 30+ models over one API | Context length and quality bar |
| Scale | Single-node focused | Inference scales with the service | Corpus growth |
Frequently asked questions
Is Chroma a server or a library?
Both. It runs embedded in your process for scripts and notebooks, and it can run as a local server that multiple clients share.
How do I keep data between restarts?
Use a persistent client pointed at a storage directory instead of an in-memory client. Collections are written to local disk and reload with the client.
Which embedding function should I use?
A local embedding model keeps data on-device; an OpenAI-compatible embeddings endpoint keeps parity with production. Never change embedding models without re-embedding the collection.
How do I avoid mixing documents from different sources?
Attach metadata such as source, date and access level to every chunk, then pass a metadata filter on each query so retrieval stays scoped.
Does Chroma support hybrid search?
Chroma focuses on vector search with metadata filtering. If you need strong keyword matching alongside vectors, pair it with a keyword index or choose a database with built-in hybrid search.
Can I use Chroma in production?
It is widely used for local and small-scale deployments. For multi-tenant scale, replication and heavy concurrency, evaluate a server-oriented vector database.
How do I move Chroma RAG to a hosted model?
Keep Chroma local and point only the generation step at an OpenAI-compatible endpoint. If you also move embeddings, re-embed because vector spaces must match.