RAG

What is the best vector database for RAG?

The best vector database depends on scale, filtering needs, hosting rules and how much operations you want to own. Managed services remove ops work, open-source engines give control, and Postgres extensions keep vectors next to relational data. Plugsky's managed RAG collections cover retrieval without running a vector store at all.

Key facts

Managed optionRAG collections with documents, queries and citations, no vector store to run
Retrieval modesKeyword, vector and hybrid search with optional reranking
Standalone vectorsPOST /v1/embeddings works with any vector database you already run
IngestionDocuments are chunked, embedded and indexed automatically per collection
FormatsPDF, DOCX, TXT, MD and HTML
Data handlingPer-collection encryption at rest; API data is not used to train models
DeploymentManaged, VPC, on-prem and air-gapped options
Product statusLive

TL;DR

  • Choose by filtering needs, scale and hosting rules, not by benchmark charts alone.
  • Managed services trade cost and data path for zero operations work.
  • Open-source engines such as Qdrant, Weaviate and Chroma can run on your own hardware.
  • pgvector keeps vectors inside PostgreSQL, which suits teams already standardised on it.
  • Plugsky's managed collections remove the vector database decision entirely.

How it works, step by step

  1. List the filters queries must apply: tenant, department, date, access level.
  2. Estimate corpus size and growth to understand index and memory needs.
  3. Decide where vectors and metadata are allowed to live: cloud, VPC or on-prem.
  4. Test two candidates with your own embeddings and queries, measuring recall and latency.
  5. Check hybrid search and reranking support before committing to a store.
  6. If retrieval is a means to an answer rather than a product, evaluate a managed RAG service first.
1List the filtersqueries must apply:tenant, department,2Estimate corpussize and growth tounderstand index3Decide wherevectors andmetadata are4Test two candidateswith your ownembeddings and5Check hybrid searchand rerankingsupport before6If retrieval is ameans to an answerrather than a

Try it yourself

Open the vector database comparison →

What a vector database adds beyond an index

At minimum, a vector database stores embeddings and runs an approximate nearest-neighbour search. In production, the harder requirements are metadata filtering, hybrid retrieval, index updates, backup and access control. RAG queries usually combine a semantic match with structured filters such as tenant, department or date, and a store that cannot filter efficiently forces you to over-fetch and discard results in application code.

Operational characteristics decide most real deployments: how the index is rebuilt, how incremental updates behave, how memory scales with vectors, and whether replication and disaster recovery meet your requirements. These are the questions to test before a migration, not after.

The main categories of vector store

Managed services such as Pinecone handle scaling and operations for you; you accept vendor dependency, a data path through their infrastructure and usage-based cost. Open-source engines such as Qdrant, Weaviate and Milvus can run in your own cloud or on-prem, which suits residency requirements but puts upgrades and capacity planning on your team. Embedded options such as Chroma are easy to start with for prototypes and small corpora.

Postgres extensions such as pgvector keep vectors in the database you already operate, with transactional consistency and SQL filtering. The trade-off is that search performance and scale follow PostgreSQL rather than a purpose-built engine.

How to run a fair comparison

Use the same embedded corpus, the same queries and the same filters for every candidate. Measure recall at the k you actually pass to the model, latency at your target concurrency, and the effort to update or delete documents. Include the cost of running the store, not just the licence or subscription.

Also test the ugly cases: deleting a document and confirming it disappears from results, changing an access filter and seeing the result set change, and restoring from backup. A vector store that is fast on a demo corpus but awkward on permissions will cost more later than any latency difference.

The managed alternative: Plugsky RAG collections

If your goal is grounded answers rather than owning retrieval infrastructure, Plugsky manages ingestion, chunking, embedding, indexing and querying behind three endpoints: POST /v1/embeddings, POST /v1/rag/collections and POST /v1/rag/query. Queries support keyword, vector and hybrid modes with optional reranking, and every response includes ranked chunks with source attribution.

Teams that already run pgvector, Qdrant or Pinecone can keep it and use Plugsky only for embeddings. Teams that want fewer moving parts can use collections directly. Compare options with the vector database comparison, then start free with plugsky-micro and plugsky-lite; current plans are on the live pricing page.

Honest comparison

OptionManaged serviceOpen-source enginePostgres extensionPlugsky RAG collections
OperationsVendor-managedYou run and upgrade itYou run PostgresFully managed
HostingVendor cloudYour cloud, VPC or on-premWherever Postgres runsManaged or private deployment
FilteringMetadata filtersPayload filtersSQL and filtersMetadata attached to documents
Hybrid searchVaries by serviceUsually availableRequires extra workKeyword, vector and hybrid built in
Best forTeams avoiding opsResidency and controlPostgres-centric stacksAnswer quality over infrastructure

Frequently asked questions

Do I need a separate vector database for RAG?

Not necessarily. A managed RAG service stores and retrieves chunks for you. A dedicated vector database makes sense when retrieval is a long-term platform concern or you already standardise on one.

Can I keep my current vector database and use Plugsky?

Yes. POST /v1/embeddings works standalone, so you can generate vectors with Plugsky and store them in pgvector, Qdrant, Pinecone or another store.

Which vector database is fastest?

Speed depends on index type, dimensionality, filters, hardware and dataset size. Measure recall and latency on your own corpus rather than relying on general rankings.

Is pgvector enough for production RAG?

For many teams it is, especially when corpora fit comfortably in PostgreSQL and filtering matters more than extreme scale. Very large or write-heavy corpora may need a dedicated engine.

Does Plugsky support hybrid search?

Yes. RAG queries support keyword, vector and hybrid retrieval with optional cross-encoder reranking.

Is there a free plan?

Yes. The free plan includes plugsky-micro and plugsky-lite with 2 API keys and no credit card, and a 14-day full-access trial is available.