Key facts
| Managed option | RAG collections with documents, queries and citations, no vector store to run |
| Retrieval modes | Keyword, vector and hybrid search with optional reranking |
| Standalone vectors | POST /v1/embeddings works with any vector database you already run |
| Ingestion | Documents are chunked, embedded and indexed automatically per collection |
| Formats | PDF, DOCX, TXT, MD and HTML |
| Data handling | Per-collection encryption at rest; API data is not used to train models |
| Deployment | Managed, VPC, on-prem and air-gapped options |
| Product status | Live |
TL;DR
- Choose by filtering needs, scale and hosting rules, not by benchmark charts alone.
- Managed services trade cost and data path for zero operations work.
- Open-source engines such as Qdrant, Weaviate and Chroma can run on your own hardware.
- pgvector keeps vectors inside PostgreSQL, which suits teams already standardised on it.
- Plugsky's managed collections remove the vector database decision entirely.
How it works, step by step
- List the filters queries must apply: tenant, department, date, access level.
- Estimate corpus size and growth to understand index and memory needs.
- Decide where vectors and metadata are allowed to live: cloud, VPC or on-prem.
- Test two candidates with your own embeddings and queries, measuring recall and latency.
- Check hybrid search and reranking support before committing to a store.
- If retrieval is a means to an answer rather than a product, evaluate a managed RAG service first.
Try it yourself
Open the vector database comparison →
What a vector database adds beyond an index
At minimum, a vector database stores embeddings and runs an approximate nearest-neighbour search. In production, the harder requirements are metadata filtering, hybrid retrieval, index updates, backup and access control. RAG queries usually combine a semantic match with structured filters such as tenant, department or date, and a store that cannot filter efficiently forces you to over-fetch and discard results in application code.
Operational characteristics decide most real deployments: how the index is rebuilt, how incremental updates behave, how memory scales with vectors, and whether replication and disaster recovery meet your requirements. These are the questions to test before a migration, not after.
The main categories of vector store
Managed services such as Pinecone handle scaling and operations for you; you accept vendor dependency, a data path through their infrastructure and usage-based cost. Open-source engines such as Qdrant, Weaviate and Milvus can run in your own cloud or on-prem, which suits residency requirements but puts upgrades and capacity planning on your team. Embedded options such as Chroma are easy to start with for prototypes and small corpora.
Postgres extensions such as pgvector keep vectors in the database you already operate, with transactional consistency and SQL filtering. The trade-off is that search performance and scale follow PostgreSQL rather than a purpose-built engine.
How to run a fair comparison
Use the same embedded corpus, the same queries and the same filters for every candidate. Measure recall at the k you actually pass to the model, latency at your target concurrency, and the effort to update or delete documents. Include the cost of running the store, not just the licence or subscription.
Also test the ugly cases: deleting a document and confirming it disappears from results, changing an access filter and seeing the result set change, and restoring from backup. A vector store that is fast on a demo corpus but awkward on permissions will cost more later than any latency difference.
The managed alternative: Plugsky RAG collections
If your goal is grounded answers rather than owning retrieval infrastructure, Plugsky manages ingestion, chunking, embedding, indexing and querying behind three endpoints: POST /v1/embeddings, POST /v1/rag/collections and POST /v1/rag/query. Queries support keyword, vector and hybrid modes with optional reranking, and every response includes ranked chunks with source attribution.
Teams that already run pgvector, Qdrant or Pinecone can keep it and use Plugsky only for embeddings. Teams that want fewer moving parts can use collections directly. Compare options with the vector database comparison, then start free with plugsky-micro and plugsky-lite; current plans are on the live pricing page.
Honest comparison
| Option | Managed service | Open-source engine | Postgres extension | Plugsky RAG collections |
|---|---|---|---|---|
| Operations | Vendor-managed | You run and upgrade it | You run Postgres | Fully managed |
| Hosting | Vendor cloud | Your cloud, VPC or on-prem | Wherever Postgres runs | Managed or private deployment |
| Filtering | Metadata filters | Payload filters | SQL and filters | Metadata attached to documents |
| Hybrid search | Varies by service | Usually available | Requires extra work | Keyword, vector and hybrid built in |
| Best for | Teams avoiding ops | Residency and control | Postgres-centric stacks | Answer quality over infrastructure |
Frequently asked questions
Do I need a separate vector database for RAG?
Not necessarily. A managed RAG service stores and retrieves chunks for you. A dedicated vector database makes sense when retrieval is a long-term platform concern or you already standardise on one.
Can I keep my current vector database and use Plugsky?
Yes. POST /v1/embeddings works standalone, so you can generate vectors with Plugsky and store them in pgvector, Qdrant, Pinecone or another store.
Which vector database is fastest?
Speed depends on index type, dimensionality, filters, hardware and dataset size. Measure recall and latency on your own corpus rather than relying on general rankings.
Is pgvector enough for production RAG?
For many teams it is, especially when corpora fit comfortably in PostgreSQL and filtering matters more than extreme scale. Very large or write-heavy corpora may need a dedicated engine.
Does Plugsky support hybrid search?
Yes. RAG queries support keyword, vector and hybrid retrieval with optional cross-encoder reranking.
Is there a free plan?
Yes. The free plan includes plugsky-micro and plugsky-lite with 2 API keys and no credit card, and a 14-day full-access trial is available.