Key facts
| Type | PostgreSQL extension adding vector columns and indexes |
| Index types | HNSW for speed and recall, IVFFlat for build trade-offs |
| Operators | Distance operators for L2, cosine and inner product |
| Dimension limits | Vector indexes have dimension ceilings; half-precision extends them |
| Filtering | Combine vector ordering with ordinary SQL WHERE clauses |
| Transactions | Embeddings live in the same transactional store as your data |
| Generation | Any local or OpenAI-compatible chat model |
| Endpoint status | Embeddings and chat are live; batch endpoints are coming soon |
TL;DR
- pgvector keeps embeddings inside PostgreSQL, so there is no new service to run.
- HNSW favours query speed and recall; IVFFlat favours build and memory trade-offs.
- Filter with regular SQL to scope retrieval to a tenant or date range.
- Mind index dimension limits and choose half-precision when needed.
- Generation stays a separate step, local or hosted.
How it works, step by step
- Install the pgvector extension in your Postgres instance and enable it.
- Create a table with a vector column sized to your embedding model.
- Embed chunks and insert them with their source metadata.
- Build an HNSW or IVFFlat index after the initial data load.
- Query nearest neighbours with the matching distance operator and filters.
- Pass the retrieved rows to a chat model for grounded generation.
- Measure recall and latency as the corpus grows, then retune the index.
Try it yourself
Open the vector database comparison →
Why Postgres is enough for many RAG systems
Most internal RAG systems do not need a specialised vector service. Documents already live in a relational database, access rules are expressed as tables and roles, and backups, migrations and monitoring already exist. pgvector adds similarity search to that environment instead of beside it.
The practical benefit is transactionality and filtering. You can insert a document, its chunks and its embeddings in one transaction, then query nearest neighbours while filtering by tenant, status or date using the same SQL you already trust. Fewer moving parts means fewer drift problems.
Indexes, operators and filtering
Two index families dominate. HNSW builds a graph and gives strong recall with fast queries, at the cost of build time and memory. IVFFlat clusters vectors into lists and needs tuning of list count and probes, but builds faster and uses less memory. Start with HNSW unless memory is the constraint.
- Match the distance operator to your embedding model: cosine for most text models, L2 or inner product where appropriate.
- Keep the declared dimension exactly equal to the embedding size, and re-embed if you change models.
- Use half-precision vectors when you need higher indexable dimensions or smaller storage.
- Filter in SQL before or alongside the vector order-by, and index those columns.
Operating pgvector at scale
Vector search degrades quietly. Recall drops as the corpus grows if index parameters are left at defaults, and latency rises with unfiltered queries over large tables. Build a fixed query set with known answers, measure recall and latency on a schedule, and retune parameters when they move.
Vacuum and analyse regularly, because embeddings are written in bulk during ingestion. For generation, keep the model call separate: pgvector stores and searches, it does not embed or generate. Plugsky provides OpenAI-compatible embeddings and chat, both live, so a Postgres-based stack can route generation to 30+ models without changing the database. Batch endpoints are coming soon. See pricing for plans.
Honest comparison
| Concern | pgvector | Plugsky embeddings and host model | Check before deciding |
|---|---|---|---|
| Storage | Inside your PostgreSQL database | Your database, unchanged | Existing Postgres usage |
| Embeddings | Computed locally or via API | Embeddings endpoint, live | Dimensions and parity |
| Indexing | HNSW or IVFFlat in Postgres | Not applicable | Corpus size and latency |
| Filtering | Full SQL plus vector ordering | Unchanged | Filter complexity and tenancy |
| Generation | Local or hosted chat model | 30+ models on one API | Quality bar |
Frequently asked questions
Do I need a separate vector database if I use Postgres?
Often no. pgvector handles large vector sets with the right index, and keeping vectors beside relational data simplifies filtering, transactions and backups.
HNSW or IVFFlat?
HNSW gives better recall and query speed at higher build time and memory cost. IVFFlat builds faster and uses less memory but needs tuning of list counts and probes.
What are pgvector's dimension limits?
The vector type accepts high dimensions, but indexed search has a ceiling. Half-precision vectors extend that ceiling, so check the current limits for your embedding model before committing.
How do I filter by tenant?
Add a tenant column, include it in the WHERE clause next to the vector order-by, and index it. This keeps retrieval scoped without a collection per tenant.
Can I generate embeddings inside Postgres?
The extension stores and searches vectors; it does not create them. Compute embeddings in your application or through an API, then insert them.
How do I keep the index fast as data grows?
Rebuild or retune index parameters as volume increases, vacuum regularly, and monitor recall against a fixed query set rather than assuming defaults stay optimal.
Can I use pgvector with a hosted model?
Yes. Keep the database where it is and route generation, or embeddings, to an OpenAI-compatible endpoint. Keep dimensions consistent if you switch providers.