Local AI

How do you build local RAG with pgvector?

pgvector adds a vector column type and similarity operators to PostgreSQL. You store embeddings alongside the rows they belong to, index them with HNSW or IVFFlat, and query nearest neighbours with SQL while filtering on ordinary columns. Retrieval stays in the database you already operate, with no separate vector service.

Key facts

TypePostgreSQL extension adding vector columns and indexes
Index typesHNSW for speed and recall, IVFFlat for build trade-offs
OperatorsDistance operators for L2, cosine and inner product
Dimension limitsVector indexes have dimension ceilings; half-precision extends them
FilteringCombine vector ordering with ordinary SQL WHERE clauses
TransactionsEmbeddings live in the same transactional store as your data
GenerationAny local or OpenAI-compatible chat model
Endpoint statusEmbeddings and chat are live; batch endpoints are coming soon

TL;DR

  • pgvector keeps embeddings inside PostgreSQL, so there is no new service to run.
  • HNSW favours query speed and recall; IVFFlat favours build and memory trade-offs.
  • Filter with regular SQL to scope retrieval to a tenant or date range.
  • Mind index dimension limits and choose half-precision when needed.
  • Generation stays a separate step, local or hosted.

How it works, step by step

  1. Install the pgvector extension in your Postgres instance and enable it.
  2. Create a table with a vector column sized to your embedding model.
  3. Embed chunks and insert them with their source metadata.
  4. Build an HNSW or IVFFlat index after the initial data load.
  5. Query nearest neighbours with the matching distance operator and filters.
  6. Pass the retrieved rows to a chat model for grounded generation.
  7. Measure recall and latency as the corpus grows, then retune the index.
1Install thepgvector extensionin your Postgres2Create a table witha vector columnsized to your3Embed chunks andinsert them withtheir source4Build an HNSW orIVFFlat index afterthe initial data5Query nearestneighbours with thematching distance6Pass the retrievedrows to a chatmodel for grounded

Try it yourself

Open the vector database comparison →

Why Postgres is enough for many RAG systems

Most internal RAG systems do not need a specialised vector service. Documents already live in a relational database, access rules are expressed as tables and roles, and backups, migrations and monitoring already exist. pgvector adds similarity search to that environment instead of beside it.

The practical benefit is transactionality and filtering. You can insert a document, its chunks and its embeddings in one transaction, then query nearest neighbours while filtering by tenant, status or date using the same SQL you already trust. Fewer moving parts means fewer drift problems.

Indexes, operators and filtering

Two index families dominate. HNSW builds a graph and gives strong recall with fast queries, at the cost of build time and memory. IVFFlat clusters vectors into lists and needs tuning of list count and probes, but builds faster and uses less memory. Start with HNSW unless memory is the constraint.

  • Match the distance operator to your embedding model: cosine for most text models, L2 or inner product where appropriate.
  • Keep the declared dimension exactly equal to the embedding size, and re-embed if you change models.
  • Use half-precision vectors when you need higher indexable dimensions or smaller storage.
  • Filter in SQL before or alongside the vector order-by, and index those columns.

Operating pgvector at scale

Vector search degrades quietly. Recall drops as the corpus grows if index parameters are left at defaults, and latency rises with unfiltered queries over large tables. Build a fixed query set with known answers, measure recall and latency on a schedule, and retune parameters when they move.

Vacuum and analyse regularly, because embeddings are written in bulk during ingestion. For generation, keep the model call separate: pgvector stores and searches, it does not embed or generate. Plugsky provides OpenAI-compatible embeddings and chat, both live, so a Postgres-based stack can route generation to 30+ models without changing the database. Batch endpoints are coming soon. See pricing for plans.

Honest comparison

ConcernpgvectorPlugsky embeddings and host modelCheck before deciding
StorageInside your PostgreSQL databaseYour database, unchangedExisting Postgres usage
EmbeddingsComputed locally or via APIEmbeddings endpoint, liveDimensions and parity
IndexingHNSW or IVFFlat in PostgresNot applicableCorpus size and latency
FilteringFull SQL plus vector orderingUnchangedFilter complexity and tenancy
GenerationLocal or hosted chat model30+ models on one APIQuality bar

Frequently asked questions

Do I need a separate vector database if I use Postgres?

Often no. pgvector handles large vector sets with the right index, and keeping vectors beside relational data simplifies filtering, transactions and backups.

HNSW or IVFFlat?

HNSW gives better recall and query speed at higher build time and memory cost. IVFFlat builds faster and uses less memory but needs tuning of list counts and probes.

What are pgvector's dimension limits?

The vector type accepts high dimensions, but indexed search has a ceiling. Half-precision vectors extend that ceiling, so check the current limits for your embedding model before committing.

How do I filter by tenant?

Add a tenant column, include it in the WHERE clause next to the vector order-by, and index it. This keeps retrieval scoped without a collection per tenant.

Can I generate embeddings inside Postgres?

The extension stores and searches vectors; it does not create them. Compute embeddings in your application or through an API, then insert them.

How do I keep the index fast as data grows?

Rebuild or retune index parameters as volume increases, vacuum regularly, and monitor recall against a fixed query set rather than assuming defaults stay optimal.

Can I use pgvector with a hosted model?

Yes. Keep the database where it is and route generation, or embeddings, to an OpenAI-compatible endpoint. Keep dimensions consistent if you switch providers.