pgvector Explained: HNSW, IVFFlat, and Hybrid Search
Teams building search over application data often want semantic retrieval without operating a separate vector service. pgvector adds vector storage and similarity search to PostgreSQL, letting developers keep embeddings beside relational records, permissions, and ordinary text indexes. This explainer covers its index choices, filtering behavior, and a practical pattern for combining vector retrieval with PostgreSQL full-text search.
What Is pgvector?
pgvector is an open-source PostgreSQL extension for storing vectors and querying them by distance. A vector is a list of numeric values, commonly an embedding produced by a machine-learning model. The extension adds a vector data type, distance operators, and index access methods for approximate nearest-neighbor search.
Unlike a standalone vector service, pgvector works inside PostgreSQL. Applications can use SQL joins, transactions, constraints, and row-level access controls alongside vector queries. The pgvector project documentation describes supported types, distance operators, and its HNSW and IVFFlat indexes. The extension does not create embeddings or generate answers: an application still needs an embedding model and, for retrieval-augmented generation (RAG), a separate language model.
The Problem pgvector Solves
An application may already store users, documents, and access metadata in PostgreSQL, while semantic search appears to require another database and a synchronization pipeline. Copies of the same records can become inconsistent, and a search result must still be checked against the user’s permissions and the current version of its source.
pgvector provides a way to keep vectors and related application data in one database. That can reduce operational components and make transactional updates straightforward. It is not automatically the right choice for every workload: very large vector collections, specialized filtering needs, or a search service’s operational features may justify a separate system. The broader vector database explainer covers the retrieval pipeline and how to choose a storage approach.
How pgvector Search Works
An exact nearest-neighbor query compares the query vector with every eligible row and orders the results by a distance operator. It is simple and produces exact results, but the work grows with the number of vectors. Approximate indexes examine a smaller part of the collection to reduce latency, trading some recall for speed.
pgvector offers two commonly used approximate index types:
| Index | Structure and query behavior | Strengths | Costs and considerations |
|---|---|---|---|
| HNSW | A multilayer graph navigated to find nearby vectors | Strong speed-recall trade-off; does not require training data before index creation | Slower to build and uses more memory; index and search settings affect recall |
| IVFFlat | Clusters vectors into lists and probes selected lists for candidates | Faster to build and uses less memory in many workloads | Needs representative data when built; too few probes can miss relevant neighbors |
| Exact scan | Computes distance for every row that passes the SQL filters | Exact results and no approximate-index tuning | Query work grows with the eligible collection size |
The query operator and index operator class must match. For cosine distance, <=> pairs with vector_cosine_ops; L2 distance uses <-> and vector_l2_ops, while negative inner product uses <#> and vector_ip_ops. The PostgreSQL index types reference explains the general role of index access methods; pgvector documents the vector-specific operators and operator classes.
An index is not a guarantee that every query will use it. PostgreSQL’s planner weighs table size, filters, and estimated costs. Check actual plans and retrieval quality with representative data rather than assuming index creation made a query faster or equally accurate.
Key Components and Design Choices
Vectors and distance. The vector dimension must match the embedding model’s output. Select one compatible embedding model and distance metric for both stored passages and queries. Changing models usually means generating new embeddings; comparing values from incompatible embedding spaces is not meaningful.
HNSW and IVFFlat tuning. HNSW’s m and ef_construction influence graph size, build work, and recall. At query time, hnsw.ef_search controls how many candidates are considered; increasing it can improve recall at additional query cost. IVFFlat’s lists determines the number of clusters and ivfflat.probes how many are searched. Higher probe counts can improve recall while reducing the speed advantage. Tune either index against measured latency and a labeled set of relevant results.
Filtering and permissions. Approximate-index filtering is applied after the index scan. If a query filters by tenant or category, a small initial candidate set may contain too few rows that survive the filter. pgvector supports iterative index scans, which can continue searching for more candidates within configured limits. Partial indexes for a few fixed values or partitioning for many values may also help. Filters are an application security boundary: apply the same authorization constraints to vector and lexical retrieval, not only after results have been combined.
Hybrid retrieval. PostgreSQL full-text search converts text to tsvector and a query to tsquery; a GIN index can find matching terms. The official PostgreSQL full-text search documentation describes this pipeline. Combining its lexical candidates with pgvector’s semantic candidates helps searches that mix conceptual questions with exact names, identifiers, or error codes. PostgreSQL’s ts_rank_cd is a text relevance function, not BM25. Score scales from different retrieval methods should not be added blindly; reciprocal rank fusion (RRF) combines ranks instead. See hybrid retrieval architectures for other fusion approaches and evaluation.
Real-World Use Cases
An internal support assistant can store article chunks, embeddings, and tenant IDs in the same database, then retrieve relevant passages with an access filter. A product catalog can use vector similarity for descriptive queries and full-text search for exact model names. A small team can add semantic lookup to an existing PostgreSQL application without deploying a separate vector database.
These designs work best when the database is already a suitable operational foundation and the team can measure the added index and query load. For broader PostgreSQL query planning and monitoring advice, see the PostgreSQL performance guide.
Getting Started: Build a Small pgvector Search
The following local setup runs PostgreSQL with pgvector in Docker. It uses three-dimensional sample vectors to demonstrate the database mechanics; those values are deliberately illustrative, not useful embeddings for a production search system.
docker run --name pgvector-demo \
-e POSTGRES_PASSWORD=postgres \
-e POSTGRES_DB=search_demo \
-p 5432:5432 \
-d pgvector/pgvector:pg17
docker exec -it pgvector-demo \
psql -U postgres -d search_demo
Create a table for document chunks. The generated tsvector column uses a fixed text-search configuration so PostgreSQL can index it:
CREATE EXTENSION IF NOT EXISTS vector;
CREATE TABLE documents (
id text PRIMARY KEY,
tenant_id text NOT NULL,
content text NOT NULL,
embedding vector(3) NOT NULL,
search_tsv tsvector GENERATED ALWAYS AS (
to_tsvector('english', content)
) STORED
);
INSERT INTO documents (id, tenant_id, content, embedding) VALUES
('guide', 'acme', 'PostgreSQL supports vector similarity search.',
'[1,0,0]'),
('index', 'acme', 'GIN indexes accelerate PostgreSQL full-text search.',
'[0.9,0.1,0]');
Add indexes for cosine-distance search and lexical matching. HNSW is convenient for a new collection because it does not need a training step. For comparison, IVFFlat should generally be created after representative rows have been loaded so its lists can be built from the data.
CREATE INDEX documents_embedding_hnsw_idx
ON documents USING hnsw (embedding vector_cosine_ops);
CREATE INDEX documents_search_tsv_idx
ON documents USING gin (search_tsv);
An IVFFlat alternative could use USING ivfflat (embedding vector_cosine_ops) WITH (lists = 100). Treat that as a starting point only: choose list and probe counts based on collection size and measured recall.
The next query retrieves candidates from both paths and combines their ranks with RRF. Each path applies the tenant condition before its results are fused. The sample query vector and search phrase are fixed for clarity; a real application supplies the embedding for the user’s query and binds both inputs safely through its database driver.
WITH params AS (
SELECT '[1,0,0]'::vector AS embedding,
websearch_to_tsquery('english', 'database search') AS terms
),
semantic AS (
SELECT d.id,
row_number() OVER (ORDER BY d.embedding <=> q.embedding) AS rank_no
FROM documents AS d
CROSS JOIN params AS q
WHERE d.tenant_id = 'acme'
ORDER BY d.embedding <=> q.embedding
LIMIT 20
),
lexical AS (
SELECT d.id,
row_number() OVER (
ORDER BY ts_rank_cd(d.search_tsv, q.terms) DESC
) AS rank_no
FROM documents AS d
CROSS JOIN params AS q
WHERE d.tenant_id = 'acme'
AND d.search_tsv @@ q.terms
ORDER BY ts_rank_cd(d.search_tsv, q.terms) DESC
LIMIT 20
),
fused AS (
SELECT id, sum(1.0 / (60 + rank_no)) AS score
FROM (
SELECT * FROM semantic
UNION ALL
SELECT * FROM lexical
) AS candidates
GROUP BY id
)
SELECT d.id, d.content, fused.score
FROM fused
JOIN documents AS d USING (id)
WHERE d.tenant_id = 'acme'
ORDER BY fused.score DESC
LIMIT 5;
Verify both paths and their plans with a query representative of the expected workload:
EXPLAIN (ANALYZE, BUFFERS)
SELECT id, content, embedding <=> '[1,0,0]'::vector AS distance
FROM documents
WHERE tenant_id = 'acme'
ORDER BY embedding <=> '[1,0,0]'::vector
LIMIT 5;
On a disposable setup, compare the results with and without the approximate index to see whether recall and latency meet the application’s needs. For filtered searches, test realistic tenant sizes and candidate limits; enable hnsw.iterative_scan or adjust hnsw.ef_search only after measuring. Keep the vector and full-text indexes maintained when source content changes, and evaluate permission checks, deletion behavior, index growth, and backup requirements before production use. The broader database indexing guide explains how indexes trade read speed for storage and write work.
Common Misconceptions
- “pgvector creates embeddings.” It stores and searches vectors; an embedding model and ingestion pipeline remain application responsibilities.
- “An approximate index returns the exact nearest rows.” HNSW and IVFFlat trade some recall for speed. Compare them with exact search on representative queries.
- “Keeping data in PostgreSQL makes filtering automatically safe.” The application still needs correct authorization predicates in every retrieval path. Index settings do not replace access-control checks.
- “Hybrid search means pgvector provides BM25.” pgvector supplies vector search; PostgreSQL full-text search supplies a separate lexical path. The application combines their ranked results.
Related Articles
- Vector Databases for RAG: How Retrieval Works explains the broader embedding and retrieval pipeline.
- Hybrid Search for RAG: Combining Vector and Sparse Retrieval compares methods for combining search results.
- PostgreSQL Performance Optimization covers query plans, indexes, and monitoring.
- Database Indexing Strategies introduces index trade-offs across database systems.
Changelog
- Initial publication.
Last updated: October 4, 2026

