Vector Databases for RAG: How Retrieval Works

Updated on
9 min read

Vector databases for retrieval-augmented generation (RAG) help an application find relevant passages in a collection of documents before an AI model answers a question. They matter to developers building assistants over private or frequently changing information, because a language model cannot reliably answer from documents it has never seen. This explainer follows the path from source files to retrieved context and shows where vector search fits alongside conventional search.

What Are Vector Databases for RAG?

A vector database stores numerical representations of data called embeddings, along with the original text and metadata needed to use the results. An embedding model maps a passage or query into a list of numbers. Texts with related meaning often have vectors that are close under a chosen distance measure, even when they use different words.

RAG combines that retrieval step with text generation: an application searches its own knowledge, supplies selected results to a language model, and asks it to answer using that context. The original RAG research describes this pattern as combining a model’s learned parameters with an external, retrievable memory. A vector database is one way to provide that memory; it does not generate answers itself.

The database also has to manage more than vectors. It needs to associate each vector with its source passage, document, access permissions, and other metadata. Systems such as pgvector add vector storage and similarity operators to PostgreSQL, while Qdrant’s documentation describes a purpose-built vector search engine. The right choice depends on the surrounding application, data volume, operations, and search requirements.

The Problem RAG Solves

A language model’s training data is not a dependable source for a company’s current handbook, product documentation, or customer-specific records. Putting an entire document collection into every prompt is usually too expensive, can exceed context limits, and may disclose irrelevant or unauthorized material. Fine-tuning can change model behavior, but it is not a convenient way to keep a model synchronized with frequently updated facts.

Retrieval narrows the material to a few passages relevant to the current question. The application can update its document index without retraining the language model, and it can attach source information so an answer can be checked. Retrieval does not guarantee correctness: if the relevant text was not indexed, ranked poorly, filtered out, or misread by the model, the generated response can still be wrong.

How Vector Retrieval Works

A RAG pipeline generally has an offline ingestion path and an online query path:

  1. Prepare documents. Extract readable text from source files, remove boilerplate, and split long documents into chunks. Chunk boundaries should preserve useful context, such as a heading with the paragraphs beneath it.
  2. Create embeddings. Send each chunk to an embedding model. Store the resulting vector with the text and metadata such as document ID, section, timestamp, and access scope.
  3. Embed a question. At query time, encode the user’s question with the compatible embedding model. The query and stored passages must use the same model and vector dimensions.
  4. Search and filter. Find nearby vectors, optionally restricting candidates by tenant, document type, date, or permissions. Similarity is a ranking signal, not proof that a passage answers the question.
  5. Build grounded context. Select a manageable set of results, include their source labels in the prompt, and ask the language model to answer from those sources. The application can return citations with the response.

For small collections, exact search compares a query with every stored vector. As the collection grows, approximate nearest-neighbor indexes reduce the number of comparisons and trade some recall for speed and lower query cost. In pgvector, for example, HNSW and IVFFlat are index options with different build, memory, and search trade-offs. The pgvector project documentation explains its distance operators and index types.

There is no single best retrieval method for every question. Dense vector search is useful for conceptual similarity, lexical search is strong when exact terms matter, and hybrid retrieval combines signals from both:

Retrieval approach What it matches Strengths Limitations Good fit
Dense vector search Semantic proximity between embeddings Can match paraphrases and related concepts May miss exact identifiers, rare names, or numbers Natural-language questions over varied wording
Lexical search (for example, BM25) Terms and their frequency in documents Strong for product names, error codes, quoted phrases, and rare keywords Synonyms and paraphrases may not match Documentation, logs, and exact-term lookup
Hybrid retrieval Dense and lexical result signals Balances semantic and exact-term matches Needs score fusion or reranking and more tuning Mixed queries over technical or domain-specific material

Hybrid search can improve coverage, but it is not automatically better. Teams need representative questions and labeled relevant passages to determine whether fusion, filters, or a reranker improve the results that matter.

For a deeper look at query construction, score fusion, and evaluation, see Hybrid Search for RAG: Combining Vector and Sparse Retrieval.

Key Components and Design Choices

Embedding model. It determines which relationships are easy to retrieve. Compare models on the language, domain, and query styles in your own corpus. When replacing a model, re-embed stored documents; vectors from different models should not be compared as if they shared the same coordinate space. OpenAI’s embeddings documentation describes embedding inputs and available models.

Chunking and metadata. Very large chunks can bury a useful detail among unrelated text; very small chunks can omit the qualifications that make a passage accurate. Keep stable source IDs and enough metadata to reconstruct context. Version or replace chunks when their source changes so stale passages do not remain searchable.

Distance metric and index. Cosine distance, Euclidean distance, and inner product are common choices, but the embedding model and database operator must agree. Approximate indexes need operational tuning: HNSW can use more memory and take longer to build, while IVFFlat requires suitable data and list configuration. Measure recall, latency, index size, and update behavior instead of choosing only by headline query speed.

Filters, reranking, and context assembly. Apply authorization constraints as part of retrieval, not after sending results to the model. A reranker can rescore a smaller candidate set with a more expensive model. Finally, trim duplicate or redundant passages and preserve source labels; filling the prompt with low-quality matches can make answers worse.

These database choices are one part of a larger application; RAG system best practices covers source freshness, chunking, prompt assembly, citations, and end-to-end evaluation.

Real-World Uses

An internal support assistant can retrieve current product manuals and incident procedures, then cite the relevant section. A legal search tool can surface clauses that express similar obligations despite different wording. An engineering assistant can search design documents and runbooks, while metadata filters keep each team’s restricted material separate. In all of these cases, the database returns candidate evidence; business rules and the language model still determine how that evidence is used.

When many requests share the same system instructions before their retrieved passages, an inference server may reuse that common prompt prefix; LLM inference optimization and KV caching explains the cache and its limits.

Getting Started with pgvector

This small example uses PostgreSQL with the pgvector extension and OpenAI embeddings. Start a local database with the pgvector Docker image, then install the Python clients:

docker run --name rag-db -e POSTGRES_PASSWORD=rag -e POSTGRES_DB=rag -p 5432:5432 -d pgvector/pgvector:pg17
python -m venv .venv

Activate the environment (.\.venv\Scripts\Activate.ps1 in Windows PowerShell, or source .venv/bin/activate on macOS or Linux), then install dependencies and set your API key:

python -m pip install openai "psycopg[binary]" pgvector
$env:OPENAI_API_KEY = "<your-api-key>"

In macOS or Linux, set OPENAI_API_KEY with export OPENAI_API_KEY="<your-api-key>" instead. Create a table whose vector dimension matches the embedding model. text-embedding-3-small returns 1,536 dimensions by default:

CREATE EXTENSION IF NOT EXISTS vector;

CREATE TABLE chunks (
  id text PRIMARY KEY,
  tenant_id text NOT NULL,
  source text NOT NULL,
  content text NOT NULL,
  embedding vector(1536) NOT NULL
);

CREATE INDEX chunks_embedding_hnsw_idx
  ON chunks USING hnsw (embedding vector_cosine_ops);

The following script embeds two example passages, stores them, and retrieves the closest results for a question. A small example does not need an approximate index for speed, but the index illustrates the production path:

from openai import OpenAI
import psycopg
from pgvector.psycopg import register_vector

client = OpenAI()
documents = [
    ("invoice-policy", "billing/retention.md", "Invoices are retained for seven years."),
    ("refund-policy", "billing/refunds.md", "Refund requests are reviewed within five business days."),
]
document_vectors = client.embeddings.create(
    model="text-embedding-3-small",
    input=[content for _, _, content in documents],
).data
question = "How long do we keep invoices?"
query_vector = client.embeddings.create(
    model="text-embedding-3-small",
    input=question,
).data[0].embedding

with psycopg.connect("postgresql://postgres:rag@localhost:5432/rag") as conn:
    register_vector(conn)
    for (doc_id, source, content), vector in zip(documents, document_vectors):
        conn.execute(
            """
            INSERT INTO chunks (id, tenant_id, source, content, embedding)
            VALUES (%s, %s, %s, %s, %s)
            ON CONFLICT (id) DO UPDATE
            SET source = EXCLUDED.source,
                content = EXCLUDED.content,
                embedding = EXCLUDED.embedding
            """,
            (doc_id, "acme", source, content, vector.embedding),
        )

    results = conn.execute(
        """
        SELECT source, content, 1 - (embedding <=> %s) AS cosine_similarity
        FROM chunks
        WHERE tenant_id = %s
        ORDER BY embedding <=> %s
        LIMIT %s
        """,
        (query_vector, "acme", query_vector, 5),
    ).fetchall()

for source, content, score in results:
    print(f"{score:.3f} {source}: {content}")

Check that the database is reachable and that the query returns the invoice passage. Before using this pattern with real users, add document ingestion and deletion, enforce authorization filters, and evaluate retrieval against questions with known relevant sources. Do not send sensitive documents to an embedding provider unless the data handling and retention terms are acceptable for your use case.

Common Misconceptions

  • “A vector database is the RAG system.” It supplies one retrieval capability. Parsing, chunking, permissions, prompt construction, generation, citations, and evaluation are separate responsibilities.
  • “The closest vector is always the right answer.” Similarity ranks candidates according to an embedding and metric. It does not establish factual correctness, authority, or freshness.
  • “A larger vector index fixes poor answers.” Bad extraction, unsuitable chunk sizes, wrong filters, mismatched embedding models, or an ineffective prompt can all cause failures that more indexed data will not fix.

Changelog

  • Initial publication.

Last updated: September 29

TBO Editorial

About the Author

TBO Editorial writes about the latest updates about products and services related to Technology, Business, Finance & Lifestyle. Do get in touch if you want to share any useful article with our community.